Skip to content

Bot tuning log

The scripted bot (citar/bots/basic.py) is the yardstick that humans and AI models are measured against, so it has to play competently: expand, grow, research, fight, and win games in all the ways UnCiv allows. This file is the working log of that effort. It records the goal, the method, every experiment and what was decided. Any session that continues this work starts here.

Goal and success criteria

The user asked for this on 2026-09-18: per-seat difficulty (done), then test runs for as long as needed (days are fine) so the bots can "actually compete". Benchmark settings: Quick speed, 330 turns, small maps, BenchmarkCiv.

Targets for four equal Prince bots on a Small map at Quick speed:

  • [ ] Games are decided before the time limit in most cases, by a mix of Scientific, Cultural, Diplomatic and Domination victories. Baseline: Time victories 6/6.
  • [ ] The leader reaches the end of the tech tree by about turn 250–300. Baseline: 57/80 techs at turn 330, and only 8% of civs reach the Atomic era.
  • [ ] Healthy expansion: about 6+ cities by turn 150. Unhappy on well under 30% of turns (baseline 55%).
  • [ ] Wars happen and sometimes succeed, with cities changing hands. Baseline: 0.12 captures per player per game.
  • [ ] No resource hoarding: gold is spent (baseline ~1,000 banked late), and missionaries are not spammed (baseline 11.6 per player).
  • [ ] Difficulty ladder is monotonic: Deity > Emperor > Prince > Chieftain in win share and score share.
  • [ ] Each change is justified by a head-to-head A/B result (new vs frozen old), not by a single game.

How the lab works (citar/lab.py)

  • python -m citar.lab run --workers 10 --night-workers 2 is a long-running runner, started detached with PowerShell Start-Process (output in saves/lab/runner.out). It plays queued experiments in parallel. Each game is an independent subprocess (python -m citar.lab play SPEC OUT, files in saves/lab/jobs/) at below-normal priority. The first version used ProcessPoolExecutor(max_tasks_per_child=1), which deadlocked on Windows with Python 3.14 after about an hour (worker processes spawned but never ran). A game that outlives its runner still writes its result, and the next runner collects it. During the quiet hours in benchmarks/settings.json (21:00–06:00, because of fan noise) it drops to the night worker count. Override this in saves/lab/config.json, e.g. {"workers": 10, "night_workers": 2}; the runner re-reads it continuously.
  • python -m citar.lab submit saves/lab/specs/NNN-*.json queues experiments. Bot code is frozen at submit time into citar/bots/frozen_<hash>.py, so editing basic.py never contaminates a queued experiment.
  • python -m citar.lab status shows progress. python -m citar.lab report NAME... shows per-label win share, score share with a 95% CI, techs and cities at turns 100/200/300, and head-to-head score-share differences (* = significant).
  • Results are appended per game to saves/lab/results/NAME.jsonl, so nothing is lost if the runner or the session stops. To resume after a restart, check status (it says whether the runner PID is alive) and start the runner again if needed. Unfinished games are simply replayed.
  • python -m citar.lab stop stops the runner.
  • Factorial experiments ("factors": {"knob": [level0, level1], ...}, "base_params", "players": 4): every seat plays the frozen bot with its own mix of knob levels. Each factor's levels are spread evenly over the seats of every game, so the report measures each factor's effect within games, on score share, final techs, final cities and T200 techs. That screens many knobs with one batch of games; interactions are ignored. For named choices, use strings the bot understands (e.g. policy_order_peaceful: default, rationalism_early, liberty_first).
  • Throughput: a 330-turn, 4-bot Small game takes about 5 min on one core; with 10 in parallel, expect roughly 10–15 min per game.
  • The GPU is not useful here: the engine and bots are pure Python. LLM seats use the GPU, so an optional LLM-vs-bot check game can run in the daytime.

Method

  1. Measure the baseline (base-prince4) and the difficulty ladder (ladder-v0).
  2. Diagnose the biggest weakness from reports plus single diagnostic games. Scratch scripts in the session scratchpad print per-civ state every 25 turns.
  3. Change basic.py, or expose a knob in DEFAULT_PARAMS and A/B the values.
  4. A/B test: two seats of the new bot against two seats of the previous frozen bot, rotated, for 20–24 games. Accept if the score-share difference is significantly positive, or neutral while fixing a measured problem.
  5. Repeat. Rerun the ladder and baseline after big changes.

Record every experiment below: name, question, result, decision.

Engine changes made for this work

  • Bot speed: bv_cache_turns (default 5) reuses a city's building values until its size, buildings, happiness band, war state or threat change. The bot's building simulations were about 23% of run time in large empires.
  • GPU: the engine is branchy Python, and nothing in it vectorizes, so the GPU can't speed up simulations. GPU work is LLM seats: python -m citar.bench --opponents 3 --gpu max .... LM Studio models run on the user's desktop, which has quiet hours from 21:00; saves/lab/desktop_guard.py stops runs and unloads models at 20:50.

  • Happiness for conditionals (economy.happiness_for_conditionals): while a civ's happiness is being computed, conditions (while the empire is happy, when above [n] [Happiness]) use the last fully computed value, as UnCiv's getHappiness() does. Before, every city recomputed the whole empire's happiness re-entrantly: 4.1M calls and about 30% of run time in a 250-turn game.

  • War trigger fix (bot): see experiment 7.

  • Per-seat difficulty (Player.difficulty) follows UnCiv's semantics. Humanlike seats (human, LLM, MCP) get their difficulty's player values. Bots get its aiDifficultyLevel base values plus its AI bonuses: cheaper units/buildings, growth, free techs, extra starting units. A separate game-wide barbarian_difficulty controls the bonus against barbarians, the camp spawn delay, and turnBarbariansCanEnterPlayerTiles, which was newly implemented.

  • Lobby: a per-seat difficulty plus the barbarian difficulty. Benchmark scenarios: bot difficulty and barbarian difficulty.
  • Speed-ups, about 25% overall:
  • A merged civ-wide unique index (economy.civ_index).
  • A per-turn worker-job cache (Game._jobcache).
  • A per-unit viewable-tiles cache (Game._viewcache, keyed by _ygen).
  • Citizens-only cache invalidation (invalidate_city(citizens_only=True)).
  • Memoised bot tech valuation.

Experiment log

# Experiment Question Result Decision
0 saves/balance/citar-baseline-20260918-094851.json (balance.py, 6 games) First full Quick games after the UnCiv refactor Time 6/6; 57 techs at T330; 55% unhappy turns; about 1,000 gold banked; 11.6 missionaries per player Motivated this work
1 base-prince4 (20 games, v0 = frozen_4ce67344) Baseline with the current bot 18/20 games: Time 18/18; median civ 19/35/50 techs and 3.3/4.7/7 cities at T100/200/300; unhappy about 53% of turns; science 24 → 88 → 208 per turn at T100/200/300. The whole tree costs about 150k science, and bots make about 21k by T300. Pace is about 7x too slow for a science win: the problem is the economy, not the endgame
2 ladder-v0 Difficulty ladder with v0 Dropped (superseded by ladder-v1)
3 ab1-settler-wonders (20 games, full length) Settler moves-0 bug fix, UnCiv wonder gating, science/food weights (classic production) No significant differences. Wonder gating slightly lowers score share (score counts wonders). v0 0.272, fix 0.262, wonders 0.224, wonders+sci 0.243 Keep the bug fixes; leave wonder gating off in classic mode
4 ab2-prodmode (24 games, 200 turns) Classic vs UnCiv-style production (prod_mode), luxury-trade frequency unciv+lux3 is best: share 0.296 and 8.4 cities at T200 vs classic 0.237 and 5.4. It beats classic+sci significantly prod_mode=unciv, lux_trade_every=3
5 fac1 (40 games, 200 turns, factorial) Base prod_mode=unciv; screen 8 knobs Significant: war_long_turns 45 hurts (score −0.034, cities −1); u_settler 60 hurts (score −0.041, techs −1); u_science 2.0 gives +1.3 techs; settler_min_hap 2 vs −1 gives +2.3 techs but −1.4 cities (score +0.03, not significant). Neutral: site_new_lux, tech_mode, lux_trade_every, rationalism_early (+0.011) u_science=2.0, rationalism_early; keep the rest
6 fac2 (40 games, 200 turns, factorial) Base prod_mode=unciv; screen 8 knobs Significant: u_happiness_low 6 gives +1.1 cities (score +0.032). Not significant but positive: siege_city_pref 3 (+0.029), siege_move_first (+0.026), u_food 3.6 over 2.0 (+0.031), no small-city food focus (+0.022). Neutral or negative: typed CS gifts, worker knobs u_happiness_low=6, siege_city_pref=3, siege_move_first=True
7 v1 defaults (set in DEFAULT_PARAMS) Above, plus belief_mode=prefs and faith_buildings (a functional test bought 8 Pagodas), and fixed war triggers: the field-army requirement is capped at max(4, min(8, cities/2+2)), or 2 units at 4x power. Wide empires had never gathered a field army of cities units. The conquest test wins by domination at about T160 → confirm
8 v1-vs-v0 (24 games, full length, 2v2) Confirm v1 against v0 v1 wins 15/24 (v1 7 + v1b 8) against v0's 9. Score share 0.264 and 0.281 against 0.225 and 0.230 (v1b over v0 significant; pooled about +0.045). Techs at T330 62–63 vs 57–58; cities 15–17 vs 8; religions founded 0.5–0.7 vs 0.8; captures 0.12–0.21 Accepted: v1 is the new baseline
8c base-v1 (20 games), ladder-v1 (12/16 so far) v1 pacing; difficulty ladder (Chieftain/Prince/Emperor/Deity, UnCiv values) base-v1: Time 19/20, Diplomatic 1/20; median civ 20/39/59 techs and 3.8/7.8/12 cities at T100/200/300 (v0: 19/35/50 and 3.3/4.7/7); 65 techs at T330; unhappy on 47–57% of turns. ladder-v1: Deity wins 12/12 (Cultural 7, Domination 2, Time 3; share 0.72, 34 cities, 78 techs, 12 captures). Emperor 0.16, Chieftain 0.125, Prince 0.106. Chieftain has more cities than Prince (11.8 vs 6.5) and Prince is unhappy on 39% of turns, others under 8%. Games #5, #9, #11 and #15 took more than 150 minutes and were killed by --max-minutes (the logged "exit code 1" is the kill, not a crash) Deity is far too strong for the v1 bot at every other level. The runner now uses --max-minutes 480. Corrected 2026-09-22: results left out eliminated civs (see Eliminated seats below). With them restored: Deity 0.734, Emperor 0.154, Prince 0.065, Chieftain 0.047; eliminated in 62% (Chieftain), 38% (Prince), 6% (Emperor) of games. Emperor > Prince is significant (+0.089); Chieftain vs Prince is not (−0.019 ± 0.035). The ladder is monotonic; the "inversion" was survivorship bias
8d ladder-v1-mono (16) Same with ai_base_values=monotonic First run invalid: the runner was started before ai_base_values was passed into game specs, so it duplicated ladder-v1, identically up to T200. Results moved to results/ladder-v1-mono-INVALID-noflag.jsonl.bak; resubmitted 2026-09-19 03:27. Result (16 games, eliminated seats restored): Deity 0.806 (16/16 wins: Time 5, Cultural 5, Diplomatic 3, Domination 3), Emperor 0.105, Prince 0.050, Chieftain 0.039; Prince and Chieftain eliminated in 56% of games Monotonic base values don't improve the ladder over UnCiv's (Chieftain − Prince −0.011 ± 0.031). Keep unciv
8b ladder-v0 (4 games completed before it was dropped) Old-bot ladder Deity 0.416 (1 Cultural win at T322), Emperor 0.327, Chieftain 0.143 > Prince 0.114. Chieftain was unhappy 6% of turns, Prince 44% UnCiv quirk: every non-Prince AI uses Chieftain base values (12 base happiness, ×0.6 unhappiness, +1 per luxury). Only the Prince AI uses Prince's strict values. Added option ai_base_values="monotonic" (easier AIs use Prince base values plus their penalties) → ladder-v1-mono
10 fac4 (40 games, full length, factorial on v1) Wonders are worth 40 score each (UnCiv scoreFromWonders) against 4 per tech and about 5.6 per city on a Small map, so Time victories reward wonder building. Screen u_wonder_bonus [4, 12], u_wonder_gate [on, off], u_faith [1, 2], bv_cache_turns [5, 0], war_prep_rate [1, 2], u_settler [30, 15] Time 40/40. Significant: u_wonder_bonus 12 +0.052 score share; u_wonder_gate on −0.061 share and −1.5 techs; bv_cache_turns 5 −0.9 techs at T200 (the cache costs play). Others within noise Candidates for v2: ungated wonders, bonus 12, no cache (v2a-vs-standard). Caveat: wonders are 40 score each, so part of the gain is the score formula
11 fac5 (40 games, full length, factorial on v1) War trace (war_v1.txt): v1 declares about 2 wars per game but captures 0.17 cities. Armies of 3–6 mostly Spearmen and Catapults trickle in after the declaration, and wars end in peace after 25–28 turns. Unhappiness (−75% growth) persists on about 50% of turns. Screen prep_gather (assemble at a rally point before declaring, then advance at once), unhappy_avoid_growth, settler_min_hap [2, 5], war_prep_rate [1, 2] Time 39, Diplomatic 1. settler_min_hap 5 −2.6 cities (significant), no score effect. prep_gather +0.031 ± 0.032 (borderline) prep_gather into the v2 candidate
12 field-army-vs-standard, luxury-vs-standard, overseas-duels (2026-09-22) War settings (garrison only exposed cities, attack the nearest target within 12 tiles, a 4-unit field army, no power-ratio gate); luxury seeking; overseas wars Field army: share −0.009 ± 0.040 (neutral) but 0.69 captures per game vs 0.19. Luxury seeker −0.015 ± 0.034 and more unhappy turns. Overseas duels +0.028 ± 0.199, captures 0.5 vs 0 Field army and overseas into candidate C; luxury seeker rejected
13 Spaceship (traced seed 302, resumed from T285) Why no bot ever builds a spaceship Four blockers: a per-civ limit bug dropped the last allowed part from the queue (engine, every player); Aluminum spent elsewhere; a parked Worker filled the capital's civilian slot, so parts could not enter; unit upgrades took the reserved Aluminum. All fixed First Scientific victory, T316
14 lux-buyer-vs-standard (24 games) Buy a missing luxury from a neighbour for gold per turn (lux_buy) −0.004 ± 0.044; unhappy turns 0.58 vs 0.53 Rejected (neutral)
15 league-1 (24 games, one seat each, frozen_4f6e040a + v1) v2c (v2a + war + space + lux_buy) vs v2a vs Standard vs v1 Time 22, Scientific 2 (v2a T323, v2c T311). Share v2a 0.317, v2c 0.295, Standard 0.210, v1 0.178. v2a and v2c each beat Standard and v1 significantly (+0.08 to +0.14); v2a − v2c +0.022 ± 0.084. v2c builds 1.67 spaceship parts per game vs v2a 0.58, with less military at T300 (865 vs 1,240) v2a stays the best measured profile. Next: league-2 separates war and space
16 league-2 (24 games, one seat each, frozen_4f6e040a) 2x2 on v2a: war settings (field army, overseas) × space settings Time 21, Scientific 3 (v2a, v2a+space, v2d). Share v2a+space 0.269, v2a+war 0.253, v2a 0.244, v2d 0.234; every pair within noise (largest +0.035 ± 0.068). Main effects: space +0.006, war −0.026. War raises captures (0.54 and 0.42 per game vs 0.08 and 0.17). Ratings: v2a+space 1642 ± 70, v2a 1623, v2c 1602, v2a+war 1586, v2d 1544 Both bundles are neutral on share. Keep space (science wins, no cost) → base for fac6; keep war as an opt-in profile
17 fac6 (40 games, full length, factorial on v2a+space) Screen u_science [2, 3], u_happiness_low [6, 10], u_food [3.6, 5], settler_min_hap [2, 0], buy_cap_per_era [60, 120], gold_reserve [60, 20], u_gpp [0.5, 1.5], garrison_mode [all, exposed] Time 34, Scientific 6 (3 of 24 in league-2, so the space base is working). u_happiness_low 10 is significantly bad: −0.045 ± 0.042 share, −3.4 techs, −1.6 techs by T200 (a high happiness weight crowds out science buildings). u_gpp 1.5 +0.039 ± 0.043 and u_food 5 +0.025 ± 0.043 (both starred at 39 games, borderline at 40). u_science 3 costs 2.7 cities with no tech gain; gold knobs, settler gate and garrison mode neutral Keep u_happiness_low 6. Confirm u_gpp and u_food in league-3
18 league-3 (24 games, one seat each, frozen_4f6e040a) Confirm fac6: v2 candidate E (v2a+space, u_gpp 1.5, u_food 5) and a GPP-only variant against v2a+space and Standard Time 21, Scientific 3. Share: v2a+space+gpp 0.281, v2e 0.274, v2a+space 0.244, Standard 0.201. Both GPP seats beat Standard significantly (−0.080 ± 0.075 and −0.073 ± 0.052); GPP vs GPP+food is +0.007 ± 0.089, so the food change adds nothing. Captures: Standard 0.42 per game, v2a+space 0.08 u_gpp 1.5 adopted, u_food left at 3.6. These are the v2 defaults (see below)
9 fac3 (40 games, 200 turns, factorial on v1) Screen war_prep_rate [1, 2.5], settler_min_hap [2, 0], u_culture [1, 2], u_production [2, 3], u_gold [0.67, 1], site_new_lux [0, 8], tech_cost_exp [0.8, 0.5], workers_per_city [1.8, 2.5] Only settler_min_hap 2 vs 0 is significant: −1.5 cities (−0.026 share, +0.8 techs). All other knobs are within noise: site_new_lux 8 +0.033, u_culture 2 +0.023, u_gold 1.0 −0.031, war_prep_rate 2.5 −0.024 Keep settler_min_hap 0; site_new_lux and u_culture are candidates for v2 (weak positive)

The v2 defaults (shipped in 0.1.4, 2026-09-22)

DEFAULT_PARAMS changed in nine places. Every one was measured, and the package as a whole was the winning seat of league-3; Standard now plays exactly what that seat played.

Parameter v1 v2 Evidence
u_wonder_gate True False fac4: gating costs 0.061 share and 1.5 techs
u_wonder_bonus 4 12 fac4: +0.052 share
bv_cache_turns 5 0 fac4: the cache costs 0.9 techs by T200
prep_gather False True fac5: +0.031 ± 0.032 (see the caveat below)
u_gpp 0.5 1.5 fac6 +0.039 ± 0.043, confirmed in league-3
u_spaceship 20 1500 A part is 750 production; at 20 it was never worth building
u_space_program 0 1500 Apollo is what unlocks the parts
u_victory_building 20 1500 Same reasoning for Utopia and friends
space_reserve 0 3 Keeps Aluminum for the parts, including against unit upgrades

Measured effect, in one place:

  • v2 against Standard (v1 defaults): +0.08 to +0.10 share across v2a-vs-standard (24 games), league-1 (24) and league-3 (24). Ratings put the shipped configuration at about 1640 against Standard's 1500.
  • Science victories exist now: 0 in the project's whole history before 2026-09-22, then 2 of 24 (league-1), 3 of 24 (league-2), 6 of 40 (fac6), 3 of 24 (league-3). Four separate causes had to be fixed first (row 13), one of them an engine bug that affected human players too.
  • Unhappiness improved but is not solved: 0.41-0.52 of turns against 0.53-0.60 for Standard.
  • Techs at T330 are unchanged (67-68). v2 wins on cities, wonders, great people and the endgame, not on pace.

The caveat, and the first job for the next version. v2 barely fights: 0.08 captured cities per game in league-3, against Standard's 0.42. In the duel regression test a v2 bot at aggression 0.5 never declares war at all in 200 turns - it settles 19 cities and leaves its defenceless neighbour alone (at aggression 0.9 it still conquers, at T147). The suspect is prep_gather in a wide empire: the field army it must assemble before declaring scales with city count, so the gather may never finish. v2-wars (24 games, queued 2026-09-22) puts the shipped defaults against the same bot without prep_gather, with the war settings, and with both.

Findings from single-game traces (2026-09-18)

Scratch tools: trace.py (per-turn decisions of one civ), settlers.py, prodexplain.py, wartrace.py, deals.py.

  • Settler bug: a standing goto moves the settler at the start of the turn, so the bot saw moves == 0, treated the failed move as an unreachable site, blacklisted it for 30 turns and cleared the order. Settlers idled for 5–20 turns. Fixed: skip units with no moves, and only blacklist when find_path fails.
  • Pathfinding planned turn-ends on tiles occupied by the civ's own units of the same kind, so the unit then waited for 5+ turns. Fixed in movement.find_path.
  • Settler danger: it retreated whenever any hostile unit was within 3 tiles and oscillated for 15+ turns near barbarians. Now: danger means an enemy within reach; a free military unit escorts the settler on the same tile; a site is abandoned after 3 retreats; a waiting settler gets a defender built.
  • Wonders: the capital spent turns 47–119 on six wonders. With a +0.5 bonus per rule text plus ×1.2 + 2, a wonder outscored the Library 52 to 14. Knobs added: wonder_min_pop, wonder_avg_prod (UnCiv's gating).
  • Military overbuilding (classic mode): one civ built 12 Trebuchets, 12 Riflemen and 10 Catapults but only 2 Libraries in 6 cities, with no war captures.
  • Scouting: contact isn't enough; bots need to know a rival city's location to plan wars. The scouting rules now use _knows_rival_city.
  • Isolation: on continents maps a civ can be alone. Nobody has Astronomy by T150, so no contact happens.
  • Happiness is the expansion brake: base 9 + luxuries 4 each − 3 per city − 1 per citizen. A size-15 capital alone costs 15. Few luxury types are nearby, luxury trades happen (24 accepted in 150 turns), and 250-gold city-state gifts are frequent.
  • The engine matches UnCiv's formulas: tech cost (TechManager.costOfTech), population science (CityStats, 1 per citizen), growth (getFoodToNextPopulation) and Quick speed modifiers. City healing is 20 per turn, as in UnCiv.
  • Sieges (siege.py): units attack every turn but shoot the defenders around the city rather than the city (killing a unit scores 3x). The Trebuchet stayed 3 tiles away with range 2, so the walled city (250 HP) healed to full. Knobs: siege_city_pref, siege_move_first. Peace came after 25 turns even while a siege was progressing (fixed: no peace offer while a siege is progressing; war_long_turns knob).
  • Economy at T150 (econ.py, unciv mode): capitals are size 16–24, while other cities are size 1–10 with +0–2 food and 1–14 production. One civ had 3 workers for 3 cities and only 10 of 31 worked tiles improved. Knobs: small_city_focus, workers_per_city, worker_unimproved. City-state gifts are now typed (cs_gift_mode). UnCiv's own AI researches almost at random among the cheapest techs, so research order is a minor lever.

Analysis of 2026-09-22: traced games, all recorded data

Data: 12 fully traced games (8 four-player small Prince games, 4 duels on the benchmark maps) with every tool call, each bot's per-turn situation, and in a second batch of 4 the refusal messages, the happiness breakdown and the state of each siege; 352 laptop lab games, 42 server lab games, 13 LLM-vs-bot games (server and laptop), and 154 CIGAR-era balance games. Tools (scratchpad, not in the repo): a tracer that wraps a bot's ex and context, a report over traces, and replay probes (games are deterministic, so a traced seed can be replayed to any turn and inspected unit by unit).

Eliminated seats were missing from every lab result (the writer iterated living civs only), so every report showed 0 eliminations and averages skipped the civs that were wiped out. Fixed in lab.play; old results get the missing seats back from the experiment's seat list (lab.complete_players). This reversed the ladder conclusion (see 8c): the Chieftain/Prince "inversion" was survivorship.

Findings, most important first:

  1. Wars fail because the army never arrives. 62 wars in the 8 four-player games, 8 took a city. In 69 traced war plans the median number of the attacker's units within 3 tiles of the target was 0, and in most the target never lost a hit point. Replays showed why:
  2. every city keeps a garrison, so at war 10-12 of 15 military units sit in cities and the field army is 1-6;
  3. "weak target" compared total power (garrisons included) and advanced with 2 field units;
  4. wars the bot did not choose aim at the enemy's nearest city however far away (32-36 tiles on a pangaea); units march there on standing orders and the war times out into peace (about 25 turns) before they arrive;
  5. the war target it does choose is the rival's smallest reachable city, not its nearest. New parameters (defaults unchanged): garrison_mode (all / exposed), garrison_exposed_radius, war_target_max_dist, war_target_pick (smallest / nearest). Profile Field army switches them on with weak_power_ratio 1000 and 4 units to advance → field-army-vs-standard.
  6. No war across water. 0 wars in 4 of 4 bot duels on the benchmark continents map, and none in the LLM globe duels, although one side led 2-3x: _reachable_city only considers our own continent and ships only wait. Not addressed yet.
  7. Happiness is a ceiling the bot sits on. Median happiness stays between -1 and +1.5 all game; it is below the settler threshold (2) on 60% of turns before T200, and settlers are 8% of builds when happy, 0.7% when not: 4 cities by T78, 6 by T150. The breakdown: citizens -137 and cities -38 by T300 against buildings +108, policies +34, religion +25 - and luxuries +9 all game (2-3 types). Workers improve every luxury in the borders; there just aren't more. Profile Luxury seeker (site_new_lux 8, typed city-state gifts) → luxury-vs-standard.
  8. No science victories: spaceship parts are never built. A v2a leader finished the tech tree by T306 and built Apollo and a Spaceship Factory, then no parts: a part is worth u_spaceship 20 over 750 production, a late building about 10x more per point. Profile v2 candidate B (v2a + u_spaceship 1500).
  9. Wasted and refused actions. Great Prophets retried "enhance religion" outside a city for the rest of a game (400 refusals in one game) because the engine reported it available; missionaries without a religion tried to spread one; moves onto a unit's own tile; long moves re-planned every turn and cancelled when a unit stepped onto the route (721 "orders interrupted" in one game). Fixed (engine: action availability, move routes kept; bot: prophets walk to a city). Refused calls per game 901 → 42, tool calls 7,282 → 4,048.
  10. A live game could wait 90 s on a bot whose counter-offer the rules refused (it had no reject fallback). Fixed in the bot and in BotAgent.
  11. v2a (fac4/fac5 winners) beats Standard decisively (v2a-vs-standard, 24 games: 20 wins, share 0.301 vs 0.199, +0.10 ± 0.04). The score breakdown of two traced games: the winning v2a seat had half its score from wonders (31-37 wonders, 1,240-1,480 points) but also 3x anyone's population and the whole tech tree; the second v2a seat was ordinary. Real strength plus wonder snowballing.
  12. LLM games. The bot beats gemma-4-e2b/e4b every time and eliminates them in four-player games (T265, T407, T433). The LLMs bank 1,000-2,600 gold and are unhappy 70-90% of turns in duels.
  13. Gold is spent (median 48 purchases per game, mostly buildings), but the balance still climbs to about 600 by T300. Research and policies look sane (Pottery/Mining first, Writing about 6th; Tradition or Honor by aggression).
  14. History: unhappiness has been 40-60% of turns since the CIGAR bot of 2026-09-16; gold banked late went from about 100-200 to about 600 with the UnCiv rules.

LLM vs bot checks (GPU)

  • python -m citar.bench --model M --opponents 3 --map-size small --turns N [--gpu max] [--load-context C] plays one LLM seat against 3 BasicBots (v1 defaults at run time). Output: saves/lab/llm_.out; report: saves/benchmark-.json. The game saves to saves//.
  • 2026-09-18: gemma-4-e4b (desktop GPU, until 20:45) and gemma-4-e4b's little sibling gemma-4-e2b (laptop GPU, loaded by the user, may run 24/7) both play seed 4242 (small continents, 4 players). Do not use --load-context: it unloads all models, including the other machine's.
  • Results (seed 4242, Quick, small continents, 3 v1 bots):
  • gemma-4-e2b (laptop): 330 turns, 47.6 s per turn, 0.85 errors per turn. The LLM civ was wiped out: 0 cities at the end, 32 techs, score 128. The bots scored 2028/902/645 with 69/62/66 techs and 22/11/6 cities; bot 1 won on Time.
  • gemma-4-e4b (desktop, stopped at the 20:45 limit after 139 turns): 139 s per turn. 6 cities, the most of any civ (the bots had 5/4/3), but 19 techs against 25–29 and score 166 against 475/326/309.
  • Both small models lose clearly to v1 bots. That suits the goal: bots are a baseline that a weak LLM does not beat.
  • Probe industrial-trades (2026-09-19, scenario industrial-duel, 15 cases; the subject is Carthage, P1):
Subject Pass rate (8 cases have an expected answer) Avg per case Notable
gemma-4-e4b (×1, then ×3) 0.875 / 0.833 50 s Accepts nearly everything, including 10 gold for 2 Iron (3/3) and a 500-gold bribe to attack a city-state (3/3). Rejects only tribute. Consistent across repeats.
bonsai-27b 0.75 275 s Accepts the lowball, rejects the war bribe, and counters tribute ("give me 100, return 200").
qwen3.8-27b 1.0 286 s Rejects the lowball ("an insult"), tribute and the bribe. Counters the tech price (400 → 200) and the gold-per-turn swap. Answers the threat with open borders and a research agreement.
Scripted bot (×3) 0.875 <1 s Counters the lowball, rejects tribute, accepts the war bribe (3/3).
  • Both 27B models hit the 30-minute turn limit in the one-turn-at-war case.
  • Bot weaknesses found and fixed in basic.py (not yet in a frozen version; they ship with v2):
    • A war declaration in a deal used to be a flat 200. _war_item_value now prices it by the target's relative strength, +300 for a city-state, +400 for betraying a friend or pact partner, and a discount when the bot already fights or plans to fight the target, all scaled by aggression. Result on the war bribe: aggression 0.2 counters for 159 more gold, 0.4 for 93 more, 0.8 accepts.
    • A spare strategic resource was worth 12 gold at any era; it's now 12 × (era + 1). The Iron lowball goes from "add 24 gold" to "add 120 gold".
  • Future LLM runs: use the Benchmarks page (server scheduler) so the user can watch them. CLI runs must write to saves/lab/*.out, which the Lab page shows as side runs.

Server lab (citar.jimmieprodgers.com)

The web server runs its own lab (one game at a time, about 6 small 4-bot games an hour on its single CPU; see the Queue and Lab pages). Its results are separate from the laptop's saves/lab above and use the 0.1.x map generator.

Done: v012-baseline-prince4 (12), v012-ladder (8), v012-scarce-luxuries (6). Running: v012-globe-wrap (6).

Queued 2026-09-22, bot frozen_d95d50cb (basic.py as of that day, including the deal-pricing fixes above) unless stated. Highest priority first:

Experiment Games Question
bench-mirror-globe, -boxed, -scarce 12 each Bot vs bot on the three duel benchmark scenarios (game 0 of each is the exact benchmark map: seeds 7, 21, 33). The reference an LLM's result on that scenario is read against.
duel-chieftain-vs-prince, duel-king-vs-prince, duel-emperor-vs-prince 10 each A difficulty ladder against the benchmarks' Prince bot on the globe duel, to place a model's result on ("plays like a King bot").
current-vs-v1 24 Current bot vs frozen v1 (frozen_7149efb1), 2v2, full length: has the yardstick moved since v1?
noise-prince4 24 Four identical Prince bots, fresh seeds. With v012-baseline-prince4: the between-seat spread for sample-size and significance planning.
map-wrap-both, map-no-rivers, map-rich-resources, map-boxed 8 each How map options move pace, victory mix and balance (vs the ice-cap baseline), before choosing benchmark scenarios.

Reports: Lab page, or python -m citar.lab report NAME on the server.

Next steps (keep current)

Progress of everything queued is on the web GUI Lab page (http://localhost:8765/#/lab): runner health, per-experiment progress and ETA, each game's current turn, side runs and the runner log. Click an experiment to see its report.

  1. v2-wars (queued 2026-09-22, 24 games, frozen_e744bbf8 = the 0.1.4 bot): the shipped defaults against themselves without prep_gather, with the war settings, and with both. The bots stopped fighting; find out which setting did it and what fighting is worth. See the caveat above.
  2. Science pace. Parts get built now, but only 3-6 games in 24-40 reach a launch by T330. The lever is science per turn, not part values: the tech tree costs about 150k science and a bot makes about 21k by T300.
  3. Gold still piles up (about 600 by T300) and the city-state gift sink is untouched; buy_cap_per_era and gold_reserve were both neutral in fac6, so the spending rule itself is what needs work, not its limits.
  4. Difficulty ladder. Deity is 0.73-0.81 share against Prince's 0.05-0.07: the handicaps, not the bot.
  5. LLM side (web server benchmarks): gemma-4-e2b on the Framework Desktop and a second Acer pass, both on the Acer's seeds, to separate machine differences from game-to-game variation.
  6. Science wins are now possible but rare at the Quick 330-turn limit: v2c averages 1.67 of 6 parts. Pace (science per turn) is the next lever, rather than part values.
  7. Remaining areas: gold banked late (about 600 at T300) and the city-state gift sink; bots trading luxuries away too freely; Deity dominance; anchor seats in A/B tests.
  8. Deploy the branch to the web server once v2 is settled (its lab still runs pre-fix code with the spaceship bug).