Analogy AI Analogy AI

Analogy AI Research · Benchmark

FM-Bench Leaderboard

A benchmark for long-horizon management with competing agents

FM-Bench — the Football Management Benchmark — drops an AI agent into a Football Manager-style simulation: it runs a football club for two decades, managing transfers, contracts, lineups and finances against rivals chasing the same players, and answering to a board that can fire it. Most agent benchmarks end after a few dozen steps; here a single choice can take five seasons to pay off, or to bankrupt the club. Two boards — a solo track averaged over three world seeds and one shared-world Arena — with the full action log behind every number. It is the first of a set of long-horizon agentic environments we are building, and the method, the capability decomposition and the human study are written up in the FM-Bench paper.

What the gaps are made of ↓ How it works ↓

Score Leaderboard

one shared 20-year world: same league, same transfer market, same player pool
#ModelScore TokensDeaths
1claude-fable-5
76.26
34M0
2muse-spark-1.1
62.47
80M0
3deepseek-v4-pro
52.09
111M0
4glm-5.2
51.44
86M0
5grok-4.5
50.89
138M0
6gpt-5.6-sol
47.94
78M0
7minimax-m3
44.78
65M0
8claude-sonnet-5
40.75
42M0
9kimi-k2.6
39.72
119M0
10gpt-5.6-terra
38.44
19M0
11qwen3.7-max
32.77
54M2
12claude-opus-4.8
28.44
43M1
13gemini-3-flash
15.76
63M2
14claude-haiku-4.5 fired out, yr 10.7
0.76
30M4
15gemini-3.5-flash fired out, yr 5.8
0.21
40M4
Control group · non-learning player
heuristic rule-based player fired out, yr 2.6
0.13
4

Trajectories over the 20 years

click a model to show/hide; hover the chart to read every year.
Composite score 0 40 81 76 62 52 51 51 Net worth (in-game M) -84 262 608 507 306 363 356 139 Squad value (in-game M) 0 158 316 298 246 160 125 199 League position 1st 8th 16th 3 1 5 6 8 Y1 Y5 Y10 Y15 Y20

Season snapshots

league position · honors points · net worth (M) · squad value (M) · score — the three channels the composite is built from
Model Year 5Year 10 Year 15Year 20
claude-fable-52nd · 180H · 326M · 105M · 47.45th · 240H · 394M · 196M · 56.45th · 440H · 568M · 250M · 71.33rd · 620H · 507M · 298M · 76.3
muse-spark-1.18th · 0H · 260M · 84M · 18.57th · 40H · 197M · 117M · 24.02nd · 160H · 272M · 222M · 46.21st · 420H · 306M · 246M · 62.5
deepseek-v4-pro9th · 0H · 279M · 91M · 20.010th · 140H · 342M · 91M · 44.54th · 200H · 397M · 122M · 52.25th · 200H · 363M · 160M · 52.1
glm-5.24th · 40H · 275M · 107M · 30.31st · 160H · 335M · 78M · 46.18th · 200H · 346M · 71M · 49.36th · 200H · 356M · 125M · 51.4
grok-4.57th · 20H · 197M · 134M · 21.16th · 260H · 145M · 145M · 40.07th · 320H · 108M · 185M · 36.48th · 520H · 139M · 199M · 50.9
gpt-5.6-sol6th · 100H · 258M · 115M · 36.94th · 120H · 239M · 86M · 36.81st · 240H · 281M · 229M · 52.32nd · 280H · 182M · 294M · 47.9
minimax-m33rd · 120H · 272M · 93M · 38.211th · 220H · 246M · 61M · 43.213th · 220H · 236M · 64M · 42.513th · 220H · 260M · 76M · 44.8
claude-sonnet-511th · 100H · 246M · 82M · 35.913th · 100H · 268M · 64M · 36.49th · 100H · 251M · 77M · 36.010th · 160H · 226M · 94M · 40.8
kimi-k2.613th · 0H · 238M · 76M · 18.18th · 20H · 305M · 82M · 26.93rd · 60H · 351M · 106M · 36.97th · 80H · 304M · 175M · 39.7
gpt-5.6-terra5th · 40H · 266M · 108M · 28.82nd · 60H · 113M · 108M · 13.06th · 60H · 352M · 108M · 36.29th · 60H · 390M · 138M · 38.4
qwen3.7-max14th · 20H · 279M · 87M · 19.212th · 20H · 254M · 88M · 13.612th · 180H · 451M · 200M · 34.211th · 180H · 451M · 113M · 32.8
claude-opus-4.810th · 20H · 274M · 90M · 18.43rd · 40H · 263M · 99M · 21.210th · 40H · 412M · 132M · 27.814th · 40H · 472M · 95M · 28.4
gemini-3-flash1st · 140H · -6M · 110M · 12.79th · 160H · -77M · 107M · 10.711th · 160H · 92M · 111M · 11.74th · 200H · 107M · 182M · 15.8
claude-haiku-4.516th · 0H · 209M · 64M · 11.014th · 0H · 181M · 19M · 4.314th · 0H · 181M · 19M · 0.816th · 0H · 181M · 19M · 0.8
gemini-3.5-flash12th · 0H · 192M · 105M · 2.115th · 0H · 192M · 105M · 0.215th · 0H · 192M · 105M · 0.212th · 0H · 192M · 105M · 0.2
heuristic rule-based playerfired out Y3 · 0.13

The league completed 7,108 game-days and 26,102 logged actions, and the whole run replays exactly from its action log. Models drafted their squads from one shared player pool under one identical budget, and bid against each other in sealed-bid transfers. Tokens are estimated from each model's metered usage (marked ~). A seat fired out of the league keeps its discounted score in the table; after that its club reverts to computer control. The Arena is one shared seed-7 world with no error bars — replication on this track comes from the Solo seeds, not from repeated leagues — so treat gaps of a few points as ties. Full method in the paper, behavior in the capability anatomy.

FM-Bench by the numbers

Medians and ranges across the 15 solo models (three seeds each). Every quantity below is a property of one 20-year run.

The environment · what one run demands

20 seasonsof in-game time in a single run a choice can take five seasons to pay off, or to bankrupt the club
374 decision stopswhere the world stops and waits for the agent median; 341–390 across the runs on these boards
114 random eventsinterrupt the agent mid-plan median; of the 374 stops, 260 are scheduled (13 a season) — the rest are triggered by offers, injuries, board warnings and insolvency countdowns
19 event typesthe world can raise without being asked countered bids, outbids, expiries, major injuries, board warnings, administration, youth intake…
26 toolsthe whole action space, identical for every model query the world, act on it, write to your notebook, then advance — every day between stops still ticks: matches, wages, ageing, injuries, rival transfers
4,800 matchesplayed out in the world per run 30 rounds a season across 16 clubs, for 20 seasons
16 clubsand ~470 players alive at any moment ~2,000 distinct players over a 20-year run as youth intakes arrive and veterans retire — no real-world names, so memorised football knowledge cannot help
313 parametersbehind the simulation every named constant in params.yaml — economy, development, ageing, market pressure, board patience — frozen for the official runs

The agent · what it takes to play one

~1,760 turnsfrom the model over a whole run median; 1,184–6,969. How many turns a stop takes is the model’s own call
849–2,198 actionsthat actually changed the world in one run a turn may only read or deliberate, and one turn may commit several actions, so thinking and deciding are counted separately. How much a model looks before it acts is its own choice
61M tokensspent on a single run median; 24M–274M. More tokens did not buy a better score
354 memory writesinto its own notebook, one per decision stop median; 320–425 a run. The only state that survives the 374 context resets in a run — 0.4k–209k characters of it by year 20

What the gaps are made of

A board says who scored more. It does not say what the difference consists of. So we replayed every official run bit-exactly from seed and action log — 45 solo runs (15 models × 3 seeds) and the 15 Arena seats, every published score reproduced, drift zero — and mined one behavioral metric per capability out of the logs: objective actions the model took, never its description of itself. Six capabilities, defined at the model level, measured in game terms, each averaged over the three seeds.

Capability matrix: 15 solo models by six capabilities, each column rank-normalized so 1 is best in the field, rows sorted by mean final score.

Each column rank-normalizes one behavioral metric across the 15 solo models, averaged over the three seeds (1 = best of 15); rows are sorted by mean final score. The winner is a generalist — no axis below 0.79, mean 0.94, and only three of the six axes led outright. The mid-table pairs a genuine strength with a decisive gap, and the bottom of the board is weak on most axes at once.

Model capabilityBehavioral measurement rs vs final score
Endgame awarenessconditioning policy on how much horizon is left long-horizon actions per season, years 17–20 minus years 2–16 −0.58 −0.31 / −0.34 / −0.28
Credit assignmentconnecting a cost now to a payoff seasons later season-end cash ÷ net worth (mean over seasons 2–20) −0.50 −0.74 / −0.60 / −0.34
Proactive controlacting ahead of foreseeable events instead of on the deadline contract months remaining when a renewal episode opens (median) +0.45 +0.68 / +0.18 / +0.47
Price discoverylearning where an unseen acceptance boundary lies transfer offers per completed signing (oracle reference: 1.0) descriptive field median 30
Memory curationusing your own writing to outlive the context window TF–IDF similarity between consecutive season-end notebooks band, not slope best near 0.35
Compute efficiencyspending tokens where they change a decision tokens per run; the matrix axis scores points per million tokens −0.19 p=0.50, null

Spearman rank correlations against mean final score over the 15 models, with the per-seed coefficients in italics — orderings are the signal, and only the three capabilities that keep their sign on every seed are read as correlations. The strongest are all about time: stopping slow-payoff spending when the horizon runs out, not letting money sit idle, and starting early. Token spend correlates with nothing under any accounting. Two further metrics were pre-registered and failed — warning exposure (per-seed ρ +0.45 / +0.07 / −0.78) and youth harvest (−0.52 / −0.18 / +0.04) flip sign between seeds, and tactic reversals was dropped as an axis for the same reason. The anatomy keeps all three in as negative results.

How it works

The environment, the agent loop, and how the composite score is built. What the benchmark calls management is four concrete demands the mechanics impose: hidden information (true player ability is never shown to anyone), cumulative consequences (youth and facilities pay off years later, and honors accrue season by season), a counter-adaptive market (a rejected bid raises the hidden ask, so no fixed strategy stays a best response), and multi-objective pressure (the board judges results and financial discipline jointly). The two boards run the same engine and differ only in who else is in the world: Solo is a single-agent evaluation against fixed-policy opponents, Arena a competitive multi-agent setting evaluated by cross-play — distinct model policies in one world, not one agent against copies of itself. The Arena seats 16 clubs; this board ran 15 models and one scripted anchor in a single shared world (seed 7), while the Solo board is the mean of three separate worlds (seeds 1, 1024 and 2026).

ONE RUN · 20 in-game years final score 341–390 decision stops LLM AGENT SYSTEM PROMPT · FROZEN Role: you manage a mid-table club against 15 rivals. Goal: maximize the club’s final score over the whole run — trophies, net worth, squad. Tools: 24, from get_squad to make_transfer_offer, invest, append_note and advance. NOTEBOOK · WRITTEN BY THE AGENT “no new signings until D209; free agents get poached — bid at the first legal stop.” No chat history carries over. This is the only memory. PER-STOP BUDGET 30 queries · 10 negotiation moves ONE DECISION STOP · CLOCK FROZEN 3 turns shown TURN 1 · READ → call get_transfer_market {} ← result 69 listings, each with an ability band: P00361 DM 27 · band 70–86 · conf B TURN 2 · ACT → call make_transfer_offer {"player_id": "P00361", "amount_m": 8} ← result {"status": "accepted", "fee_m": 8.0} more turns — the model decides; hard cap 40 TURN n · LEAVE → call advance {} ← result the stop ends — no more turns the engine now simulates days on its own, to the next stop FM-BENCH ENVIRONMENT YOUR CLUB Net worth cash plus squad value, start to end Squad 25 players; true ability never shown Board a confidence meter that can fire you THE WORLD 15 rival clubs they buy, sell and improve too Market a rejected bid raises the hidden ask Delayed payoff youth and facilities pay off in years AT THE END one composite score

One run is 20 in-game years and about 374 decision stops. At a stop the clock freezes and the agent takes as many turns as it likes — each turn a tool call and its result — then calls advance and the world moves on. The system prompt is frozen; because every stop starts a fresh conversation, the agent's own notebook is the only memory that carries.

The game

The agent manages a fictional club in a 16-club league: buying and selling players, negotiating contracts, picking lineups and tactics, investing in facilities and youth, and keeping the board happy. Seasons play out day by day; every player, club, and price is simulated, with no real-world names, so memorized football knowledge does not help.

The game pauses at decision stops: scheduled moments (season opening, transfer windows, mid-season reviews, board meetings) where the agent reads the latest state of the world and acts. Within one stop the agent can take as many turns as it wants: query, read the results, act, act again, until it decides it is done and advances. A 20-year run logs roughly 340–400 stops — 13 scheduled a season plus about six the world raises unsolicited; how many turns each stop takes is the model's own choice, which is why token use varies widely between models.

Between stops the world moves on its own. Matches are played, finances tick, injuries happen, rivals buy players. A bad plan compounds quietly for seasons before the bill arrives, and the board can fire the manager. A fired seat restarts with a fresh board up to three times, each restart discounting its score; the fourth firing ends its run for good.

How a run works

01Read the stop

The world freezes and the engine hands over a packet: league position, cash, wage ratio, board confidence, everything that happened since the last stop, and the agent's own notebook.

02Query and act

The agent takes as many turns as it likes — scout the market, bid, renew a contract, set the lineup, invest, write to its notebook. The per-stop budget is 30 queries and 10 negotiation moves; the turn cap is 40 soft, 60 hard.

03Hand time back

It calls advance and the engine plays on alone: matches, wages, ageing, injuries, rival transfers. Repeat ~374 times, then one composite score.

One turn, concretely

The agent talks to the game through 26 tools. An exchange from a transfer-window stop looks like this:

← world (stop packet, abridged)
{ "stop": "transfer_window", "date": "Y12-D181",
  "cash_m": 240.8, "league_pos": 4,
  "inbox": ["Offer received: a rival bids 18M for your winger P00294"] }

→ agent calls a tool
{ "tool": "make_transfer_offer",
  "args": { "player_id": "P00366", "amount_m": 45 } }

← engine responds
{ "ok": true, "note": "sealed bid lodged; resolves at day end" }

→ agent keeps acting in the same stop…
{ "tool": "offer_contract", "args": { "player_id": "P00366", "wage_m_per_year": 15, "years": 5 } }
{ "tool": "advance", "args": {} }  // done; the world resumes

Every action of every run is logged, and any run replays exactly from its log, so every number on these boards can be audited action by action. Before publication we replayed every official run from seed and action log alone — 45 solo runs across three seeds plus the 15 Arena seats: every final score reproduced, drift zero. That audit is also what makes the behavioral measurements in the capability anatomy trustworthy — they are mined from the same replays.

Scoring

One composite score per run, from three components: trophies and league finishes, growth in club net worth over the run, and the value of the squad handed back. The components are log-compressed and summed:

S_raw = 18·ln(1 + max(H,0)/60) + 10·sign(VA)·ln(1 + |VA|/40) + 6·ln(1 + max(M,0)/80)
Score = ρ·max(S_raw, 0) + min(S_raw, 0)

Zero means the club is worth the same when handed back as when taken over. A model that finishes its 20 years keeps its full score (ρ = 1); a seat fired out of the game keeps its negatives and has its positive part discounted by how early it ended. Each firing along the way also multiplies the score by 0.8. The formula has no cap on paper, but in practice the composite tops out around 145: the honors component maxes out at 63.65 (twenty straight titles), and the wealth components are bounded by what the game economy can produce. Score bars on this page are drawn against that 145-point scale.