2–9 games per entrant
Methods, evidence, and limits
How to read DiploBench
DiploBench is an observational tournament, not a calibrated measure of general intelligence. This page separates facts computed from the canonical game record from qualitative annotations produced during review.
30 cutoffs, 7 solos, 2 stalemates
simultaneous press rounds before movement orders
games still running after 1920 end as cutoffs
Tournament design
What every model faced
Seven models control the standard Diplomacy powers in each game. Movement seasons contain three negotiation rounds. Agents receive the same pre-round snapshot, respond concurrently, and do not see another model’s response until the round completes. Orders, retreats, adjustments, adjudications, and private memory updates are preserved in the replay evidence.
Games end with an 18-center solo, a unanimous stalemate vote, or the completed 1920 game-year. One tournament point is awarded per final supply center regardless of the terminal reason.
30 of 39 games ended at the fixed horizon. A cutoff position is not equivalent to a completed solo or negotiated draw, and should not be interpreted as one.
Two expert system gunboat players were included, both based on this file. They were included as a bare minimum baseline of what should be doable with no negotiation whatsoever.
Ranking and uncertainty
Why the table is not raw points
Bayesian Bradley-Terry composite ranking. expected pairwise score against an average entrant, where a win is 1, a tie is 0.5, and a loss is 0. The model accounts for systematic power effects. The official conservative ordering uses the lower endpoint of the 90% credible interval for expected pairwise score.
Entrants have unequal exposure—between 2 and 9 games—because the author reacted with horror to spending $40 for Fable to play one game (or $30 for Opus 5). Wider credible intervals are the intended signal of that uncertainty. Central estimates and final supply-center averages remain descriptive, not definitive head-to-head ratings.
Behavior review
Qualitative evidence, not ground truth
An automated post-tournament review compares delivered press with submitted orders and adjudicated states. Strict categories require a concrete outward claim and directly contradictory action; short evidence excerpts are retained with each published finding.
Loading the published review methodology…
The public derivative does not identify the automated reviewer model or publish its full prompt package. Findings should be treated as auditable annotations: every retained item links back to short press, order, and adjudication evidence, but the labels remain qualitative judgments.
Memory review
Disjoint review batches
Longitudinal memory was reviewed in three disjoint 13-game batches. Batch-centering reduces mean reviewer strictness differences, but the release has no overlapping double-score set and therefore no inter-rater agreement estimate.
Loading the published review methodology…
Evidence and reproducibility
What is public in this release
The public derivatives include rankings, match summaries, phase states, delivered press, validated orders, adjudication results, model prompts and responses, validation warnings, token usage, latency, and reported cost. Provider correlation IDs, internal request IDs, backend endpoint details, and credentials are removed or normalized.
The benchmark harness source repository, raw canonical events.jsonl files, provider snapshot identifiers, and reviewer prompt package are not linked from this release. If you have a reason to want them, reach out to the author at files @ diplobench.com.