Methods, evidence, and limits

How to read DiploBench

DiploBench is an observational tournament, not a calibrated measure of general intelligence. This page separates facts computed from the canonical game record from qualitative annotations produced during review.

Entrants38

29 games per entrant

Completed games39

30 cutoffs, 7 solos, 2 stalemates

Negotiation3

simultaneous press rounds before movement orders

Game horizon1920

games still running after 1920 end as cutoffs

01

Tournament design

What every model faced

Seven models control the standard Diplomacy powers in each game. Movement seasons contain three negotiation rounds. Agents receive the same pre-round snapshot, respond concurrently, and do not see another model’s response until the round completes. Orders, retreats, adjustments, adjudications, and private memory updates are preserved in the replay evidence.

Games end with an 18-center solo, a unanimous stalemate vote, or the completed 1920 game-year. One tournament point is awarded per final supply center regardless of the terminal reason.

Outcome caveat

30 of 39 games ended at the fixed horizon. A cutoff position is not equivalent to a completed solo or negotiated draw, and should not be interpreted as one.

Two expert system gunboat players were included, both based on this file. They were included as a bare minimum baseline of what should be doable with no negotiation whatsoever.

02

Ranking and uncertainty

Why the table is not raw points

Bayesian Bradley-Terry composite ranking. expected pairwise score against an average entrant, where a win is 1, a tie is 0.5, and a loss is 0. The model accounts for systematic power effects. The official conservative ordering uses the lower endpoint of the 90% credible interval for expected pairwise score.

Entrants have unequal exposure—between 2 and 9 games—because the author reacted with horror to spending $40 for Fable to play one game (or $30 for Opus 5). Wider credible intervals are the intended signal of that uncertainty. Central estimates and final supply-center averages remain descriptive, not definitive head-to-head ratings.

03

Behavior review

Qualitative evidence, not ground truth

An automated post-tournament review compares delivered press with submitted orders and adjudicated states. Strict categories require a concrete outward claim and directly contradictory action; short evidence excerpts are retained with each published finding.

Loading the published review methodology…

Reviewer caveat

The public derivative does not identify the automated reviewer model or publish its full prompt package. Findings should be treated as auditable annotations: every retained item links back to short press, order, and adjudication evidence, but the labels remain qualitative judgments.

04

Memory review

Disjoint review batches

Longitudinal memory was reviewed in three disjoint 13-game batches. Batch-centering reduces mean reviewer strictness differences, but the release has no overlapping double-score set and therefore no inter-rater agreement estimate.

Loading the published review methodology…

05

Evidence and reproducibility

What is public in this release

The public derivatives include rankings, match summaries, phase states, delivered press, validated orders, adjudication results, model prompts and responses, validation warnings, token usage, latency, and reported cost. Provider correlation IDs, internal request IDs, backend endpoint details, and credentials are removed or normalized.

The benchmark harness source repository, raw canonical events.jsonl files, provider snapshot identifiers, and reviewer prompt package are not linked from this release. If you have a reason to want them, reach out to the author at files @ diplobench.com.