Menu

A GAME IN ONE PROMPT.
BUT IS IT FAITHFUL?

A2Z GameSpec-Bench measures how faithfully coding agents implement
the rules, visuals, and interactions in game design documents.

The benchmark at a glance

Seonho Lee*†1Wonryeol Jeong*†1Alberto Cereser†1Inha Kang‡1,2Hyeonjong Kim1Seungmin Kwak‡1,3Dongmin Park1

¹ KRAFTON   ² KAIST   ³ Korea National University of Arts

FROM SPECIFICATIONTO THE INTENDED GAME.

Coding agents must preserve the designer’s intent
across the rules, visuals, and interactions
of a complete game.

GAME DESIGN
DOCUMENTS.

Long-form specifications, with requirements that work together. We measure how faithfully coding agents implement them.

GAME DESIGN DOCUMENT

Cloud Cast

Game capture from Cloud Cast
Cloud Cast

OUR GAMES
BUILT BY AGENTS.

* Games generated by GPT-5.6-Sol.

100Long-form game design documents
~51,433MAX Tokens
6,912Outcome requirements across all GDDs
3,240Invariants specified across all GDDs
DESIGN FIDELITY

WHAT TO VERIFY.
HOW TO TEST.

Dependency-aware contracts define the requirements. Test policies exercise the game to gather evidence.

From design to evidence
Game designThe intended experience
Coding agentFrom specification to implementation
A generated game displayed on a monitor
Generated gameSource project + runnable build
What to verifyDependency-Aware
Contract
r₁r₂r₃r₄r₅ Rules · Invariants · DependenciesFixed across builds and revisions
Contract-guided evaluation
Source codeInspect the implementation
How to testTest Policies
Scenario-based
replay
Fixed inputs and timing
Adaptive
playtest
Inputs from live observations
Gather evidenceAcross three axes.
CodeFramesTraces rRequirementfeedback
Revise.
Evaluate again.
The same contract.

A coding agent turns a long-form game design document into a source project and playable build.

SOURCE CODE
TEST POLICY
Source-code analysis

Follow the rules.Check the connections.

Inspect the specified conditions, triggers, and effects. Weight rule scores by their downstream reach.

EVIDENCE Code references
RULE WEIGHT0.501.00
Rule weight grows with downstream reach Illustrative dependency graph using the paper’s source-scoring weights. Charge leads to Ready. Ready and Target each lead to Attack. Attack leads to Damage and Hit VFX. The rule weights are 1.00 for Charge, 0.93 for Ready and Target, 0.84 for Attack, and 0.50 for Damage and Hit VFX. Circle area represents rule weight. These are illustrative dependencies, not measured rule scores. r₁ 1.00 Charge r₂ 0.93 Ready r₃ 0.93 Target r₄ 0.84 Attack r₅ 0.50 Damage r₆ 0.50 Hit VFX
Rule weight grows with downstream reach Illustrative dependency graph using the paper’s source-scoring weights. Charge leads to Ready. Ready and Target each lead to Attack. Attack leads to Damage and Hit VFX. The rule weights are 1.00 for Charge, 0.93 for Ready and Target, 0.84 for Attack, and 0.50 for Damage and Hit VFX. Circle area represents rule weight. These are illustrative dependencies, not measured rule scores. 1.00 Charge 0.93 Ready 0.93 Target 0.84 Attack 0.50 Damage 0.50 Hit VFX
Greater downstream reach, greater weight.
Explore the evaluation framework
BENCHMARK RANKINGS

THE LEADERBOARD.

Compare coding agents by overall GDD Fidelity or explore each evaluation axis.

Score
Game design
Coding-agent results. GDD Fidelity is on a zero to one hundred scale. Use the split buttons to compare all, Small, and Big designs.
RankCoding agentGDD Fidelity

Results from Table 2. Rankings follow the selected score and game-design split.

Explore the results and evaluation

FROM DESIGN
TO PLAY.

Explore the benchmark
A2Z GAMESPEC-BENCH

FROM RUNNABLE.
TO FAITHFUL.

Coding agents can generate runnable games, yet still struggle to preserve the intended rules and their dependencies.

A2Z GameSpec-Bench evaluates design fidelity across code, visuals, and play, and provides targeted feedback for more faithful game development.

At a glance
THE GAME DESIGN

Game introduction

Your goal

How it plays

A short introduction based on the game brief and design document.