Menu
A2Z GAMESPEC-BENCHAT A GLANCE

EVALUATING
DESIGN FIDELITY.

How faithfully can coding agents implement a game design specification? We examine the same requirements through source code, rendered scenes, and adaptive playtesting.

The evaluation framework
01 / MOTIVATION
DESIGN FIDELITY

Does the whole design
hold together?

Long-form GDDs connect requirements through conditions, state changes, and dependencies. A2Z GameSpec-Bench evaluates how faithfully coding agents implement these requirements and their relationships.

02 / THE BENCHMARK

WHAT TO VERIFY.
HOW TO TEST.

Dependency-aware contracts define the requirements. Test policies exercise the game to gather evidence.

Dataset

100 GDDS.
TWO SCOPES.

Small and Big differ in game scope. Both contain detailed, self-contained specifications.

Small50 GDDs

A compact core task

One main action or a small set of mechanics, with clear success, failure, and feedback.

Game design documentBloom or Weed
  1. 01
    Click or ignore

    Click an active flower. Let a weed expire without clicking.

  2. 02
    Show the outcome

    A correct click bursts into petals. A correct ignore shows a mint check.

  3. 03
    Finish the run

    At 45 seconds, discard any active icon unscored and show Results.

One repeated decision
IconClick / ignoreFeedbackAccuracy
14,085Mean tokens / GDD
54.2Mean outcome requirements
Big50 GDDs

Broader, interacting systems

Multiple systems and more content, with requirements that carry across the game.

Game design documentSSITGIM
  1. 01
    Combat and resources

    Valid bells spend spirit power and raise corruption. Damaging sword hits restore power.

  2. 02
    Corruption and control

    Maximum corruption temporarily swaps movement inputs, then attack and bell inputs.

  3. 03
    Ritual and traversal

    Deplete boss vigor, then complete the rite to unlock underwater movement.

Connected systems
Spirit powerCombatCorruptionBoss vigorRiteWaterway
26,297Mean tokens / GDD
84.0Mean outcome requirements

Per-document means from Table S3. Selected GDD requirements and relationships are summarized.

From design to evidence
Game designThe intended experience
Coding agentFrom specification to implementation
A generated game displayed on a monitor
Generated gameSource project + runnable build
What to verifyDependency-Aware
Contract
r₁r₂r₃r₄r₅ Rules · Invariants · DependenciesFixed across builds and revisions
Contract-guided evaluation
Source codeInspect the implementation
How to testTest Policies
Scenario-based
replay
Fixed inputs and timing
Adaptive
playtest
Inputs from live observations
Gather evidenceAcross three axes.
CodeFramesTraces rRequirementfeedback
Revise.
Evaluate again.
The same contract.

A coding agent turns a long-form game design document into a source project and playable build.

Dependency-aware contract

From GDD to a fixed contract.

A GDD passage becomes a rule with conditions, triggers, and effects. Shared state connects the rules.

  1. 01Bind GDD entriesShared vocabulary
  2. 02Link prerequisitesRead/write relationships
  3. 03Complete rulesExtract invariants
  4. 04Validate & freezeCompare three generations
GAME DESIGN DOCUMENT
Place Sensors, Find the Cause game screenPlace Sensors,
Find the Cause
GDD excerpts · RT-PLACE / RT-TRACE
Rule P2Place a sensor
Conditions
PLACE, an empty eligible node, fewer than two sensors, trace unused.
Trigger
Toggle the node.
Effects
Add the sensor, sort sensorNodes, and play the placement cue.
Linked to GDD row P2.

Rule dependencies

Reads fromWrites to

Placement

P1Remove sensor

Trace

T1Start trace
T2Visit relays
T3Visit output
T6Advance time

Inference

I1Correct cause
I2Wrong cause

Focus

F1Hold on blur
F3Restore focus
F5Resume trace
F6Return to title
P2 → P4 · sensorNodes

P2 updates the sensor set. P4 reads its count and requires two sensors before Run.

Invariant · per sessionAt most one traversal.

Accepted before inspecting builds.
One contract across agents and revisions.

01Source-code analysis

Follow the rules.Check the connections.

Inspect the specified conditions, triggers, and effects. Weight rule scores by their downstream reach.

Conditions · Triggers · Expected effectsEVIDENCE Code references
RULE WEIGHT0.501.00
Rule weight grows with downstream reach Illustrative dependency graph using the paper’s source-scoring weights. Charge leads to Ready. Ready and Target each lead to Attack. Attack leads to Damage and Hit VFX. The rule weights are 1.00 for Charge, 0.93 for Ready and Target, 0.84 for Attack, and 0.50 for Damage and Hit VFX. Circle area represents rule weight. These are illustrative dependencies, not measured rule scores. r₁ 1.00 Charge r₂ 0.93 Ready r₃ 0.93 Target r₄ 0.84 Attack r₅ 0.50 Damage r₆ 0.50 Hit VFX
Rule weight grows with downstream reach Illustrative dependency graph using the paper’s source-scoring weights. Charge leads to Ready. Ready and Target each lead to Attack. Attack leads to Damage and Hit VFX. The rule weights are 1.00 for Charge, 0.93 for Ready and Target, 0.84 for Attack, and 0.50 for Damage and Hit VFX. Circle area represents rule weight. These are illustrative dependencies, not measured rule scores. 1.00 Charge 0.93 Ready 0.93 Target 0.84 Attack 0.50 Damage 0.50 Hit VFX
Greater downstream reach, greater weight.
TEST POLICY

How the game is exercised.
How execution evidence is collected.

02Scenario-based replay

Replay the scenario.
Find the visual evidence.

Replay each shared scenario with its prescribed inputs and timing. Retrieve relevant frames to assess the specified visual responses.

Reeling gauges42% catch progress with A/D and J/K cues
Cloudstorm readabilityTiming gauges remain visible during the storm
Catch summaryPerfect grade, 90.3 cm, and +108 coins
Paper appendix · Selected screenshots
03Adaptive playtest

Observe. Act.
Test what comes next.

Normal play explores the game. Targeted adversarial tests supply preconditions to check rules not yet verified.

FROM THE GDD
“Destroy all seals to open the Rift Gate.”

Explore normally, then test unverified rules.

Abyssal Chain

Normal play

Try normal controls

Testing

Targeted adversarial test

Last seal weakened

Testing
Code references+Selected frames+Recorded tracesRequirement-level feedback
03 / THE RESULTS

RUNNABLE IS COMMON.
FAITHFUL IS HARDER.

Evaluate the game against its intended design. GDD Fidelity combines source-code, replay, and adaptive-playtest scores with equal weight.

Score
Game design
Coding-agent results. GDD Fidelity is on a zero to one hundred scale. Use the split buttons to compare all, Small, and Big designs.
RankCoding agentGDD Fidelity

Results from Table 2. Rankings follow the selected score and game-design split.

Source and runtime fidelity

All100 GDDs

All: source and runtime fidelityNine coding agents across 100 GDDs. Dots mark the exact scores; icon badges are offset labels. Both axes share the same scales across all three panels. Dashed contours show overall fidelity of 40, 60, and 80.20406080406080100Runtime fidelityClaude-Fable-5.1. All: Source 86.63; Runtime 72.13.5.1Claude-Opus-5. All: Source 83.16; Runtime 69.24.5Claude-Opus-4.8. All: Source 61.05; Runtime 54.65.4.8GPT-6-Astra. All: Source 79.25; Runtime 67.73.6GPT-5.6-Sol. All: Source 62.39; Runtime 61.52.5.6GPT-5.5. All: Source 68.47; Runtime 60.63.5.5GLM-5.3. All: Source 68.36; Runtime 49.26.5.3DeepSeek-V4-Pro. All: Source 64.72; Runtime 43.28.V4Kimi-K2.7. All: Source 49.06; Runtime 34.65.2.7

Small50 GDDs

Small: source and runtime fidelityNine coding agents across 50 GDDs. Dots mark the exact scores; icon badges are offset labels. Both axes share the same scales across all three panels. Dashed contours show overall fidelity of 40, 60, and 80.20406080406080100Runtime fidelityClaude-Fable-5.1. Small: Source 89.73; Runtime 79.35.5.1Claude-Opus-5. Small: Source 87.92; Runtime 77.42.5Claude-Opus-4.8. Small: Source 72.53; Runtime 65.61.4.8GPT-6-Astra. Small: Source 84.15; Runtime 75.74.6GPT-5.6-Sol. Small: Source 78.52; Runtime 69.92.5.6GPT-5.5. Small: Source 81.01; Runtime 69.15.5.5GLM-5.3. Small: Source 78.61; Runtime 61.10.5.3DeepSeek-V4-Pro. Small: Source 75.66; Runtime 60.09.V4Kimi-K2.7. Small: Source 60.81; Runtime 46.72.2.7

Big50 GDDs

Big: source and runtime fidelityNine coding agents across 50 GDDs. Dots mark the exact scores; icon badges are offset labels. Both axes share the same scales across all three panels. Dashed contours show overall fidelity of 40, 60, and 80.20406080406080100Runtime fidelityClaude-Fable-5.1. Big: Source 83.54; Runtime 64.91.5.1Claude-Opus-5. Big: Source 78.40; Runtime 61.07.5Claude-Opus-4.8. Big: Source 49.56; Runtime 43.69.4.8GPT-6-Astra. Big: Source 74.35; Runtime 59.73.6GPT-5.6-Sol. Big: Source 46.26; Runtime 53.12.5.6GPT-5.5. Big: Source 55.92; Runtime 52.11.5.5GLM-5.3. Big: Source 58.10; Runtime 37.42.5.3DeepSeek-V4-Pro. Big: Source 53.77; Runtime 26.46.V4Kimi-K2.7. Big: Source 37.31; Runtime 22.59.2.7406080
Dependency-weighted source fidelity
Source uses dependency-weighted scoring (α = 0.5). Runtime is the mean of Replay and Adaptive Playtest. All averages the Small and Big split means. Dashed lines mark overall GDD Fidelity of 40, 60, and 80.
04 / FROM EVALUATION TO REVISION

FIND THE MISMATCH.
FIX THE GAME.

Requirement-level evidence gives agents specific feedback on what to repair. The contract stays fixed while the game improves.

+10.9%

Relative GDD Fidelity gain
over self-revision

Same starting builds.
Two revision rounds.

50 Big GDDs · GPT-5.6-Sol

Self-revision67.0
Source + Replay + Playtest74.3

REVISION EXAMPLES.

Cancel treatment

Night Shift RX

GDD
REQUIREMENTSummary
WHEN

Cancel treatment with Escape, then type without selecting a patient again.

EXPECTED

The patient stays waiting and unselected.

Self-revision
Night Shift RX: Typing resumes treatment without a new selection.
Active prompt
Enlarged active prompt from the same screenshot

Typing resumes treatment without a new selection

Source + Replay + Playtest
Night Shift RX: Patient remains waiting.
Waiting symptom
Enlarged waiting symptom from the same screenshot

Patient remains waiting

Test details

Both archived art-applied builds are after two revision rounds. From the same full-ward preset, the test selects a patient, begins diagnosis, presses Escape, and types the symptom again. Self-revision reaches the prescription step without a new selection. The feedback build keeps the patient waiting. Only real keyboard inputs were used, with no state injection or game-code edits.

05 / BEYOND 2D

DESIGN FIDELITY.
BEYOND 2D.

The same requirements connect 3D gameplay to recorded test evidence.

Spin dash

REQUIREMENT

Releasing a charged spin dash launches Sonic along the aimed direction.

Satisfied

Sonic charges, launches forward, and returns to running.

06 / THE RESEARCH
A2Z GAMESPEC-BENCH

HOW FAITHFULLY CAN CODING AGENTS GENERATE GAMES FROM GAME DESIGN SPECIFICATIONS?

Seonho Lee*†1 · Wonryeol Jeong*†1 · Alberto Cereser†1 · Inha Kang‡1,2 · Hyeonjong Kim1 · Seungmin Kwak‡1,3 · Dongmin Park1

¹ KRAFTON   ² KAIST   ³ Korea National University of Arts

* Equal contribution · † Core contribution
‡ Work done during an internship at KRAFTON

Read the overview

Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naïve prompting. Game development provides a demanding testbed, as long-form Game Design Documents describe requirements that must work together across game logic, visual rendering, and player interactions. A2Z GameSpec-Bench evaluates end-to-end game development using 100 long-form GDDs. Each document is turned into a dependency-aware contract containing rules, constraints, and prerequisite relations. Source-code inspection, scenario-based replay, and adaptive playtesting connect evidence to the same requirements. Evaluations reveal gaps between runnable games and faithful implementations. Requirement-specific feedback improves GDD Fidelity by a 10.9% relative gain over self-revision after two rounds.

CITATION

@misc{lee2026a2zgamespecbench,
  title = {{A2Z GameSpec-Bench}: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?},
  author = {Lee, Seonho and Jeong, Wonryeol and Cereser, Alberto and Kang, Inha and Kim, Hyeonjong and Kwak, Seungmin and Park, Dongmin},
  year = {2026}
}

ACKNOWLEDGMENTS

We sincerely thank Kyungdo Park, Janghoon Ju, Jaeuk Kim, Inkyu Park, Myungseok Oh, Yujin Hong, and Inyoung Cho from the KRAFTON AtoZ team for helpful discussions and support throughout this project. We also thank Kangwook Lee and JuneSig Sung from KRAFTON for their support, and Jiho Choi from KAIST for valuable advice and feedback on this work.

The original website template was provided by Myungseok Oh. The A2Z team adapted the design and developed this project website.