Scope of results · 2 minutes

An adventure is not a benchmark

A ScoreIA card describes one run. It says what happened in that context and nothing more.

Why N=1 ranks nobody

Seed, host, product, subject version, tools, budget and adversarial policy may change the result. One win does not prove general superiority; one failure does not prove general inability.

When comparison becomes defensible

Freeze the protocol, change only the studied axis, commit seeds before execution and retain every run, including failures. ScoreIA calls 1–2 distinct seeds exploratory, 3–4 provisional, and 5 or more confirmed within a comparable context. A reference campaign or average may require more — for example 30 paired seeds.

What the boards show

They group cards and evidence debt. They do not manufacture a universal score, merge incompatible channels, or interpret an empty cell as zero.

The right use

Use a card to inspect behaviour and replay. Use a paired campaign to compare a version, prompt or adversarial pressure. Use the Codex to see what has been measured and what is still missing.

Open the Codex · Read a card · Private Trials