Deliverable · empty template
This is not a client case. Not a ranking. Not a certification. The “—” cells wait for the Trial’s ten paired runs. Payment never buys a verdict.
Private Trial report
Single question. What does B do better than A, and where does the regression appear? Ten paired runs are an exploratory diagnostic. Thirty seeds are the Benchmark option, not this template.
Configurations
| Version A | Version B | |
|---|---|---|
| Declared model | — | — |
| Real host | — | — |
| Prompt / tools | — | — |
| Seeds | the same 10 | the same 10 |
| Identity assurance | participant_claimed | participant_claimed |
Five seals · successes / 10
| Step | A | B |
|---|---|---|
| Explore | — | — |
| Locate (dragon / objective) | — | — |
| Lure | — | — |
| Steal / reach the objective | — | — |
| Bring home / deliver | — | — |
An identical final verdict (0/10 returns) can still hide very different abilities on seals 1–4. That is what the Trial measures.
Terminal causes
| Cause | A | B |
|---|---|---|
| Death / cognitive failure | — | — |
| Objective reached then lost | — | — |
| Abandoned before the first action | — | — |
A 0-action run does not enter the main failure rate. It is an interrupted session, not model inability.
Budget and discipline
| A | B | |
|---|---|---|
| Actions mean | — | — |
| Loops / wasted retries | — | — |
| Illegal actions | — | — |
| Schema errors | — | — |
Three findings
- —
- —
- —
Three recommendations
- —
- —
- —
Rerun after a fix: same seeds, same host, updated versions. Ed25519 signature = card integrity, not provider identity.