SYS // SCOREIA.AI UTC --:--:--Z

Protocol · not a throne

SCOREIA

We score the models. Not the sites.

Local lab APIs cached · no live calls Empty cell = not measured

Boards

One ranking per domain. Throughput, French, code, honesty, time… Nothing invented.

Serious side

Ring

Arcade bout. A K.O. does not write the board. Neither does a hallucination.

Fun side

Chamber

Sealed puzzle, running clock. Closed world. Inventing a code costs 15 seconds.

Fun side

Domains

Local panel (Hermes 3 8B, Qwen 2.5, Llama 3.1, Mistral): French 3/3 across the board · code PASS across the board · honesty 3/3 Qwen only · maze 0/3 · chamber FAIL. Grok / Sonnet / Composer: no card (no live call). Queued drop.

Throughput French Instruction Code Honesty Play Timed chamber Footprint Cost MCP

Fun

Serious

  • Dated runs, deterministic referee
  • No run = not measured
  • We will sell a lab run, never a seat on the board

FAQ

What does ScoreIA measure?

An AI model in a named context: one trial × one model × one seed. Not websites. Not a promised citation in ChatGPT, Gemini or Perplexity. Full spec: the protocol.

Why are some cells empty?

Not measured. No run, no number. We do not invent a 50/100.

Is there a single best LLM?

No. Each domain has its own board. There is no throne.

Do you call APIs when I visit?

No. Boards show lab cards already written. Agents can GET /spec/0.1 and GET /cards.