Models in a world of consequences
UNIMATRIx
Simulation as a Benchmark.
Evaluate language models through decisions in a simulated world. Fixed, versioned recipes set the conditions; the candidate model changes. Compare the resulting performance across eight domains.
↳ How it works
A recipe defines the cases, roles, seeds, peers and scoring rules. Models face the same scenarios, and their actions produce measurable outcomes across a sequence of simulation steps.
Completed runs retain their observations, responses and events for inspection. Rankings group results by recipe revision, engine fingerprint and compute reporting track, so comparisons share the same conditions.
↳ Explore the leaderboards
Loading recipe details…
- Cases
- —
- Domains
- —
- Scripted peers
- —
Loading illustrative leaderboards…
↳ Reading the scores
Scoring and comparison notes
Scores range from 0 to 100; higher is better. The aggregate gives each domain equal weight. Within a domain, levels are equally weighted, with metrics weighted according to the recipe. Every case must finish before a real run receives an aggregate score.
Standard V1 contains 192 cases; Compact V1 contains eight and has a separate leaderboard. The bundled recipes are experimental definitions. Compact V1 uses a single seed and cannot estimate variation across seeds.
This page uses fictional models and illustrative domain scores. Its overall demo score is the arithmetic mean of the eight domain scores. It does not represent a live benchmark run.