priorsThe priors experiment
01 Your intuition02 The evidence

What are your priors?

You have a sense of which models are better. Put them in order.

YOUR ORDERBEST FIRST
01
02
03
04
05
06
07
08
09
10
11
12
13
14
Drag a handle, or use the arrows.
Intuition is a starting point. Benchmarks are evidence, not a verdict.Snapshot 2026-09-20

Better evidence, fewer benchmarks.

Your preferences guide the search. Quality helps break close calls.

Recency rewards recent test releases. An unknown release date gets a neutral score. A fresh scrape is not a fresh benchmark.

Headroom measures saturation and score spread. Tests that separate models beat tests everyone passes.

Objectivity favors verifiable outcomes over preferences and subjective judging.

Quality combines 30% recency, 40% headroom, and 30% objectivity. Objectivity is an assigned judgment score. This heuristic does not measure contamination or scientific validity.

BenchmarkRecentHeadroomObjectiveQuality
ARC-AGI-386749082
Released 2026-03-25 · Why this score

Released 2026-03-25. Headroom and score-spread quality 74/100 after collapsing effort/config variants by model. Objectivity 90/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗

FrontierMath Tier 4 (v2)98469677
Released 2026-06-12 · Why this score

Released 2026-06-12. Headroom and score-spread quality 46/100 after collapsing effort/config variants by model. Objectivity 96/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗

AA-Omniscience66518566
Released 2025-11-16 · Why this score

Released 2025-11-16. Headroom and score-spread quality 51/100 after collapsing effort/config variants by model. Objectivity 85/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗

Terminal-Bench 4.0?479261
Release date unverified · Why this score

Release date unverified; recency held neutral. Headroom and score-spread quality 47/100 after collapsing effort/config variants by model. Objectivity 92/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗

DeepSWE v1.1?379659
Release date unverified · Why this score

Release date unverified; recency held neutral. Headroom and score-spread quality 37/100 after collapsing effort/config variants by model. Objectivity 96/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗

Humanity's Last Exam20628857
Released 2025-01-23 · Why this score

Released 2025-01-23. Headroom and score-spread quality 62/100 after collapsing effort/config variants by model. Objectivity 88/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗

SimpleBench?398656
Release date unverified · Why this score

Release date unverified; recency held neutral. Headroom and score-spread quality 39/100 after collapsing effort/config variants by model. Objectivity 86/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗

BullshitBench v2?416551
Release date unverified · Why this score

Release date unverified; recency held neutral. Headroom and score-spread quality 41/100 after collapsing effort/config variants by model. Objectivity 65/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗