✳ ARC-AGI-386749082
Released 2026-03-25 · Why this score
Released 2026-03-25. Headroom and score-spread quality 74/100 after collapsing effort/config variants by model. Objectivity 90/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗
✳ FrontierMath Tier 4 (v2)98469677
Released 2026-06-12 · Why this score
Released 2026-06-12. Headroom and score-spread quality 46/100 after collapsing effort/config variants by model. Objectivity 96/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗
✳ AA-Omniscience66518566
Released 2025-11-16 · Why this score
Released 2025-11-16. Headroom and score-spread quality 51/100 after collapsing effort/config variants by model. Objectivity 85/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗
✳ Terminal-Bench 4.0?479261
Release date unverified · Why this score
Release date unverified; recency held neutral. Headroom and score-spread quality 47/100 after collapsing effort/config variants by model. Objectivity 92/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗
✳ DeepSWE v1.1?379659
Release date unverified · Why this score
Release date unverified; recency held neutral. Headroom and score-spread quality 37/100 after collapsing effort/config variants by model. Objectivity 96/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗
✳ Humanity's Last Exam20628857
Released 2025-01-23 · Why this score
Released 2025-01-23. Headroom and score-spread quality 62/100 after collapsing effort/config variants by model. Objectivity 88/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗
✳ SimpleBench?398656
Release date unverified · Why this score
Release date unverified; recency held neutral. Headroom and score-spread quality 39/100 after collapsing effort/config variants by model. Objectivity 86/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗
✳ BullshitBench v2?416551
Release date unverified · Why this score
Release date unverified; recency held neutral. Headroom and score-spread quality 41/100 after collapsing effort/config variants by model. Objectivity 65/100 is an assigned judgment score based on grading method. Quality is a 30% recency, 40% headroom-and-spread, 30% objectivity heuristic, not proof of validity or freedom from training contamination. Source ↗