Tony Nguyen · research note · 2026-08-03
Great on paper. Great for your task?
A leaderboard tells you who won its measurement. Not yours. GPT-5 shows why: near the top of capability benchmarks, yet significantly lower in blind human preference comparisons — a split documented by Scale AI itself (a company whose neutrality is a story of its own). Before you bet on a model, find out who measures it. When I compared paid subscriptions for coding agents I left this question deliberately open. This is not a ranking. It’s a map — the four big players that measure and rank models, and what exactly they measure. None of them is “the right one”. Each measures something different, differently, for a different purpose.
The map
Four giants, four different truths
Artificial AnalysisOpen
What it measures: quality (Intelligence Index) · price $/1M tokens · speed & latency
Independent analysis of models and API providers — quality, price, speed. I already pointed to it as the quality source in my first note.
ArenaOpen
What it measures: crowdsourced blind battles → Elo/Bradley-Terry rating
Formerly Chatbot Arena / LMArena out of Berkeley (2023), now a $1.7B company (Jan 2026). Coding leaderboard as of 2026-08-01: 1.56M votes across 379 models. Now sells “AI Evaluations” to vendors.
Vals AIOpen
What it measures: enterprise evals · composite Vals Index (finance, coding, legal…)
A startup selling evaluations for enterprises. Their Vals Index as of 2026-08-01 is topped by Claude Fable 5 at 75.14%.
Epoch AIOpen
What it measures: model database · training compute · FrontierMath (v2, 338 problems)
Self-described independent nonprofit. Notable AI Models (3,500+ models since 1950), compute estimates, the FrontierMath math benchmark. Funded by grants and commissioned evals (OpenAI, DeepMind, xAI).
What is not here
SWE-bench, Terminal-Bench, SEAL, HELM, LiveBench and the OpenRouter rankings are deliberately missing. Those are individual benchmarks and market signals more than aggregators — and, more importantly, they are the subject of a separate note on coding-agent benchmarks I am preparing.
The rule is simple: a leaderboard is always someone’s measurement. Before you trust a number, find out who measures it, how, and what they get out of it.
Facts checked against each player’s public pages as of 2026-08-03.