Tony Nguyen · research note · 2026-08-03

Great on paper. Great for your task?

2026-08-03Related: Paid subscriptions for AI agentsČíst česky

A leaderboard tells you who won its measurement. Not yours. GPT-5 shows why: near the top of capability benchmarks, yet significantly lower in blind human preference comparisons — a split documented by Scale AI itself (a company whose neutrality is a story of its own). Before you bet on a model, find out who measures it. When I compared paid subscriptions for coding agents I left this question deliberately open. This is not a ranking. It’s a map — the four big players that measure and rank models, and what exactly they measure. None of them is “the right one”. Each measures something different, differently, for a different purpose.

Four giants, four different truths

  1. Artificial AnalysisOpen

    What it measures: quality (Intelligence Index) · price $/1M tokens · speed & latency

    Independent analysis of models and API providers — quality, price, speed. I already pointed to it as the quality source in my first note.

  2. ArenaOpen

    What it measures: crowdsourced blind battles → Elo/Bradley-Terry rating

    Formerly Chatbot Arena / LMArena out of Berkeley (2023), now a $1.7B company (Jan 2026). Coding leaderboard as of 2026-08-01: 1.56M votes across 379 models. Now sells “AI Evaluations” to vendors.

  3. Vals AIOpen

    What it measures: enterprise evals · composite Vals Index (finance, coding, legal…)

    A startup selling evaluations for enterprises. Their Vals Index as of 2026-08-01 is topped by Claude Fable 5 at 75.14%.

  4. Epoch AIOpen

    What it measures: model database · training compute · FrontierMath (v2, 338 problems)

    Self-described independent nonprofit. Notable AI Models (3,500+ models since 1950), compute estimates, the FrontierMath math benchmark. Funded by grants and commissioned evals (OpenAI, DeepMind, xAI).

SWE-bench, Terminal-Bench, SEAL, HELM, LiveBench and the OpenRouter rankings are deliberately missing. Those are individual benchmarks and market signals more than aggregators — and, more importantly, they are the subject of a separate note on coding-agent benchmarks I am preparing.

The rule is simple: a leaderboard is always someone’s measurement. Before you trust a number, find out who measures it, how, and what they get out of it.

Facts checked against each player’s public pages as of 2026-08-03.