Intermediate
What are AI Benchmarks?
How the industry measures model capability — and why scores need context.
Benchmarks are standardized test sets used to compare models: MMLU (general knowledge), HumanEval and SWE-bench (coding), MATH (math reasoning), MMMU (multimodal), and many more.
They're useful but gameable — models can be trained on benchmark-like data (contamination), and a high score on one test doesn't guarantee real-world usefulness. Newer evaluation favors agentic tasks and human preference (like LMArena).
Read benchmark claims critically: check who ran the test, on what version, and whether independent evaluations agree. A single chart from a launch blog is marketing; replicated results are evidence.
Key points
- Standardized tests: MMLU, SWE-bench, MATH, MMMU
- Useful but gameable and contamination-prone
- Human-preference evals complement static tests
- Treat launch charts as marketing until replicated
