The Wider Lens logoThe Wider Lens
← All topics

Intermediate

What are AI Benchmarks?

How the industry measures model capability — and why scores need context.

Benchmarks are standardized test sets used to compare models: MMLU (general knowledge), HumanEval and SWE-bench (coding), MATH (math reasoning), MMMU (multimodal), and many more.

They're useful but gameable — models can be trained on benchmark-like data (contamination), and a high score on one test doesn't guarantee real-world usefulness. Newer evaluation favors agentic tasks and human preference (like LMArena).

Read benchmark claims critically: check who ran the test, on what version, and whether independent evaluations agree. A single chart from a launch blog is marketing; replicated results are evidence.

Key points

  • Standardized tests: MMLU, SWE-bench, MATH, MMMU
  • Useful but gameable and contamination-prone
  • Human-preference evals complement static tests
  • Treat launch charts as marketing until replicated