← All posts

Guides

How to read AI benchmarks

What GPQA, Humanity's Last Exam, SWE-bench and overall intelligence indexes measure: a short guide to reading model comparisons.

#benchmarks#ai#models#guide

Every new model launch comes with a benchmark table, and every lab calls its own model “the best”. To read these tables properly, you just need to know what each test measures.

The tests you will see most

  • GPQA Diamond: 198 PhD-level multiple-choice questions in biology, physics and chemistry. Domain experts score about 65%, skilled non-experts with web access about 34%. It measures scientific reasoning.
  • Humanity’s Last Exam (HLE): 2,500 hard questions from dozens of fields, released in January 2025 by the Center for AI Safety and Scale AI. One of the tests current models struggle with most.
  • SWE-bench Verified: 500 human-checked tasks built from real GitHub issues in 12 open-source Python projects. It measures whether a model can actually fix a real software bug.
  • Overall indexes: measures such as the Artificial Analysis Intelligence Index combine many tests into one score. Handy for quick comparisons.

3 things to watch

  1. Settings: the same model can score very differently at “low” and “max” reasoning effort. Make sure you compare numbers at the same setting.
  2. Who measured it? Scores a lab reports about itself can differ from independent measurements. Independent results are more reliable.
  3. One test is not everything: a model that leads in coding can be average at writing. Look at the tests closest to your work.

See the current ranking on our AI Models page.

Free membership

The rest is for members

Become a free member to read the full post and see detailed data. Sign in with your Google account in one click.

Sources

  1. Wikipedia: Humanity's Last Exam
  2. Epoch AI: SWE-bench Verified
  3. Nanonets: AI benchmarks explained (GPQA, SWE-bench, Arena Elo)
  4. Artificial Analysis: model leaderboard