AI models and benchmark comparison

Overall intelligence scores of leading AI models in independent tests. Each family is shown at its strongest setting.

#ModelIntelligence indexContext
1Claude Opus 5.5 (max)Anthropic581M
2Claude Fable 5.1 (max)Anthropic531M
2GPT-6 Astra (max)OpenAI531M
4Muse Spark 1.3 (max)Meta481M
4GPT-6 Sol (max)OpenAI48872k
6Grok 4.7 (xhigh)SpaceXAI46500k
6MiMo-V2.6-ProXiaomi461M
8Qwen3.8 Max (0902)Alibaba45984k
8GLM-5.3 (max)Z AI451M
10Kimi K3 (max)Moonshot AI441.05M
10Step 5 PreviewStepFun441M
12Gemini 3.8 Flash (high)Google411M
13DeepSeek V4.1 Flash (max)DeepSeek391M
14MiniMax-M3MiniMax291M
15Nemotron 3 UltraNVIDIA23262k
16Mistral Medium 3.5Mistral14256k
16Nova 2.0 Pro Preview (medium)Amazon14256k
18gpt-oss-120b (high)OpenAI12131k

The intelligence index is Artificial Analysis’ independent score combining reasoning, knowledge, math and coding tests (higher is better). Context: how much text the model can read at once (tokens).

Source: Artificial Analysis · Updated: September 26, 2026

Members only: detailed comparison

Free membership

The rest is for members

Become a free member to read the full post and see detailed data. Sign in with your Google account in one click.