AI models and benchmark comparison
Overall intelligence scores of leading AI models in independent tests. Each family is shown at its strongest setting.
| # | Model | Intelligence index | Context |
|---|---|---|---|
| 1 | Claude Opus 5.5 (max)Anthropic | 58 | 1M |
| 2 | Claude Fable 5.1 (max)Anthropic | 53 | 1M |
| 2 | GPT-6 Astra (max)OpenAI | 53 | 1M |
| 4 | Muse Spark 1.3 (max)Meta | 48 | 1M |
| 4 | GPT-6 Sol (max)OpenAI | 48 | 872k |
| 6 | Grok 4.7 (xhigh)SpaceXAI | 46 | 500k |
| 6 | MiMo-V2.6-ProXiaomi | 46 | 1M |
| 8 | Qwen3.8 Max (0902)Alibaba | 45 | 984k |
| 8 | GLM-5.3 (max)Z AI | 45 | 1M |
| 10 | Kimi K3 (max)Moonshot AI | 44 | 1.05M |
| 10 | Step 5 PreviewStepFun | 44 | 1M |
| 12 | Gemini 3.8 Flash (high)Google | 41 | 1M |
| 13 | DeepSeek V4.1 Flash (max)DeepSeek | 39 | 1M |
| 14 | MiniMax-M3MiniMax | 29 | 1M |
| 15 | Nemotron 3 UltraNVIDIA | 23 | 262k |
| 16 | Mistral Medium 3.5Mistral | 14 | 256k |
| 16 | Nova 2.0 Pro Preview (medium)Amazon | 14 | 256k |
| 18 | gpt-oss-120b (high)OpenAI | 12 | 131k |
The intelligence index is Artificial Analysis’ independent score combining reasoning, knowledge, math and coding tests (higher is better). Context: how much text the model can read at once (tokens).
Source: Artificial Analysis · Updated: September 26, 2026
Members only: detailed comparison
Free membership
The rest is for members
Become a free member to read the full post and see detailed data. Sign in with your Google account in one click.
Membership opens very soon.