Benchmark comparison
Clavue official models on the same table as industry references — pick a model, feel the deal.
Internal Clavue evaluation suite (2026-07). Methodology aligns with public agent benchmarks; not an independent third-party audit. Peer figures are approximate public references for context only.


| Benchmark | Auto | Clavue | Clavue 2.1 | Clavue 2.1 Fast | Clavue 2.1 Pro | Clavue 2.1 Rev | Clavue 2.1 Search | Clavue Img | Clavue OCR | Clavue Music | Clavue TTS | DeepSeek V4 Flash | GPT-5.4 mini | Claude Sonnet 4.6 | GLM-class frontier |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SWE-Bench Pro | 49.8 | 54.1 | 57.8 | 53.2 | 59.4 | 56.0 | — | — | — | — | — | 52.6 | 47.9 | 48.3 | 51.2 |
| SWE-Bench Multilingual | 70.2 | 74.8 | 77.2 | 72.5 | 78.5 | 76.0 | — | — | — | — | — | 73.3 | 71.0 | 75.9 | 74.1 |
| Terminal-Bench v2.1 | 54.0 | 59.5 | 64.5 | 56.0 | 67.0 | 62.0 | — | — | — | — | — | 62.0 | 55.8 | 71.2 | 58.4 |
| Tau3-banking | 18.5 | 22.0 | 26.5 | 20.0 | 28.0 | 24.5 | — | — | — | — | — | 23.0 | 11.3 | 30.5 | 21.8 |
| MCP-Atlas | 58.0 | 64.2 | 70.5 | 61.0 | 72.0 | 68.0 | — | — | — | — | — | 69.0 | 55.2 | 66.7 | 62.5 |
| SkillsBench | 46.5 | 51.0 | 55.8 | 49.5 | 57.5 | 54.0 | — | — | — | — | — | 53.5 | 44.8 | 54.4 | 48.2 |
| WideSearch | 68.0 | 73.5 | 76.8 | 71.0 | 78.0 | 74.0 | — | — | — | — | — | 74.4 | 70.2 | 79.5 | 72.0 |
| BrowseComp | 64.0 | 70.5 | 75.0 | 67.5 | 76.5 | 72.0 | — | — | — | — | — | 73.2 | — | 74.0 | 68.5 |
| IFBench | 72.5 | 76.0 | 80.5 | 74.0 | 81.5 | 82.0 | — | — | — | — | — | 79.0 | 69.0 | 57.0 | 71.0 |
| SysBench | 91.2 | 93.0 | 94.2 | 92.0 | 94.8 | 93.8 | — | — | — | — | — | 93.9 | 93.3 | 94.8 | 92.4 |
| MRCR-128k | 78.0 | 86.5 | 90.2 | 82.0 | 91.5 | 89.0 | — | — | — | — | — | 88.5 | 56.1 | 92.5 | 84.0 |
| Multi-IF | 83.0 | 85.5 | 87.8 | 84.2 | 88.5 | 88.0 | — | — | — | — | — | 86.2 | 84.6 | 84.8 | 85.1 |
Balanced default routing for everyday agent work.
General product model for agents and IDE chat.
Flagship quality for hard reasoning and long agent loops.
Lower latency sibling of Clavue 2.1 for tight feedback loops.
Higher ceiling for long-horizon multi-agent work.
Review-oriented model for diffs, PRs, and verification.
Global web search specialty for agents (not free-form chat).
Official image generation.
Official document OCR.
Official song generation.
Official broadcast speech.