← GPT-6.1 Sol Guide · Updated Sept 30, 2026
AI Model Guide · Sept 2026
OpenAI's claimed results (Sept 2026 announcement), translated out of benchmark-speak. These are vendor numbers — useful for comparison, not gospel.
| Benchmark (what it tests) | Claimed result | Plain meaning |
|---|---|---|
| DeepSWE v1.1 — real coding tasks | Matches Astra at ~1/5 cost; +6.4pp over GPT-6 Sol | Codes about as well as the flagship, much cheaper |
| GDP.pdf — answers from complex PDFs | Beats Opus 5.5 at <½ cost; near Astra at ~1/5 cost | Strong at reading messy business documents |
| AutomationBench — multi-step business workflows | +2.2pp over Opus 5.5 at ~⅓ cost; +4.8pp over GPT-6 Sol | Good office-automation agent material |
| OSWorld 2.0 — using computer apps | +7pp over GPT-6 Sol; within 2.1pp of Astra at ~1/7 cost | Noticeably better at clicking around software |
| Terminal-Bench Science — data analysis, simulations | 2× GPT-6 Sol's score at <½ cost ($5.47 vs $23+ per task) | Cheap science grunt work |
| Factuality — share of answers with an error (hard prompts) | Errors 11.4% → 7.7% at low effort (−32%) | Fewer hallucinations, still not zero |
Sources: OpenAI announcement with links to each benchmark's own page. See also: Sol vs Luna vs Astra.