← GPT-6.1 Sol Guide · Updated Sept 30, 2026

AI Model Guide · Sept 2026

GPT-6.1 Sol Benchmarks, Explained Simply

OpenAI's claimed results (Sept 2026 announcement), translated out of benchmark-speak. These are vendor numbers — useful for comparison, not gospel.

GPT-6.1 Sol benchmark claims vs GPT-6 Sol, Astra and others (OpenAI, Sept 2026)
Benchmark (what it tests)Claimed resultPlain meaning
DeepSWE v1.1 — real coding tasksMatches Astra at ~1/5 cost; +6.4pp over GPT-6 SolCodes about as well as the flagship, much cheaper
GDP.pdf — answers from complex PDFsBeats Opus 5.5 at <½ cost; near Astra at ~1/5 costStrong at reading messy business documents
AutomationBench — multi-step business workflows+2.2pp over Opus 5.5 at ~⅓ cost; +4.8pp over GPT-6 SolGood office-automation agent material
OSWorld 2.0 — using computer apps+7pp over GPT-6 Sol; within 2.1pp of Astra at ~1/7 costNoticeably better at clicking around software
Terminal-Bench Science — data analysis, simulations2× GPT-6 Sol's score at <½ cost ($5.47 vs $23+ per task)Cheap science grunt work
Factuality — share of answers with an error (hard prompts)Errors 11.4% → 7.7% at low effort (−32%)Fewer hallucinations, still not zero

Data sources & limitations

Sources: OpenAI announcement with links to each benchmark's own page. See also: Sol vs Luna vs Astra.