Chat with Sol, Terra and Luna right here. Switch tiers mid-conversation, compare them against Claude and Gemini, and pay for one subscription instead of five.
The number marks the generation; the names are capability tiers that advance on their own cadence. Most people should start on Terra.
State of the art on coding, computer use and long-horizon agentic tasks — and it gets there with fewer tokens than the models it beats.
Half the flagship price while performing competitively with the older GPT-5.5. For most day-to-day work you won't notice the gap.
The fastest and most affordable of the three, and still ahead of several previous-generation frontier models. Built for volume.
Published benchmark results, including the ones where GPT-5.6 doesn't come first. Higher is better in every row.
| Benchmark | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | GPT-5.5 | Claude Opus 4.8 | Gemini 3.1 Pro |
|---|---|---|---|---|---|---|
| Terminal-Bench 2.1· CLI workflows | 88.8% | 87.4% | 84.7% | 85.6% | 78.9% | 70.7% |
| OSWorld 2.0· computer use | 62.6% | 50.2% | 45.6% | 47.5% | 54.8% | — |
| BrowseComp· agentic browsing | 90.4% | 87.5% | 83.3% | 84.4% | 84.3% | 85.9% |
| Agents' Last Exam· long-horizon work | 52.7% | 50.4% | 50.3% | 46.9% | 45.2% | 32.1% |
| SWE-Bench Pro· real codebases | 64.6% | 63.4% | 62.7% | 59.4% | 69.2% | 54.2% |
| GPQA Diamond· graduate science | 94.6% | 92.9% | 92.3% | 93.6% | 92.0% | 94.3% |
| FrontierMath Tier 1–3 | 89.0% | 84.9% | 78.6% | 85.3% | 80.0% | 59.6% |
| MMMU Pro· multimodal | 83.0% | 80.7% | 78.4% | 81.2% | — | 80.5% |
| Toolathlon· tool use | 58.0% | 53.1% | 53.4% | 55.6% | 59.9% | 48.8% |
Figures from OpenAI's GPT-5.6 announcement, July 9 2026. Note the honest picture: Claude Opus 4.8 still leads on SWE-Bench Pro and Toolathlon.
The headline change isn't just a higher score — it's reaching that score with far fewer output tokens, which lowers your cost per finished task.
The ultra setting runs four agents at once by default, splitting a demanding task across parallel workstreams to reach a stronger answer.
Given only high-level direction it produces interfaces that hold together, then inspects the rendered result and fixes visual problems before handing it back.
Reads a reference deck's layouts, typography and spacing rules, then applies them consistently to new material, including equations and financial models.
Substantially better recall across very long inputs — 90.7% on GraphWalks BFS at 256k, against 73.7% for GPT-5.5.
Ships with layered protections and a reasoning monitor. Note: safeguards are tuned conservatively, so some benign requests get caught.
The model is the same. What differs is what surrounds it.
The honest answer for most people is Terra. It costs half of Sol and performs competitively with GPT-5.5, which was a frontier model until recently. Unless your work is genuinely difficult, you will struggle to tell the difference in everyday writing, analysis and light coding.
Reach for the flagship when the task is long-horizon and expensive to get wrong: refactoring across a real codebase, driving a computer through a multi-step workflow, or research that runs for hours. The gap is widest on computer use, where Sol scores 62.6% on OSWorld 2.0 against Terra's 50.2%.
Volume work. Classification, summarising a hundred documents, first drafts you'll rewrite anyway. At $1 per million input tokens it's roughly a fifth of Sol's price, and it still scores 84.7% on Terminal-Bench 2.1.
Max gives the model more time to reason, check itself and revise. Ultra goes further, coordinating four agents in parallel by default. A practical habit: draft on Terra, and escalate to Sol with max only when the first attempt isn't good enough.
On SWE-Bench Pro, Claude Opus 4.8 scores 69.2% against Sol's 64.6%. On Toolathlon, Opus 4.8 edges ahead at 59.9% versus 58.0%. No single model wins everything — which is the whole argument for having them all in one place.