There is no single winner, and anyone telling you otherwise is reading one company’s marketing page.
GPT-6 Astra vs Claude, GPT-6 Astra launched 4 September 2026. Claude Fable 5.1 landed three days earlier, on 1 September. Both vendors published benchmarks. Put the two sets side by side and a clear pattern shows up — along with a methodology problem that makes the most-quoted comparison unreliable.
The Short Version
| If you need… | Pick |
|---|---|
| An agent that drives software for hours | GPT-6 Astra |
| Cybersecurity or reverse engineering | GPT-6 Astra |
| Maths and science reasoning | GPT-6 Astra |
| General knowledge work and reasoning | Claude Fable 5.1 |
| Agentic coding | Too close to call |
| Cost per task on long jobs | Claude Fable 5.1 |
Where Astra Clearly Wins
| Benchmark | Astra | Fable 5.1 | Opus 5 |
|---|---|---|---|
| ScreenSpot-Pro | 92.7% | — | 87.3% |
| AutomationBench | 41.4% | 31.4% | 26.9% |
| BenchCAD | 95.9% | 84.3% | 82.1% |
| Terminal-Bench Science | 64.6% | 52.6% | 30.0% |
| FrontierMath Tier 4 | 97.6% | 87.8% | 73.2% |
| ExploitBench | 100% | — | 70% |
| SRE-Bench | 88.0% | — | 12.5% |
The cybersecurity gap is not close. SRE-Bench measures reverse engineering binaries without source code: Astra 88%, Opus 5 12.5%. That is a different league, and it is why OpenAI classified Astra as Critical for cyber capability under its Preparedness Framework.
FrontierMath Tier 4 is nearly as stark — 97.6% against 87.8%, on a benchmark OpenAI says Astra effectively saturates.
Where Claude Wins
| Benchmark | Astra | Fable 5.1 | Opus 5 |
|---|---|---|---|
| Humanity’s Last Exam (w/ tools) | 57.2% | 65.0% | 63.6% |
| Artificial Analysis Intelligence Index | 61.2 | 65.7 | 63.1 |
| Artificial Analysis Coding Agent Index | 67.0 | — | 68.1 |
| FrontierCode 1.1 Main | 53.3% | 50.9% | 53.4% |
Humanity’s Last Exam is the clearest one — Fable 5.1 is nearly eight points ahead. And on the Artificial Analysis Intelligence Index, an independent aggregate rather than either vendor’s own eval, Astra places fourth: behind Fable 5.1, Opus 5 and Fable 5.
These figures are from OpenAI’s own comparison tables. They published the ones they lose on, which is worth acknowledging.
Coding Is Genuinely Too Close to Call
| Benchmark | Astra | Fable 5.1 | Fable 5 | Opus 5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 55.8% | 44.5% | 52.6% |
| DeepSWE v1.1 | 74.1% | 67.4% | 69.9% | 73.7% |
| FrontierCode 1.1 Extended | 64.5% | 63.6% | 64.9% | 63.6% |
| FrontierCode 1.1 Main | 53.3% | 50.9% | 53.5% | 53.4% |
| AA Coding Agent Index | 67.0 | — | 67.2 | 68.1 |
Astra takes Terminal-Bench and DeepSWE. Claude takes both FrontierCode variants and the independent aggregate. The margins are one to two points in most rows, which is noise territory.
Pick on price, speed and how it fits your workflow. The capability difference is not what will decide it.
The Comparison Problem Nobody Mentions
Here is the part that should make you sceptical of every “Astra vs Claude” chart you see this week, including the ones above.
Anthropic states plainly that Fable 5.1 was evaluated with its production safeguards enabled. Where those safeguards intervened, the model scored zero — explicitly on OSWorld 2.0 for both Fable 5.1 and Fable 5, and on AutomationBench for Fable 5. Anthropic’s own words: this “likely reduces the performance of Fable 5.1 and Fable 5 on these benchmarks.”
Meanwhile OpenAI ran Astra without production safeguards on its cybersecurity evaluations, which is how you get a 100% ExploitBench score.
So on computer-use and cyber benchmarks you are frequently comparing a model with its brakes on to one with its brakes off. That is not a like-for-like result, and neither vendor is hiding it — it is just in the footnotes rather than the headline.
Anthropic adds another wrinkle: its OSWorld 2.0 figures use the benchmark authors’ August 2026 task release, which is not directly comparable to earlier OSWorld results. That is why Fable 5.1 shows a dash rather than a score in OpenAI’s table.
And the ARC-AGI-3 Number
Astra scored 99.9% on ARC-AGI-3. Opus 5 scored 30.2%. GPT-5.6 Sol scored 7.8%.
A spread that wide is a signal to check the footnote, and the footnote says Astra ran on OpenAI’s own responses API harness with two settings changed. The ARC Prize Foundation did verify the result independently — Greg Kamradt said Astra reached “human parity” on action efficiency — so the achievement is real. But 99.9% against 7.8% across different harnesses is not the clean comparison it looks like.
Price
| GPT-6 Astra | |
|---|---|
| Input / 1M | $10 |
| Output / 1M | $50 |
| Cached input / 1M | $1 |
| Over 272K input tokens | $20 in / $75 out |
| Fast mode | 2x speed, 2x price |
Astra is roughly 2.5x what GPT-5.6 Sol cost. OpenAI’s counter-argument is cost per task rather than cost per token — it reports Astra completing OSWorld tasks in about 40 minutes where Sol took 75.
Anthropic is making the same argument from the other direction. Cognition said it is moving its Opus 5 traffic in Devin to Fable 5.1 “at a lower cost per task,” and Anthropic positions Fable 5.1 as roughly twice as fast as Opus 5 using half the tokens.
Both are true. Both are marketing. Run your own workload before believing either.
Safety Behaviour Differs Sharply
On OpenAI’s internal computer-use safety benchmark, where lower is better:
| Model | Unsafe outcomes (lower better) |
|---|---|
| GPT-6 Astra | 2.4% |
| Claude Fable 5.1 | 9.5% |
| Claude Opus 5 | 11.5% |
| Claude Fable 5 | 18.3% |
| GPT-5.6 Sol | 22.0% |
Astra leads here by a wide margin — though this is OpenAI’s own internal benchmark, so treat it accordingly.
Worth balancing against something OpenAI discloses in the same section: Astra’s written reasoning is harder to monitor than Sol’s. Fewer written steps means less to inspect. OpenAI says it takes the decline seriously.
Which Should You Actually Use?
Astra if you are building agents that operate software for long stretches, doing security work, or pushing on maths and science. The computer-use and cyber leads are large enough to be decisive.
Claude Fable 5.1 if your work is general knowledge work, reasoning, or writing. It wins the independent intelligence aggregate and Humanity’s Last Exam, and the cost-per-task story is strong.
Either, for coding. They trade rows within a couple of points. Decide on price, latency and which harness you already use.
Frequently Asked Questions
Is GPT-6 Astra better than Claude?
At computer use, cybersecurity, maths and science — yes, clearly. At general reasoning, no: Claude Fable 5.1 beats Astra on Humanity’s Last Exam and on the Artificial Analysis Intelligence Index, where Astra places fourth.
Which is better for coding, Astra or Claude?
Neither, decisively. Astra wins Terminal-Bench 4.0 and DeepSWE; Claude wins both FrontierCode variants and the Artificial Analysis Coding Agent Index. Margins are one to two points.
Why does Claude score lower on computer-use benchmarks?
Partly because Anthropic evaluated Fable 5.1 with production safeguards enabled, and tasks where those safeguards intervened scored zero. Anthropic says this likely reduces its published scores.
How much does GPT-6 Astra cost?
$10 per million input tokens, $50 per million output. Cached input is $1. Requests over 272,000 input tokens are billed at $20 and $75.
Is the 99.9% ARC-AGI-3 score real?
The ARC Prize Foundation verified it independently. But Astra ran on OpenAI’s own harness with two settings changed, so comparing it directly to competitors’ scores on different harnesses overstates the gap.
Which model is safer?
On OpenAI’s internal computer-use safety benchmark, Astra produced unsafe outcomes 2.4% of the time against Fable 5.1’s 9.5%. That is OpenAI’s own benchmark, so read it with that in mind.
Which has the bigger context window?
Astra offers 1,050,000 tokens with 128,000 output, though pricing rises past 272,000 input tokens.
The Honest Answer
The interesting thing about this generation is not that one model won. It is that the two labs have diverged. OpenAI built something that operates computers and finds zero-days. Anthropic built something that reasons and writes.
Both published benchmarks where they lose, and both footnoted their methodology. That is more transparency than these launches usually carry — and it means the useful comparison is the one you run on your own workload, not the one in either press release.
For a fuller breakdown of what Astra does, see our GPT-6 Astra explainer.
Figures sourced from OpenAI’s GPT-6 Astra announcement and Anthropic’s Claude Fable 5.1 page, including both companies’ published methodology notes.




