-
Chris Alarcon - 04 Sep, 2026
GPT-6 Astra vs Claude Fable 5.1: Same Price, Two Boards Claude Still Wins
Quick answer: these two models now cost exactly the same and are good at different things. GPT-6 Astra wins computer use, math, cybersecurity and long context. Claude Fable 5.1 wins the two broadest "how smart is it" boards, including one published by OpenAI itself. Both are $10 per million input tokens and $50 per million output. The sticker price stopped being the decision. Every headline yesterday said Astra swept. OpenAI's own launch table does not say that. The rows nobody quoted OpenAI published a comparison table on the Astra launch page with Claude Fable 5.1 in a column. Two lines in that table go to Claude. Every number below is transcribed from that page, checked September 4, 2026. Where OpenAI's footnotes qualify a Claude score, that is flagged in the next section.Benchmark (OpenAI's own table) GPT-6 Astra Claude Fable 5.1Artificial Analysis Intelligence Index v4.1.1 61.2 65.7Humanity's Last Exam (with tools) 57.2% 65.0%Terminal-Bench 4.0 57.9% 55.8%Terminal-Bench Science 0.1 64.6% 52.6%FrontierMath Tier 4 (v2) 97.6% 87.8%GPQA Diamond 96.0% 93.7%AutomationBench 41.4% 31.4%DeepSWE v1.1 74.1% 67.4%FrontierCode 1.1 Main 53.3% 50.9%ARC-AGI-2 95.0% 90.0%Computer use safety, internal (lower is better) 2.4% 9.5%The Artificial Analysis Intelligence Index is the closest thing the industry has to a single general-capability number. OpenAI put it in its own launch table, and it hands Claude a 4.5 point lead. Artificial Analysis, quoted independently:Sits beside GPT-5.6 Sol in Intelligence: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback).That is not a small footnote. On the broadest board, the newest OpenAI model landed level with the previous OpenAI model. Read the footnotes before you read the table OpenAI ran these numbers. That does not make them wrong, but the footnotes change how you should read four rows.BenchCAD: OpenAI's footnote 5 says Claude's scores reflect three modifications to the eval. ScreenSpot-Pro and ExploitGym: footnote 17 says the Fable scores reported are actually from Mythos, described as "Fable with fewer safeguards." That is a different model. HealthBench Professional: footnote 11 says OpenAI independently evaluated the Claude models itself, using GPT-5.4 as grader, with Opus 5 substituted when Fable 5.1 refused. Three science evals: footnote 12 says Claude Fable 5 and 5.1 are excluded from LifeSciBench, GeneBench Pro and MedChemBench because they refuse the majority of questions.None of that is scandalous. Vendors benchmark their own launches. But "state of the art across the board" is a press summary, not what the table says. The 99.9% asterisk The number that traveled furthest was Astra saturating ARC-AGI-3 at 99.9%. It deserves the least weight of anything in the launch.Two problems, both disclosed by the people who ran it:Harness. Per the ARC Prize blog, the 99.9% came from OpenAI's custom "Provider Adapter harness" at about $19K. The default ARC-AGI harness scored 62.7% at about $26K. OpenAI's own footnote 1 confirms it ran a modified responses API harness. That is a 37 point spread depending on plumbing. No opponent. Claude Fable has no published ARC-AGI-3 result. The headline comparison of the launch is not a comparison.The Provider Adapter, in ARC Prize's words, "preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work." That is a genuinely interesting engineering result. It is not a raw intelligence score. Simon Willison, writing the day it shipped, put the whole launch in one honest sentence: Astra "appears to score higher than Fable on most of OpenAI's self-reported benchmarks." Self-reported is doing a lot of work in that sentence. Where Astra genuinely pulls ahead Strip the noise and Astra's wins are real, and they cluster. Computer use. This is the headline capability, not the benchmark. Astra scores 59.3% on Agents' Last Exam against 55.5% for Claude Opus 5, and OpenAI reports it hitting 72.6% on OSWorld 2.0 in roughly 47% less time per task than GPT-5.6 Sol. Math and science. FrontierMath Tier 4 at 97.6% against 87.8%. Terminal-Bench Science at 64.6% against 52.6%. These are not close. Cybersecurity. 100% on ExploitBench, 88.0% single-attempt on SRE-Bench binary reverse engineering. OpenAI says this crosses the Critical threshold in its own Preparedness Framework, and it has restricted the model accordingly at launch. Long context. 100% on OpenAI's eight-needle test at 256K to 512K tokens, and 96.3% at 512K to 1M. Token efficiency. This one matters more than the scores. On Agents' Last Exam, OpenAI reports Astra using about 65% fewer output tokens than Opus 5 at the highest-scoring settings. Where Claude Fable 5.1 holds The broad boards. Intelligence Index and Humanity's Last Exam, both from OpenAI's own table. Coding, closer than the headline. Astra edges Fable 5.1 on FrontierCode 1.1 Main, 53.3% to 50.9%. But look one column over in OpenAI's table: Fable 5 scores 53.5%, ahead of Astra. On the Artificial Analysis Coding Agent Index, Astra posts 67.0 against 67.2 for Fable 5 and 68.1 for Opus 5. Coding is a wash. Cost per task at equal quality. Artificial Analysis again: "Per task, the model is less than half the cost of Claude Fable 5, for the same score." That is the one place Astra's efficiency turns into a real advantage, and it is worth more than the trophy rows. What I actually threw at it on day one I pay for both stacks, and Astra is one day old as I write this.I did not run a benchmark. I ran my actual work, which is the only test I trust. The one that changed my mind about Astra was a prospecting task. I asked it to find companies in a specific niche worth reaching out to, and it went and searched forums and threads on its own without me telling it where to look. What came back was roughly a month old, where that kind of query usually surfaces threads from many months back. That is the real upgrade. Not the score, the autonomy. Then I did the thing I would actually recommend: I ran the output back through Claude Fable to QA it, Fable rewrote the prompt, and I sent it back at Astra. The results from that second pass were promising enough that it is now how I would run this by default. The honest other side: I have used Fable more, and Fable still feels more familiar and more dependable to me for building. Fable 5.1 has also done some genuinely dumb things this week, mostly losing track of tools it definitely had access to. Minor, but real. My read on the benchmark inversion: I assumed the scoreboard said Astra beat Fable outright. It does not, and that matches my hands-on impression rather than contradicting it. Fable still feels better at the deep work. Astra feels newer at getting things done by itself. The decision rule I would give someone today Do not switch providers over this. Route work instead.The job OpenDeep building, writing in your voice, long careful work Claude Fable 5.1Anything the AI should do on its own across apps and the web GPT-6 AstraResearch and volume where you need lots of runs GPT-6 AstraReviewing and QAing another model's output Claude Fable 5.1Math, science, anything with a checkable right answer GPT-6 AstraYou can only pay for one, and you are a generalist ChatGPT (usage and resets decide it, not the score)The thing I keep coming back to: smartest is no longer the flex. A model that burns tokens to win a board is worth less to me than a leaner one that finishes the task. That is the direction both companies are being pushed, and this launch is the clearest evidence of it so far. Verdict Same price, split by job, and the split is the useful output. Astra is the better agent. Fable is still the better thinker on the broadest measures OpenAI itself published, and it is my daily driver for building. If you were about to cancel Claude because of a headline, read the table first. Working out which subscription this actually affects? Which ChatGPT plans get Astra, Claude Max vs ChatGPT Pro, and Claude vs ChatGPT for the everyday call. More of how I run both at Claude at Work. Published September 4, 2026, one day after GPT-6 Astra shipped. Every benchmark figure was transcribed that day from OpenAI's launch page, and the ARC-AGI-3 harness figures from Simon Willison's write-up citing the ARC Prize blog, both linked above. These are vendor-run numbers on a vendor page and OpenAI's footnotes qualify several Claude scores; read the footnotes section before quoting any row. Benchmarks and prices move fast and this page will need re-checking.