KO

Articles · AI Design

Claude Fable 5.1 vs GPT-6 Astra, by the numbers

Two frontier models released two days apart. Official specs, third-party evaluations and early reactions from designers, side by side with sources.

· 5 min read

Translated from the Korean original. Read in Korean

In the first week of September, two frontier AI models came out two days apart: Anthropic's Claude Fable 5.1 on September 1–2, and OpenAI's GPT-6 Astra on September 3. Both were announced as the smartest yet, and, as it happens, their price sheets are identical.

Asked which is better, it's hard to answer in one line. Lay out the official specs, the independent evaluations and the reactions of people who tried them first, one after another, and you'll see why.

First, a disclaimer. Many of the benchmarks were measured by the companies that made the models. Each table notes who measured what. Also, I'm building this site with Claude Code, so I only used numbers that have a source. All figures are as of mid-September 2026 and are changing fast.

Specs: same price, different cache

Claude Fable 5.1GPT-6 Astra
ReleasedSep 1–2, 2026Sep 3, 2026
Input price (per 1M tokens)$10$10
Output price (per 1M tokens)$50$50
Cache reads (per 1M tokens)$0.25$1
Context window1M tokens1.05M tokens
Max output128K tokens128K tokens
Knowledge cutoffJune 2026April 2026

Sources: Anthropic model docs, OpenAI model docs. Bold marks the advantage.

On the table alone, they're practically twins. The differences are in two places.

For work where the model rereads the same long material, the picture changes. The cache price, what you pay to reload content it has already read, is a quarter as much for Fable 5.1 as for Astra. Attaching a brand guideline PDF and asking dozens of questions about it is exactly this kind of work. And Fable 5.1 knows about the world two months later.

There are no official numbers to compare speed. Anthropic classifies Fable 5.1 as "slower" among its own models.

Same overall score, more than double the cost

On the overall index from independent evaluator Artificial Analysis (September 9, v4.3), the two models are tied for first at 53 points. Their coding agent index is also the same, at 62.

The difference is in what it cost to reach the same score. On average, Fable 5.1 spent $7.63 per index task and Astra spent $3.26. The token prices are the same; the cost differs because of how much they say. Per task, Fable 5.1 used about 78,000 output tokens and Astra about 27,000.

Cost and output to reach the same score. A shorter bar means less was used.
Cost and output to reach the same score. A shorter bar means less was used.

One caution. This index was revised twice in the week after Astra came out, and in that time Astra climbed from fifth to first. It's too early to call these settled numbers.

Change the kind of work and the ranking changes

EvaluationWhat it measuresFable 5.1AstraMeasured by
Code Arena: WebDevBuilds a web app; people pick the better of two1758 (2nd)1800 (1st)Arena user votes
Agent ArenaLong tasks using tools1st2ndArena user votes
Design Arena 3DBuilding 3D scenes1423 (3rd)1481 (1st)Design Arena user votes
Terminal-Bench 4.0Coding tasks in the terminal55.8%57.7%OpenAI announcement
DeepSWE v1.1Long, real-world development tasks67.4%74.1%OpenAI announcement
Humanity's Last Exam (with tools)Expert-level questions65.0%57.2%OpenAI announcement

Sources: Runtime Wire (Arena, Sep 9–13), Better Stack (Design Arena), Vellum (OpenAI's announcement table).

In head-to-heads building web screens, Astra leads; in agent head-to-heads involving long tool use, Fable 5.1 leads. Even in vote-based evaluations where people choose directly, the results split like this.

The row worth noticing is the last one. Even in OpenAI's own announcement table, Fable 5.1 scores higher on expert-level questions.

It's also worth knowing that the two companies choose different tests to announce. Anthropic published a SWE-bench Pro score of 81.2% for Fable 5.1, but OpenAI didn't publish Astra's score on the same test. Cases where the same test's scores sit side by side are rarer than you'd think.

How designers reacted

When OpenAI introduced Astra, it put visual judgment in front-end design up front: give it a sketch or reference and it turns it into a working screen, refining layout, typography and spacing.

The verdicts from designers who tried them first don't lean one way. One comparison (Eidos Design, September 7) summed it up like this.

  • Astra: The first draft is already close to finished. Handles layers and space well, strong at 3D and CAD. But it has a habit of adding when it should subtract, overfilling the screen, and when it's wrong, it's confidently wrong.
  • Fable 5.1: The restrained one, better at refining a screen you'll actually ship. But it lacks boldness, so it needs another pass.
  • Both: No real taste.

A comparison that split five design tasks between them (Better Stack, September 14) reached a similar conclusion. For web design, Kimi K3 actually did best; Astra won 3D, and Fable 5.1 won 2D games. UI components were about even across the three, and none of them produced mobile app designs usable as is.

One more. Artificial Analysis noted that on a measure of the quality of documents like presentations, Astra scored lower than its predecessor, GPT-5.6 Sol. Worth knowing if you make a lot of slides.

In short

The price sheets are the same, but the actual cost depends on how you use them. For work where you attach long material and ask many times, Fable 5.1 with its cheaper cache has the edge; for short, one-shot answers, Astra, which talks less, has the edge.

In evaluations, Astra leads at quickly producing web screens and 3D, while Fable 5.1 leads at long tasks involving tool use. The approach recommended by designers who tried them first lines up with these results: explore first with Astra, finish and refine with Fable 5.1.

And both models are less than a month old. The rankings could change again in a few weeks. In the end, you only know if it fits your work by running it on your work.

AI tools are collected by type under AI tools. Where to fit any tool into your workflow, I wrote about separately in Why thirty AI drafts don't make the work any shorter.

More articles

7
2026. 09. 28 A free image model that even does transparent backgrounds. Can you use it for work?

While companies shift to cheap open models, Alibaba's new image model Qwen-Image-2.1 shut the door on commercial use. The difference between being able to download a model and being allowed to use it for work.

AI Design · 4 min
2026. 09. 27 AI got half price in a week. Will design tools follow?

On September 22, Opus 5.5, GPT-6 Sol and Luna all launched on the same day and the price sheets changed dramatically. What that means, in numbers, for the credits and subscriptions designers pay for.

AI Design · 4 min
2026. 09. 25 Why thirty AI drafts don't make the work any shorter

Exploration and mass production go to AI; direction and final checks stay with people. A boundary I only saw after months of use.

AI Design · 4 min
2026. 09. 25 Putting "make it feel right" into words

Whether it's a client or a prompt box, you end up explaining style in words. A way of breaking style into six layers.

Design · 4 min
2026. 09. 25 The references we save and never open again

Twenty-plus boards, and still I start every project by searching from scratch. How to collect references you can actually pull back out.

Design · 3 min
2026. 09. 25 Where Korean designers talk to each other

When you want to read, when you want to show your work, when you need an answer fast. Korea's design communities, sorted by what they're good for.

Community · 3 min
2026. 09. 25 Can you use AI-generated images in client work?

Copyright, terms of service, lookalikes, contracts. What to check in practice, in order. Written from Korea, with the US for comparison.

AI Design · 4 min