The two most capable models of the moment shipped in the same week. They charge the same per token, handle a million tokens of context, and both claim to be the best at writing code. The short answer: on code quality they are tied. What separates them is the cost per task and how they behave inside an agent loop.
The basics of each model, taken from the official documentation
| GPT-6 Astra | Claude Fable 5.1 | |
|---|---|---|
| Released | Sept 3, 2026 | Sept 1, 2026 |
| API ID | gpt-6-astra | claude-fable-5-1 |
| Context window | ~1.05M tokens | 1M tokens |
| Max output | 128K tokens | 128K tokens |
| Input / output (per million) | $10 / $50 | $10 / $50 |
| Cache read (per million) | $1 | $0.254× cheaper |
| Own coding agent | Codex | Claude Code |
| Also available on | ChatGPT, Amazon Bedrock | Bedrock, Google Cloud, Foundry, GitHub Copilot |
Independent measurements (Artificial Analysis) kept separate from vendor-published numbers.
If both charge $10 / $50, why does one end up cheaper?
Output is the expensive half, and Astra writes around 27K tokens per task versus 78K for Fable 5.1. On the Intelligence Index, Astra costs $3.26 per task and Fable 5.1 $7.63.
Fable 5.1 charges $0.25 per million for cache reads, four times less than before. In agent sessions that run for hours, that helps a lot.
200K tokens of fresh input, 1.8M read from cache, and each model's typical output
| Item | GPT-6 Astra | Fable 5.1 |
|---|---|---|
| Fresh input (200K) | $2.00 | $2.00 |
| Cache reads (1.8M) | $1.80 | $0.45 |
| Output | $1.35 (27K) | $3.90 (78K) |
| Total | $5.15 | $6.35 |
Simplified estimate: cache writes are not included. It shows that Fable's cheap cache does not fully offset writing nearly three times as much.
There is no absolute winner: it depends on how you write code
The same request with each vendor's official Node.js SDK
On Fable 5.1, a tool_choice of type "any" or "tool" returns a 400 error. Use "auto" with strict: true, or structured outputs.
Editing earlier messages invalidates thinking blocks. Claude Code and the Agent SDK handle it for you; if you build the messages array yourself, check it.
For small changes it may rewrite the entire file. Ask for targeted edits and you will save output tokens.
Raise the effort level for one hard step and lower it for routine ones, without losing the prompt cache.
For everyday work, something cheaper is often enough
Anthropic recommends starting here and moving up to Fable only if you need it.
Fast for completions and small changes.
30B parameters, runs offline on a 24 GB GPU.
The most useful thing you can do is run both against your own repository for a week. Benchmarks show the trend, but your codebase casts the deciding vote.
Picking the model is only the first step. Learn how to give it context with Claude Code skills and how to work with AI agents in your own projects.