Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026, and both models are live now on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Input and output pricing is unchanged from Claude Fable 5 at $10 and $50 per million tokens, but cache reads dropped 75 percent to $0.25 per million tokens. If you already call Claude Fable 5, three changes are breaking and will fail requests rather than degrade quietly: forced tool use now returns an error, thinking blocks are bound to the model that produced them, and editing earlier turns invalidates thinking blocks. All figures below were verified against Anthropic’s official model documentation and pricing pages in September 2026.
What Shipped on September 1, 2026
Claude Fable 5.1 uses the model ID claude-fable-5-1 and is available to all customers. Claude Mythos 5.1 uses claude-mythos-5-1 and is invitation only, offered to participants in Anthropic’s Project Glasswing. The two models share identical specifications, pricing, and capabilities. The only practical differences are who can call them and one enforcement detail covered below.

Both models carry a 1 million token context window as the default and maximum, at standard per-token pricing across the whole window, with a 128,000 token maximum output on the synchronous Messages API. Adaptive thinking is always on, steered by the effort parameter, which defaults to high. The reliable knowledge cutoff is June 2026. Anthropic commits to not retiring either model sooner than September 1, 2027.
Anthropic positions Fable 5.1 as a specialist rather than a default. The official guidance is to start with Claude Opus 5 for most workloads and reach for Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evaluations on Opus 5 at higher effort still fall short. That matters for cost, because Fable 5.1 is twice the price of Opus 5 on both input and output.
On benchmarks published with the launch, Fable 5.1 scored 52.6 percent on Terminal-Bench-Science 0.1 against 24.7 percent for Fable 5 and 29.0 percent for Opus 5. It reached 55.8 percent on Terminal-Bench 4.0, 60.9 percent on Humanity’s Last Exam without tools, and 73.4 percent on CursorBench 3.2.0. Treat vendor-published benchmarks as directional and re-run your own evaluations before switching production traffic.
The Three Breaking API Changes
These are the changes that return HTTP 400 errors. Everything else in the release is additive or a behavior shift you can tune with prompting.
1. Forced Tool Use Returns a 400 Error
Claude Fable 5.1 and Mythos 5.1 do not support forced tool use. Setting tool_choice to {"type": "any"} or {"type": "tool", "name": "..."} returns a 400 invalid_request_error with the message that type “tool” and “any” are not supported for this model. The same validation applies to the token counting endpoint, so a pre-flight token count will fail too. Only {"type": "auto"}, which is the default, and {"type": "none"} remain valid.
The reasoning is that thinking is always on for these models, and a forced tool call would skip it. The model would then write its working-out into the tool arguments instead, which lowers argument quality.
The fix is not to bolt on a retry. Leave tool_choice at auto, name the tool explicitly in your instruction text, for example telling the model to use the get_weather tool to answer, and enforce your schema with strict tool use by setting strict: true, or move the schema to structured outputs. Anthropic notes Fable 5.1 follows explicit tool instructions reliably. One caveat worth checking: in a CMEK organization, structured outputs including strict: true are not available on Claude Fable models, so you rely on the instruction alone.
2. Earlier Models Cannot Read Fable 5.1 Thinking Blocks
Every thinking block now records which model produced it, and preservation runs in one direction only. Claude Fable 5.1 reads thinking blocks from earlier models, including Opus 5, Fable 5, and older Claude models. No earlier model reads Fable 5.1’s blocks. A conversation that moves up onto Fable 5.1 keeps its reasoning. A conversation that moves down from Fable 5.1 loses it for the turns that run on the older model.
This does not error. When a request carries a block the target model cannot read, the API drops it before the model sees it, and you are not billed for those input tokens. The silent part is the risk. If you run a router, a client-side retry, or a classifier refusal fallback that switches models mid-conversation, the downstream model re-plans without that reasoning, which can raise cost and latency on the first turn after the switch, and you will see nothing in the response explaining why.
To make the drop visible, send the thinking-binding-controls-2026-08-01 beta header. Responses then carry a top-level input_transformations array naming each dropped block with reason: "model_binding_mismatch". Log it. This is the cheapest observability win in the whole release.
3. Editing Earlier Turns Invalidates Thinking Blocks
This is the change most likely to break a custom agent loop. Modifying anything before a Fable 5.1 thinking block, meaning the system prompt, the tools array, or an earlier message, causes an error on the next request. The rejection is a 400 whose message states that the block is bound to a different conversation. The error is permanent for that request body, so an automatic retry loop will never clear it.
Critically, enforcement depends on your account age. The check is enforced for new accounts created on or after August 31, 2026. For accounts created earlier, the API records the mismatch but acts on it only when the request explicitly sets thinking.block_binding.prefix_mismatch_behavior. Anthropic has said it plans to enforce the check for every account on future models. Claude Mythos 5.1 does not run this check at all.
That split creates a specific trap. If you ship a tool, SDK, or framework that users run with their own API keys, your development key is probably on an older account while your users are on new ones. Your users hit the check before you do. Test with the beta header enabled before you launch.
These patterns invalidate every later thinking block:
- Editing, reordering, or removing an earlier turn while keeping later ones.
- Injecting per-request text into an earlier turn, such as a reminder or status line, that you remove on the next request.
- Rebuilding the top-level
systemprompt ortoolsarray between requests in the same conversation. - An image or document URL that serves different bytes on a later request. The check covers the bytes, not the URL, so a rotating signed URL for the same file is fine.
- Removing a thinking block from anywhere other than the start of the run.
These patterns keep later blocks valid:
- Removing a leading run of thinking blocks, oldest first.
- Letting server-side compaction or context editing trim the history.
- Moving
cache_controlmarkers. - Changing
effort,max_tokens, or any request parameter outsidesystem,tools, andmessages.
Anthropic’s managed surfaces already handle this for you. Claude Code, claude.ai, Claude Managed Agents, and the Claude Agent SDK keep the prefix intact. If your code builds the messages array itself, you own the problem.
Pricing, Specs, and How Fable 5.1 Compares
Pricing for Fable 5.1 and Mythos 5.1 is identical to Fable 5 on every line except cache reads. Cache hits and refreshes are priced at 0.025 times the base input price on these two models, compared with the standard 0.1 multiplier every other Claude model uses. For long agentic sessions that repeatedly re-read a cached prefix, that is the single largest cost lever in the release. Cache writes and the 512 token minimum cacheable prompt length are unchanged.
| Item | Claude Fable 5.1 / Mythos 5.1 | Claude Fable 5 | Claude Opus 5 |
|---|---|---|---|
| Base input | $10 / MTok | $10 / MTok | $5 / MTok |
| Output | $50 / MTok | $50 / MTok | $25 / MTok |
| Cache read | $0.25 / MTok | $1.00 / MTok | $0.50 / MTok |
| 5m cache write | $12.50 / MTok | $12.50 / MTok | $6.25 / MTok |
| 1h cache write | $20 / MTok | $20 / MTok | $10 / MTok |
| Batch input / output | $5 / $25 per MTok | $5 / $25 per MTok | $2.50 / $12.50 per MTok |
| Context window | 1M tokens | 1M tokens | 1M tokens |
| Max output | 128K tokens | 128K tokens | 128K tokens |
| Default effort | high |
high |
high |
| Knowledge cutoff | Jun 2026 | Not applicable here | May 2026 |
| Availability | All customers / invite only | All customers | All customers |
Two operational constraints travel with these models. Both carry a mandatory 30-day data retention minimum and are not available under zero data retention unless expressly authorized by Anthropic, the same as Fable 5 and Mythos 5. Both are also classified as Covered Models. If your compliance posture requires zero data retention, Fable 5.1 is not an option without a specific agreement.
Separately, and useful for anyone budgeting a model mix, Anthropic confirmed that Claude Sonnet 5’s $2 and $10 per million token pricing, originally announced as introductory pricing through August 31, 2026, is now the standard price. The increase to $3 and $15 that had been scheduled for September 1, 2026 will not occur.
What Else Changed and What to Adopt
Five changes are purely additive. Three sit behind beta headers, so nothing breaks if you ignore them.
Per-message effort lets you change the effort level mid-conversation without invalidating the prompt cache, by sending a role: "system" message carrying only output_config. Raise it for a hard step, lower it for routine ones. It requires the mid-conversation-output-config-2026-07-01 beta header and works on Fable 5.1, Mythos 5.1, and Opus 5 on the Claude API.
Turn-scoped system messages solve the exact pattern that breaks under change number three. Set clear_at: "next_user_message" on a role: "system" message and its text carries system-prompt authority for the current turn only, then stops rendering. The message stays in messages and you keep sending it back verbatim, so nothing earlier changes, the prompt cache keeps matching, later thinking blocks stay valid, and a cleared message costs no input tokens. It requires the mid-conversation-system-clear-at-2026-08-21 beta header. If you currently inject and delete per-turn reminders, this is your replacement.
Progress updates between tool calls are exposed by setting thinking.display to "updates" with the thinking-display-updates-2026-08-18 beta header. Under the default "omitted", progress updates come back empty like reasoning, so a long agentic turn can look silent to your users. This matters more than it sounds, because Fable 5.1 writes fewer progress updates than Fable 5, especially at higher effort.
Content provenance is automatic and requires no code change. Text generated by both models carries Anthropic’s statistical text watermark on every platform. Supported image, video, and audio files produced through the code execution tool carry signed C2PA Content Credentials when retrieved through the Files API. Anthropic states the watermark adds no tokens or hidden characters and carries no information about you or your organization.
Beyond the API surface, several default behaviors changed without any code change on your side. Parallel tool calling is more variable, so Fable 5.1 may issue one tool call per turn where Fable 5 batched several, which costs extra round trips and wall-clock time without reducing answer quality. At low effort the model answers from memory more often instead of calling a search or retrieval tool. When editing text files it is more likely to rewrite the whole file than make a targeted edit. Its prose is denser in places, it uses bold, headers, and lists less than earlier Claude models, and when summarizing documents it is more likely to reproduce source passages without marking them as quotations. Each has a prompting fix in Anthropic’s model-specific prompting guide.
A practical migration order: swap the model ID, remove every tool_choice of type any or tool and move schema enforcement to strict tool use or structured outputs, then run a normal multi-turn session with the thinking-binding-controls-2026-08-01 header and prefix_mismatch_behavior: "drop_block" while logging input_transformations on every response. An empty array on every turn means your history is intact. Because setting that field opts the request into enforcement, this test works from any account regardless of when it was created. In CI, set "error" instead so a history edit fails the run. Finally, re-tune effort away from the default high, add a batching instruction to agent loops, and re-run your evaluations. Refusal handling, fallback, fallback credit, and token counts carry over unchanged, and the permitted fallback targets for Fable 5.1 are Claude Opus 4.8 and Claude Opus 5.
One more thing that did not change and still catches people: prefilling the assistant response returns a 400 error, and non-default temperature, top_p, or top_k values also return a 400 error. Extended thinking with budget_tokens and thinking: {"type": "disabled"} both error as well. Omit thinking or send {"type": "adaptive"}.
Claude Fable 5.1 and Mythos 5.1 FAQ
Do I need to change my code to use Claude Fable 5.1?
Only if you use forced tool use or edit conversation history. If you call Fable 5 with tool_choice at auto or none and you pass thinking blocks back unchanged in an append-only history, swapping the model ID from claude-fable-5 to claude-fable-5-1 is the entire migration. If you use Claude Managed Agents, no changes beyond the model name are required. The risk concentrates in custom code that builds the messages array itself, particularly agent harnesses that inject per-turn reminders or rebuild the tools array between requests.
How can I get access to Claude Mythos 5.1?
You cannot self-serve it. Claude Mythos 5.1 is offered by invitation only to approved customers in Anthropic’s Project Glasswing. It shares Fable 5.1’s specifications, pricing, model behavior, and the same September 1, 2027 earliest retirement date, so there is no capability advantage to chasing it if you already have Fable 5.1 access. To request access, contact your Anthropic, AWS, or Google Cloud account team.
Is Claude Fable 5.1 actually cheaper than Claude Fable 5?
Base input and output prices are identical at $10 and $50 per million tokens, so a single stateless call costs exactly the same. The savings come entirely from cache reads dropping from $1.00 to $0.25 per million tokens. That means the discount scales with how much of your traffic is cached prefix re-reads, which is why the benefit is largest on long-running agentic sessions and smallest on one-shot requests. Model your own cache hit ratio rather than assuming a headline percentage. Also note that Fable 5.1 remains twice the price of Claude Opus 5, which Anthropic recommends as the starting point for most workloads.
Clear, fact-checked guides on money, tech and everyday decisions.
Related reading
- Google AI Mode Now Tracks Flight Prices and Books Hotels
- Background AI Agents in 2026: What Claude Cowork and Gemini Spark Actually Do
- Google Pics Launches: Google’s AI Answer to Canva
More on this topic: all AI Tools articles · full article index