RouterHub

Claude Fable 5.1: Long-Horizon Intelligence for Coding and Knowledge Work

RouterHub Team · Updated 2026-09-02

Claude Fable 5.1 official launch visual

Claude Fable 5.1 is now available on RouterHub under the model ID anthropic/claude-fable-5.1.

Anthropic built Fable 5.1 for work that remains difficult over time: codebase-scale engineering, agents that operate for hours, computer-use workflows, complex research, and professional deliverables assembled from documents, tables, diagrams, and tools.

The model accepts text, images, and files and returns text, giving teams one model to evaluate across engineering, analysis, and visually grounded workflows.

What makes Claude Fable 5.1 different?

Short prompts are not the main challenge Fable 5.1 is trying to solve. Its focus is sustained execution—keeping the objective intact while planning, using tools, checking results, recovering from failures, and continuing through multiple stages.

Agentic coding across an entire codebase

Claude Fable 5.1 is designed for engineering work that extends beyond an isolated function or code snippet.

Anthropic highlights features spanning an entire codebase, difficult debugging, code review, performance work, and multi-day autonomous sessions. The model can write tests to check its own changes and use visual feedback to compare an implementation with its intended design.

Useful evaluation tasks include:

  • Investigating and fixing a bug whose cause spans multiple files
  • Implementing a feature with tests and a reviewable handoff
  • Refactoring a large section of a codebase without losing behavioral constraints
  • Reproducing a front-end design and visually checking the result
  • Running a longer engineering task with limited supervision

Agents that keep working

Fable 5.1 is also built for workflows that take hours and cross several applications.

The model can plan a sequence of work, choose and use tools, respond to intermediate failures, and provide progress updates while it runs. That makes it relevant to browser-based agents, research pipelines, operational backlogs, and other workflows where reliability across many steps matters more than the quality of a single answer.

Knowledge work with visual context

For document-heavy work, Fable 5.1 can interpret charts, tables, diagrams, images, and files as part of the same analytical task.

Potential applications include financial analysis, research synthesis, legal and policy review, architecture documentation, and executive-ready reports. The model’s visual capabilities can also support verification—for example, comparing a generated interface with a design or checking whether a chart supports the written conclusion.

What do Anthropic’s benchmarks show?

Claude Fable 5.1 official benchmark table

Anthropic’s published results show gains over Fable 5 across all seven evaluation areas included in its launch table.

Evaluation area Benchmark Fable 5.1
Agentic scientific research Terminal-Bench-Science 0.1 52.6%
Agentic coding Terminal-Bench 4.0 55.8%
Knowledge work GDPval-AA v2 1853
Computer use OSWorld 2.0 77.9% partial / 41.7% strict
Multidisciplinary reasoning Humanity’s Last Exam 60.9% without tools / 65.0% with tools
Business workflows AutomationBench 31.4%
Agentic coding CursorBench 3.2.0 73.4%

Several improvements are especially notable:

  • Terminal-Bench-Science rises from 24.7% on Fable 5 to 52.6% on Fable 5.1.
  • Terminal-Bench 4.0 improves from 42.0% to 55.8%.
  • GDPval-AA v2 increases from 1723 to 1853.
  • AutomationBench improves from 17.1% to 31.4%.
  • CursorBench 3.2.0 rises from 70.5% to 73.4%.
  • OSWorld 2.0 improves from 72.9% to 77.9% under partial scoring and from 36.1% to 41.7% under strict scoring.

The additional 60.9% result shown beneath the Terminal-Bench 4.0 score in Anthropic’s graphic belongs to Claude Mythos 5.1, not Claude Fable 5.1. The applicable Fable 5.1 result is 55.8%.

These are provider-published results rather than independent RouterHub measurements. Benchmark outcomes depend on the harness, available tools, scoring rules, model configuration, and safeguards used during evaluation. Anthropic also notes that the Terminal-Bench-Science figures have a standard error of approximately 3.5–4.5 percentage points and that its August 2026 OSWorld task release is not directly comparable with results from earlier releases.

What should teams test first?

The best starting point is not a generic chat prompt. Choose a task that exposes whether the model can preserve context, make progress, and verify its own work.

1. A repository-level engineering task

Give the model an issue that requires investigation, changes across several files, tests, and a clear handoff. Measure whether it identifies the root cause and whether its changes survive review.

2. A long-running tool workflow

Test a sequence that includes several applications or tools, intermediate decisions, and at least one recoverable failure. Track completion rate, recovery behavior, and how clearly the model reports progress.

3. A document-heavy analytical deliverable

Combine files, tables, diagrams, and numerical reasoning in one task. Evaluate the accuracy of the analysis, the traceability of its conclusions, and the amount of human correction required.

4. A visually verifiable coding task

Ask the model to implement an interface from a reference and inspect the rendered result. This tests whether visual understanding contributes to a better final implementation rather than merely a plausible first draft.

5. A multi-stage knowledge-work project

Use a task that moves from research and synthesis to a finished report or recommendation. Check whether the model remains coherent across stages and produces a concise, reviewable deliverable.

Run each task more than once and record completed-task rate, latency, consistency, failure recovery, and reviewer effort. Benchmarks can identify promising capabilities; representative workloads show whether those capabilities transfer to production.

Claude Fable 5.1 is not simply positioned as a stronger short-form assistant. Its real test is whether it can stay useful when the work becomes longer, messier, and more consequential.

Put Claude Fable 5.1 to the Test

Evaluate demanding coding, agentic, and knowledge-work tasks through RouterHub.