Fable 5.1: 52.6% on Terminal-Bench-Science, Up From 24.7%
Anthropic released Claude Fable 5.1 and Mythos 5.1, claiming a 52.6% score on the new Terminal-Bench-Science 0.1 benchmark, up from 24.7% for Fable 5 and 29.0% for Opus 5.
What it is
Claude Fable 5.1 and Claude Mythos 5.1 are updated versions of Anthropic's top-tier models, announced by Anthropic and tested by Simon Willison with his usual pelican-drawing benchmark.
What it does
Anthropic markets Fable 5.1 for coding, knowledge work, and long-running problem-solving, highlighting a 52.6% score on Terminal-Bench-Science 0.1, a benchmark first announced August 27th, compared to 24.7% for Fable 5, 29.0% for Opus 5, and 22.4% for a competitor. Claude Code 2.1.257 already ships it as the default Fable model.
Why it matters
A 28-point jump on a benchmark that is less than a week old is the kind of number worth reading skeptically, especially since Drew Breunig already argued that Fable's pricing broke the usual cheaper-model-catches-up assumption; a big capability jump without a price cut changes that calculus further.