Claude Fable 5: Capability, Price, Government, and the Trillion-Dollar Question
Four questions are being answered with each other's evidence. We separate the machine, the price, the government entanglement, and the money — marking what is verified, what the vendor merely states, and what is genuinely contested. A map, not a verdict.
There is a model that, depending on who is describing it, is either the most capable artificial intelligence ever released to the public, a cleverly rationed luxury good, or the load-bearing asset under a trillion-dollar bet that Treasury analysts have quietly compared to the dot-com bubble. All three descriptions are being offered in the same week, by serious people, about the same object. It is called Claude Fable 5.
The difficulty in saying anything true about it is not a shortage of information. It is that three different stories are being told at once, and each is being used as evidence for the others. The person dazzled by the demonstrations says: the capability is real, so the price and the scarcity must be justified. The person alarmed by the valuation says: look at the hype machine, so the capability must be inflated. Both are looking at something real. Neither inference follows.
So we separate the layers — the machine, the price, the government, and the money — and hand you a map rather than a verdict. Where the evidence is thin, we say so. Where the fact rests only on the company’s own word, we label it. And a disclosure first, because our method requires it: this analysis was drafted with the help of an Anthropic model, about Anthropic. To earn any trust at all, every favorable conclusion was written up as a claim to be attacked and handed to a different model family, outside Anthropic, to break — a pass that did real damage and forced several corrections you will see named below.
The machine, and the thing people keep getting wrong
There are two models, not one. The company describes the public model, Fable 5, as “a Mythos-class model that we’ve made safe for general use.” Mythos 5 is “the same underlying model as Fable 5, but with the safeguards lifted in some areas” — in some areas, the fewer-safeguard sibling, not an unrestricted one, and not self-serve or generally for sale. It reaches vetted partners through a trusted-access program run with the US government, for cyber-defenders and critical-infrastructure operators. Much of the “god-model” mythology online is describing the Glasswing-only sibling’s cyber capability and silently attaching it to the model an ordinary developer can buy. The company’s own “same underlying model” framing invites that blur; the operational reality for a buyer is the safeguarded version, whose classifiers can refuse a flagged request and route it to the previous-generation model.
The most substantiated user observation — that the model has a personality that resists instruction, drifts from tight specs, and performs better with less direction — turns out to be documented, not folklore. Anthropic’s own guidance says prompts written for prior models are “often too prescriptive” and will “reduce output quality,” and describes a model that narrates more, acts more autonomously, runs far longer, and should be handed a goal rather than a procedure. The street wisdom (“give it minimal context and let it run”) and the complaint (“it has a personality I can’t suppress”) are possibly two expressions of the same design shift. What that documentation does not do is prove every “personality” claim — some of what users feel is documented design, and some is humans reading intent into a system.
The real question about the model is not whether it is capable. It is marginal: what does it add on top of what you already have? For someone starting from a vague goal, its autonomy is high-value. For someone who has already built architecture, taste, and a stack of other models cross-checking every output, its autonomy competes with judgment they have already assembled — and whether it earns its price is genuinely in doubt.
(One correction to a circulating claim: the “~30% more tokens” figure is real but describes the jump from older models to the current tokenizer, which already happened. Fable uses the same tokenizer as its cheaper sibling; the cost comparison between them is governed by posted price — which is exactly double.)
The benchmarks, and the ruler that broke
On the most-cited software-engineering benchmark, an independent site scores Fable 5 at 95%, its cheaper sibling just under 89, the largest competitor just under 83. Top of the board — but the competitor in third place has publicly stopped reporting on that benchmark, calling it contaminated and its grading flawed. A top score on a discredited instrument is a data point, not a proof.
The capability case does not need it. Other independent benchmark families, not publicly disowned in the same way, point the same direction — on hard terminal work, on full web-app generation (above 90%, though even top models still fail some apps badly enough to be unusable), on code migration. But the texture matters: on the official terminal-work leaderboard, the leading competitor in its own harness essentially ties Fable in Anthropic’s, inside the margin of error. That does not prove model and harness contribute equally — but it proves the harness materially changes the ranking. “Which model is best” has quietly become “which model, plus which scaffolding, for this task.”
So the honest answer to “is it a step change, and for what” is a map. Frontier-class for long-horizon autonomous coding — though not uniquely so. Strongly supported for building a prototype from a vague description, though not turnkey-reliable. Supported for large migrations, with cost asterisks. Likely but not independently proven for screenshot-to-code. Constrained on cybersecurity in the public model. Mixed at best for short edits and strict-voice work, where the half-price sibling is the rational default. And a real but incremental lead on the broad composite index — where, read closely, Fable’s leading score is the model with its previous-generation fallback in the loop, not a pure measurement.
The price, and the mechanism
Per token, Fable costs exactly twice its cheaper sibling — $10/$50 against $5/$25 per million, input/output. You pay double at rest, before any agent loop compounds it. And the loops compound: the model’s design thesis is the long autonomous run, which generates enormous unseen text — planning, tool calls, retries, verbose self-checking. The migration benchmark is concrete: Fable led on pass rate and cost roughly nineteen times the leading competitor per task (about $115 versus $6). For a one-off migration, nothing; for a workload run ten thousand times, the difference between viable and bleeding.
The pricing mechanism people are asking about is really about access. After the June shutdown, the model returned July 1, included in subscription plans up to half the weekly allowance and only through July 7; select Enterprise plans had included access through the same date, otherwise credits. As of now, it converts to usage credits — a flagship that arrived bundled turning into a metered good you top up and draw down. Whether that is capacity management or manufactured scarcity is a question we hold for the money.
The government, and the word “capture”
For eighteen days in June, you could not buy this model. Not an ordinary outage: Anthropic says the shutdown followed a safeguard-bypass report and, on June 12, a US export-control directive requiring restrictions on foreign nationals. Unable to verify nationality in real time, Anthropic says it suspended access for all users to ensure compliance. Controls lifted June 30; the model returned July 1.
That fact is the hardest thing in the story for any tidy narrative. A private company’s newest, highest-priced flagship was switched off, globally, for more than two weeks, by an act of the state — its commercial availability constrained by that act.
Around it sits the government-linked program through which the fewer-safeguard sibling is distributed. Anthropic reports partners finding more than 10,000 high-or-critical vulnerabilities, roughly 150 new organizations (an expansion on an earlier ~50, so closer to 200 total), 15+ countries, and up to $100M in credits. Those are company figures about a largely non-public program. An independent security-tracking firm found the public record far thinner — on the order of 75 vulnerability records mentioning the company, ~40 credited, exactly one explicitly tied to the program. Responsible disclosure could explain the gap; it does not verify the private count.
Is it capture? The case that something capture-shaped is happening is strong: the company advocates for frontier-AI regulation, and compliance-heavy regulation structurally favors incumbents. But the simplest charge — that it has captured the regulator — runs into those eighteen days. A captured regulator does not switch your product off. That is evidence against total, simple capture, not against the subtler thing: a regulated strategic-asset arrangement in which state and a few labs become mutually entangled, raising the walls around the incumbents regardless of intent. The better-supported label is “regulated strategic-asset oligopoly risk” — and the facts support treating that as a live risk, not an established structure.
The money, and the sharper analogy
At the end of May, the company raised $65B at a $965B post-money valuation, stating run-rate revenue had crossed $47B — a multiple just over twenty. All verified against the announcement. But that $47B is company-wide revenue, not this model’s; using it to prove demand for Fable specifically overreaches.
In early July, a leak: career Treasury analysts drafted a document warning that AI firms are “more deeply entrenched in the U.S. economy than their dotcom predecessors” and that a downturn “would send shockwaves throughout the entire economic ecosystem,” likening the moment to the dot-com bubble. A Treasury spokesperson called the draft “unvetted and not representative.” Both true at once: the concern exists inside the government, and it has been officially disavowed. The physical commitments are staggering — a single data-center lease tied to the company is reported near $19B for a 400-megawatt campus.
The sharpest popular framing is the Beanie Baby: manufactured scarcity propped up by fear of missing out. Separate the mechanism from the asset. As a description of the mechanism — bundled access converting to credits on a fixed date, an eighteen-day blackout, a torrent of FOMO-inducing demos — it has real purchase. As a description of the asset, it fails: a Beanie Baby had no productive use, and this model has measured, benchmarked utility. The sharper analogy this piece proposes is the fiber-optic overbuild of the late nineties: the internet was real and enormous capital was still destroyed, because more was built and borrowed against than the near-term economics could support. That is one proposed frame, not the only correct one — and it is precisely the question the Treasury analysts raised and their own agency declined to endorse.
The map
Independently verified or externally reported: the exact 2× price and the July 7 shift to metered credits; the 95% top benchmark score and its public disownment by a rival, with other independent benchmarks corroborating the underlying strength; the $65B raise at ~$965B against $47B company-wide run-rate; the ~$19B / 400 MW lease; the November congressional risk letter; the leaked Treasury draft and its disavowal.
Anthropic states (true that they state it, not independently confirmed): that under 5% of sessions hit the safety fallback; that retained data is not used to train; that Fable and Mythos are the same model with safeguards lifted in some areas; the 10,000-vulnerability / ~200-organization Glasswing scale; and the behavioral profile — more autonomous, longer turns, resistant to instruction — documented by the company and consistent with practitioner reports without proving every claim.
Not knowable from the sources: the architecture and training; whether the viral demos are typical or lucky; whether the model earns its doubled price and ~19× migration cost inside a QA-layered workflow (where the honest default is skepticism, not a flat no); whether the scarcity is capacity, safety, government pressure, or strategy; and whether a trillion-dollar valuation is vindicated by margins not yet visible or remembered as this cycle’s overbuild.
Genuinely contested: whether it is a step change (domain-dependent — see the map above); whether the government entanglement is capture or something subtler; whether the valuation is a bubble or a real technology carrying too much capital. On these, the responsible position is not to pick a side and armor it. It is to hold the map — and to notice, every time a new claim arrives, which layer it belongs to. Because the most common error in this conversation is to answer a question about the machine with a fact about the money, or to dismiss a question about the money with a fact about the machine.
The verdict is yours. Produced with politicalscience.ai methodology; full sourced brief, epistemic tiers, and the adversarial-integration log are on file.