Skip to content
· 3 min read

Claude Platform cost: cache hits beat a cheaper model

Claude Platform cost dropped up to 73% in Anthropic's Sept 8, 2026 tests via cache and effort. Own the agent; pay usage, not my subscription.

An orderly operations desk with a handwritten Claude Platform cost ledger, stacked identical index cards, and a marker flowchart in warm window light.
Article language

Showing original language

Anthropic’s own support-agent tests dropped Claude Platform cost by 14.6% after a single prompt audit, and one retail benchmark fell 73% once prompt caching was placed correctly. That landed September 8, 2026, in a Claude Platform note by Lance Martin. It is not a new cheaper SKU. It is a reminder that most of the bill is wasted work, not the model name on the invoice.

If you already own a Telegram agent, this is operating advice.

What actually drives Claude Platform cost?

Claude Platform cost is mostly prefill, tool loops, and “think harder” settings — not a seat you pay me. Cache misses, leftover prompt rituals, and max effort on a booking FAQ are where the money goes.

Before Claude answers, it processes the prompt into working state. Anthropic calls that prefill, and it is the expensive part of input. Prompt caching stores that state. The next request that starts with the same prefix reads it back instead of recomputing. Cache reads are billed at a fraction of full input. On Claude Fable 5.1, Anthropic listed prompt-cache reads at $0.25 per million tokens versus $1.00 on Fable 5.

The catch is mechanical. The cache is pinned to one model. The prefix has to be byte-exact. It expires. A timestamp in the system prompt, a reordered tool list, or a mid-conversation effort change can miss it. If an agent blocks on a tool longer than the TTL, the rewrite can cost 1.25× normal input. I keep tools, policy, and the CRM field map first; the live conversation after.

Do leftover “be thorough” prompts raise the bill?

On frontier Claude models, leftover “verify twice” and “be maximally thorough” instructions raise Claude Platform cost by adding extra tool calls. Anthropic’s prompt audit cut cost 14.6% and raised accuracy 5.3% on a customer-support set.

They planted those anti-patterns while migrating a support benchmark from Opus 4.8 to Opus 5. Opus 5 followed the old coaching more tightly: verification duplicated order lookups, and “be maximally thorough” turned into dozens of knowledge-base searches. Strip the rituals and the extra tool calls disappear.

Should every owner-operator job run at max effort?

No. Effort is a dial, not a quality badge. Anthropic’s Fable 5 numbers on FrontierCode Diamond went from 11.5% at $5.35 per task on low effort to 30.9% at $19.00 on max — about 2.7× the score for about 3.5× the cost.

On Humanity’s Last Exam without tools, Fable 5.1 scored about 53% at low effort for about $0.30 per question and about 61% at max for about $2.23. The last step added about half a point for 46% more cost, inside run-to-run noise. Anthropic said Fable 5.1 at low effort matched Fable 5 at high effort on CursorBench 3.2 at a third of the cost.

That matches AI agent cost control: a money budget and an action budget. Cache and effort keep the model bill inside the budget. Hard stops still sit outside the model.

What this does not change for a small-business buyer

Claude Platform cost tuning does not answer your phone or replace a $2,000–$4,000 owned Telegram agent. It only changes how expensive the model layer is once the workflow already exists.

I still deploy the same shape I use for AI for small business: one trigger, one action, one system of record, a human for exceptions. LegalBench dropped about 58% with caching, low effort, and batch. tau2-bench retail dropped about 73% with explicit cache breakpoints. OfficeQA Pro went from $136.20 to $64.87. Those are their evals, not your Jobber queue.

A Telegram AI Agent is $2,000–$4,000 once. You own the setup. You still pay Anthropic for tokens. Those tokens should not be wasted on cache misses and leftover “think step by step” instructions. If you want that wiring mapped to your CRM, send the short audit form — I reply with the replacement map within 24 hours.

Related operator notes

Keep reading

No-pressure first step

Not sure which one fits?
Get a free 20-min audit.

Bring one workflow you'd want automated. I'll tell you which deployment fits — and which doesn't — in twenty minutes. No pitch deck, no follow-up sequence. Useful even if you don't buy.

  • A real plan, not a sales call

    Which surface (Telegram, Discord, Slack, phone) fits your team, and which one doesn't.

  • Honest "don't buy this" if it applies

    If a $99/month SaaS solves it, I'll tell you which one and how.

  • A timeline + price range

    When I could deploy, what it'd cost, and what you'd own at the end.