Prompt Caching: Stop Paying Full Price for the Same Prompt
Prompt caching fixes that. The provider remembers the part of your prompt that never changes, and the next time you send it, you pay a small fraction of the normal price for that part.
Think of a teacher grading a stack of essays. The old way, the teacher rereads the entire grading rubric before every single essay. Same rubric, same twelve pages, every time. The new way, the teacher reads the rubric once, keeps it in mind, and only reads the new essay.
That is prompt caching. The rubric is your system prompt and instructions. The essay is the new message. The provider keeps the rubric "in mind" for a few minutes, and as long as you keep sending the same rubric, it skips the reread and charges you less.
How It Works on the Provider's End
When a model reads a prompt, it does real work on every token. It builds up an internal state as it goes, often called the KV cache (short for key-value cache), which is how the model keeps track of everything it has read so far. That work is the expensive part of processing input.
If two requests start with the exact same opening, the internal state for that opening comes out the same both times. Instead of recomputing it, the provider can hold on to it in fast memory and pick up where it left off. That is the general mechanism providers describe publicly. I am not an insider on how any one company runs its hardware, but the idea is simple enough to follow.
Why Would They Discount It?
This is the question that made the whole thing click for me. Why would a company charge you less?
Because it costs them less. Reusing work they already did takes far less compute than doing it again. Passing that savings on to you is also a nudge. It rewards developers for structuring prompts in a way that is cheaper for the provider to serve. Stable material first is cheaper for everyone.
On Anthropic's Claude API, the numbers look like this for Claude Haiku 4.5:
| Price per million tokens | |
|---|---|
| Normal input | $1.00 |
| Writing to the cache (5-minute cache) | about $1.25 |
| Reading from the cache | about $0.10 |
| Output | $5.00 |
Writing to the cache costs a little more than normal input, about 1.25 times. Every read after that costs about a tenth. If you reuse a cached prompt even twice, you come out ahead.
My Own Wake-Up Call
I have a voice assistant that runs all day. Every sentence it hears gets sent to Claude Haiku along with a long set of routing instructions, about 38,000 characters, which works out to roughly 10,500 tokens per call. The instructions tell the model what the sentence is and where it should go.
Over about a week, from September 17 to 25, that added up to about 1,183 calls and about 12.4 million input tokens. The total came to roughly $13.31, about a penny per sentence. That includes plenty of room chatter the assistant heard and then ignored. Not one of those calls used caching.
Our Toolbox. Every platform, plugin, and service we use and recommend, in one place. See the list →
I found out the hard way. The API credit ran out in the middle of the day and the assistant went quiet.
With caching on the stable part of those instructions, the same week would have cost roughly $2.
The Gotcha: One Changed Byte Breaks the Cache
Caching is a prefix match. The provider compares your prompt from the very first byte, and the moment anything differs, everything after that point is a miss.
On the Claude API, the prompt is assembled in a fixed order: tools first, then the system prompt, then the messages. You mark where the cacheable part ends with a breakpoint, and everything before it has to be byte-for-byte identical from call to call.
That is exactly where my setup would have tripped. The routing instructions included the current time and the last few lines of conversation. Both change on every call. Caching the whole block as-is would miss every single time, and I would be paying the 1.25x write price on every request instead of saving anything.
The fix is ordering. Put the stable material first, place the breakpoint right after it, and put the changing material (the time, the recent conversation, the new sentence) after the breakpoint.
Need help with this? Create a ticket today →
How to Do It on Your End
- Find the part that never changes. System instructions, examples, reference docs, tool definitions. That is your cache candidate.
- Move everything that changes to the end. Timestamps, user names, recent messages, the current request. Nothing dynamic belongs above the breakpoint.
- Check the minimum size. Very short prefixes are not cached. The minimum depends on the model, somewhere between about 512 and 4,096 tokens.
- Mark the breakpoint. On the Claude API, add
cache_control: {type: "ephemeral"}to the last block of the stable section. You can use up to four breakpoints if you have layers that change at different rates. - Keep it warm. The default cache lasts about five minutes and refreshes every time it is hit. Steady traffic keeps it alive on its own.
- Verify it. Read
usage.cache_read_input_tokensin the response. If it stays at zero across repeated calls, something in your prefix is changing. A timestamp is the classic culprit.
This Is Different From a Claude Subscription
One thing worth separating out. If you use Claude through a Pro or Max subscription, for example inside Claude Code, you are not paying per token. That usage counts against your plan limits, not against API credit. Prompt caching as a billing lever matters when you call the API directly and pay per million tokens.
The Takeaway
Prompt caching is one of the easiest cost wins in AI development. It asks you to do one thing: put what stays the same at the top and what changes at the bottom. The downside is that it is fragile, and a single stray timestamp quietly turns it off, so check the usage numbers instead of assuming.
If you are running anything that calls a model all day, look at what you are resending on every call. I hope it saves you as much as it would have saved me.
See Also
See Also
Some links in this article are affiliate links. If you purchase through them, we may earn a commission at no extra cost to you. This helps support our content.
This article blends original content, AI-assisted drafting, and human oversight. How I write.
Stay Updated
Get notified when new content is published.
No spam. Unsubscribe anytime.