Token accounting for AI agents: tokens, context windows, and caching from first principles
A model has no memory - everything that looks like memory is your code resending it. Get that one sentence right and you understand tokens, context windows, and prompt caching, and why an agent's bill balloons.
On this page
- Five words to learn before you read on
- One API call: the model remembers nothing
- Many chat turns: you resend everything, every time
- What it means for one turn to have many steps
- Four common questions, answered straight
- Working one real turn step by step
- Two things to read out of these numbers
- The context window: the capacity of a single call
- Which box is the ring in Claude Code
- The shape of the two lines says it all
- Three proofs it is not a cumulative number
- The biggest item people forget
- Prompt caching: you pay again, but 10x cheaper
- What it does
- Three kinds of tokens, three prices
- The most important rule: it matches by PREFIX
- Where the two providers differ
- Things that break the cache and tell you nothing
- Three architectural rules, more important than marking the cache
- Data tables
- Windows and prices by model
- Claude minimum cache thresholds
- Claude cache rules
- What change breaks which cache (Claude)
- OpenAI cache rules
- How ChatGPT differs from the API
- Mapping to zalo-agent
- Reading the three numbers on the dashboard
- Glossary
A model has no memory. Everything that looks like memory is your code resending it. Get that one sentence right and the whole story of tokens, context windows, and caching falls into place.
How to read this First time through, go in order from Part 1 - each part adds exactly one new idea, with a diagram beside it. Coming back to look something up? Jump straight to Part 4 - all tables, no prose.
Part 1 - Foundations
Five words to learn before you read on
Most misunderstandings about tokens come from mixing up the five words below. Four of them nest inside each other like Russian dolls - the outer one holds many of the inner.
Only the innermost box - the request - compares to the context window. The three outer boxes are cumulative totals; they can be far larger than the window and never overflow.
| Word | In short | Example |
|---|---|---|
| Token | The smallest unit a model reads and writes. Not a character, not a word. | "Hello there" ~ 2-3 tokens |
| Request | One package of data sent to the API. Holds everything the model sees that time. | 10,600 tokens |
| Step | One API call. If the model asks to use a tool, you need another step. | This turn has 2 steps |
| Turn | One full question-and-answer with the user. One or more steps. | 21,653 tokens |
| Session | The whole conversation. Many turns. | 84,601 tokens |
The one sentence for this section The three numbers 10,600 / 21,653 / 84,601 above are all tokens, all from the same real conversation, and none of them is wrong. They measure three different things. Only the first - the request - has anything to do with "will it overflow the window".
One API call: the model remembers nothing
This is the single most important thing in the whole piece. A model has no memory between calls.
Picture each API call as handing a folder to an expert who has never met you before. They read the entire folder, write an answer, then forget everything. Next time you come back, it is a brand-new person - you have to hand over the whole folder again, plus the new part.
The technical term for this is stateless - it keeps no state. True for both Claude and OpenAI.
The consequence to remember Everything that looks like a bot's "memory" - remembering your name, remembering yesterday, remembering which tool it just called - is not in the model. It lives in your code, and it is resent on every call.
Many chat turns: you resend everything, every time
Because the model forgets everything, if you want it to know the past, your code has to resend it. There is no other way.
This is exactly the problem prompt caching was created to solve. More on that in Part 3.
// What turn 3 actually sends - NOT just "Thanks"
messages: [
{ role: "user", content: "My name is An" },
{ role: "assistant", content: "Hi An!" },
{ role: "user", content: "What's my name?" },
{ role: "assistant", content: "An." },
{ role: "user", content: "Thanks" } // only this line is new
]Part 2 - The heart of it
What it means for one turn to have many steps
This is the part people get tangled up in most, so let's go slowly.
A model cannot run tools. It can only say that it wants to call a certain tool with certain arguments. Running the tool is your code's job. And for the model to learn the result, you have to call the API one more time - because the model forgot everything from last time.
So a turn that needs a tool will have at least two API calls. Each of those calls is a step.
This loop repeats until the model stops asking for tools. In the example above it stops after 2 steps.
Four common questions, answered straight
Are multiple steps the same API call?
No. Each step is a separate, fully independent API call. A 2-step turn = 2 calls. A 5-step turn = 5 calls.
Is it the same model?
Same model (unless your code deliberately switches), but not the same working session. Think of it as asking the same expert twice, with their memory wiped between the two.
Within that turn, does the model see enough context from the earlier steps?
Yes - but not because the model remembers. It's because your code gathers it and sends it along. Step 2 receives everything step 1 sent, plus two new pieces: the model's tool request, and the tool result. If your code forgets to include them, the model behaves as if it never called a tool at all.
How do tokens grow across steps?
They grow multiplicatively, not by a small addition. Each step pays again for the entire context, and the context keeps swelling because tool results accumulate.
This is why a tool that returns long content (reading a web page, reading a document) is far more expensive than it looks: you pay for it on every remaining step of the turn.
Images are the most expensive item One image costs on the order of thousands of tokens. Reload 3 old images into every turn, and if the turn runs 8 steps, you pay for those 3 images 8 times. Measured on a real bot: a turn with images ~ 12,600 tokens, a text-only turn ~ 2,600.
Working one real turn step by step
The numbers below are copied straight from the trace of a running Zalo bot - not a made-up example. The user sends "remind me at 11 please", the bot schedules it and replies.
- Step 1. Your code assembles the folder and sends it: 10,600 tokens in. The model reads it, writes a short line plus a tool call: 196 tokens out. This step costs: 10,600 + 196 = 10,796.
- Your code runs the tool
schedule_task, creating the entry in the database. This step is 0 tokens - it never touches the model. - Step 2. Your code re-gathers everything from step 1, plus the tool request and the tool result, and sends it: 10,803 tokens in. The model writes the final answer: 54 tokens out. This step costs: 10,803 + 54 = 10,857.
- The model asks for no more tools, and the loop stops. Turn total: 21,653 tokens.
Two things to read out of these numbers
One. Step 2 sends 10,803, only 203 tokens more than step 1. Because it resends step 1's 10,600 verbatim, plus two new pieces. The context barely changed - but you paid for it a second time.
Two. The turn total of 21,653 is nearly double the actual context. Not because the folder doubled in size, but because the same folder was billed twice.
This bot's default context window is 128,000, and it only uses up to 70% to leave headroom. The 10,803-token request is taking about 12% of the budget - nowhere near overflow.
The most common number-reading mistake You see 84,601 tokens on the dashboard and worry "about to overflow the 128,000 window". No. That is money already spent on the whole conversation, summed across many turns. The number to watch for overflow is 10,803 - the context of a single request.
The tell: the cumulative number only goes up, even when you wipe the context clean. The request number rises and falls with conversation length.
The context window: the capacity of a single call
The context window is the maximum capacity of ONE request. Not of the whole conversation.
This is why a real bot typically only uses up to 70% of the declared window - the reserve is for the reply and for estimation error.
Why a bot "forgets" old things It is not that the model has a poor memory. It is that your code trimmed the history before sending, to make it fit the window. An agent's memory is a decision your code makes - to remember longer, change how you trim, not which model you use.
Which box is the ring in Claude Code
Claude Code shows a context indicator, and it grows after each message. The natural question: so is it counting the whole session?
No. It is the REQUEST box - the innermost box of the Russian-doll diagram at the top.
It grows after each message, but not because it accumulates. It's because that one request keeps swelling: each turn, your code appends the new message to the message list, so the next API call carries more text than the last. A single thing, getting bigger.
The shape of the two lines says it all
Illustrative scale, not real numbers. What matters is the shape: the cumulative line only climbs and passes the ceiling harmlessly, while the request line drops straight down when context is compacted.
Three proofs it is not a cumulative number
| Observation | If it were cumulative | Reality |
|---|---|---|
| Fresh session, nothing typed yet | should be 0 | already 10-20%: system prompt, tool definitions, CLAUDE.md, MCP schema |
Run /compact or /clear | no change - spent money isn't refunded | drops |
| After 6 turns of about 10K each | over 60K | might still be around 12K |
The deciding test Whatever can drop is measuring capacity. A cumulative number never drops.
An analogy: the context ring is how full one suitcase is - a fixed capacity; add more and it fills, take some out and it empties. Cumulative tokens are the total weight you carried all day - never decreases, and being many times the suitcase's capacity is perfectly normal. Both rise during an ordinary session. That is exactly why they are so easy to confuse.
The biggest item people forget
On every API call, Claude Code carries: the system prompt, all tool definitions, the contents of CLAUDE.md, the MCP schema, the full conversation history, and the result of every tool it has run.
That last item is the expensive one. A command that reads a 2000-line file stays in the context and gets resent on every subsequent turn - the same accumulation mechanism as the step section, except here it accumulates across the whole session rather than one turn. That is why the ring jumps after a few file reads, even when you only typed a short line.
The practical upshot
/clearbetween two unrelated tasks is a habit worth having: it throws out the pile of old tool results that are being carried back and forth every turn.
Part 3 - Saving money
Prompt caching: you pay again, but 10x cheaper
Parts 1 and 2 leave one obvious problem: the same pile of text is sent over and over - once per step, once per turn. Prompt caching exists to solve exactly that.
What it does
The provider saves the computed state of the front of your prompt. Next time you send a prompt whose front is byte-for-byte identical, they reload the saved state instead of recomputing from scratch. You get a discount and a faster reply.
The most common misconception, stated plainly Caching does not make old tokens free. They are still counted, still billed on every turn - just about 10x cheaper. And they still take up room in the context window as usual. Caching makes things cheaper, not bigger.
Three kinds of tokens, three prices
| Token kind | Price factor | Meaning |
|---|---|---|
| Uncached | 1.0x | Full price. The new part of each request. |
| Cache write | 1.25x | A little more expensive, paid once at creation. |
| Cache read | 0.1x | 10x cheaper. The reward on every later call. |
True for both Claude and OpenAI (GPT-5.6 and later). Claude also offers a 1-hour cache, at a 2.0x write price.
The most important rule: it matches by PREFIX
Cache matches on the front of the prompt, not on separate scattered chunks. Change one byte in the middle and everything from there on loses its cache - even the parts behind it that didn't change.
The shortcut rule: what never changes goes first, what changes every time goes last. True for both providers.
Where the two providers differ
| Claude | OpenAI | |
|---|---|---|
| How to enable | Mark it yourself with cache_control | Automatic, no code change |
| Choose the cut points | Yes, up to 4 | No, the system chooses |
| See the result in | cache_read_input_tokens | cached_tokens |
The full details are in the reference part. The practical difference: Claude lets you decide what is worth caching, OpenAI handles it for you - in exchange, Claude can optimize deeper, and can also get it wrong more often.
Things that break the cache and tell you nothing
Every item below fails silently - no error, no warning, just a bigger bill. This is a list to grep for in your prompt-building code.
| Code pattern | Why it breaks |
|---|---|
Date.now() in the system prompt | The front changes every request |
randomUUID() placed early | Every request becomes its own prefix |
JSON.stringify() without sorting keys | Key order is unstable, bytes differ |
| Putting a user id in the system prompt | Everyone gets their own prefix, nothing shared |
if (flag) system += "..." | Each flag combination is a different prefix |
tools = buildTools(user) | Tools sit at the very front - break here and all is lost |
Three architectural rules, more important than marking the cache
- Freeze the system prompt. Don't put date, time, mode, or user name in it. If you need to inject dynamic context, inject it at the end of the message list - a message at turn 5 breaks nothing before turn 5.
- Don't switch tools or model mid-stream. Tools sit at the very front; adding, removing, or reordering a tool wipes the whole cache. Cache is also bound to the model, so switching models loses it.
- Secondary calls must reuse the exact front of the main call. Summarization, context compaction, sub-agents - if you rebuild the system or tools even slightly differently, you miss the parent call's cache entirely.
A quick way to check Run two back-to-back requests with an identical front, then look at
cache_read_input_tokens(Claude) orcached_tokens(OpenAI). Still 0 means something is definitely breaking the prefix - diff the bytes of the two built prompts to find it.
Part 4 - Reference
Data tables
Windows and prices by model
| Model | Window | Input $/1M | Output $/1M |
|---|---|---|---|
| Claude Opus 5 | 1M | $5.00 | $25.00 |
| Claude Sonnet 5 | 1M | $3.00 | $15.00 |
| Claude Haiku 4.5 | 200K | $1.00 | $5.00 |
| GPT-5.6 Sol | ~1.05M | $5.00 | $30.00 |
| GPT-5.6 Terra | ~1.05M | $2.00 | $12.00 |
| GPT-5.6 Luna | ~1.05M | $0.20 | $1.20 |
Claude prices from official Anthropic docs. GPT-5.6 prices and windows from third-party aggregators (Aug 2026) - re-check the official pricing page before costing anything real. GPT-5.6 also has a long-context tier: past 272K input tokens the whole request is billed at about 2x input, about 1.5x output.
Claude minimum cache thresholds
| Model | Minimum prefix |
|---|---|
| Claude Opus 5, Fable 5, Mythos 5 | 512 |
| Opus 4.8, Sonnet 5, Sonnet 4.6 / 4.5, Opus 4.1 / 4, Sonnet 4 | 1,024 |
| Opus 4.7, Mythos Preview, Haiku 3.5 | 2,048 |
| Opus 4.6, Opus 4.5, Haiku 4.5 | 4,096 |
The trap most worth remembering in this whole piece This threshold does not rise evenly across model generations. A 3,000-token prompt caches on Opus 5 and Sonnet 4.5, but silently does not cache on Opus 4.6 or Haiku 4.5. Below the threshold there is no error - just
cache_creation_input_tokens: 0and a bill 10x higher.
Claude cache rules
| Item | Rule |
|---|---|
| Render order | tools -> system -> messages |
| Max cut points | 4 per request |
| TTL | 5 minutes (default) or 1 hour |
| Write price | 1.25x (5 min) / 2.0x (1 hour) |
| Read price | 0.1x |
| Break-even | 2 requests (5-min TTL) / 3 requests (1-hour TTL) |
| Look-back window | Up to 20 content blocks - tool-heavy turns can exceed this and miss silently |
| Parallel requests | Cache is only readable after the first response starts streaming |
// The formula to remember when reading Claude's usage
total prompt = input_tokens // the uncached part
+ cache_creation_input_tokens // the part just written to cache
+ cache_read_input_tokens // the part read from cache
// Looking at input_tokens alone and assuming a small context is a mistake.What change breaks which cache (Claude)
| Change | tools cache | system cache | messages cache |
|---|---|---|---|
| Tool definitions (add/remove/reorder) | lost | lost | lost |
| Switch model | lost | lost | lost |
| System prompt content | kept | lost | lost |
tool_choice, images, toggling thinking | kept | kept | lost |
| Message content | kept | kept | lost |
Which means changing tool_choice or toggling thinking per request does not lose the tools and system cache at all. Only switching tools and switching models forces a full rebuild.
OpenAI cache rules
| Item | Rule |
|---|---|
| How to enable | Automatic, no declaration |
| Minimum threshold | 1,024 tokens (GPT-5.6+). Older models: 1,024 - 2,048 |
| Prefix increment | Older models: in multiples of 128 tokens |
| Read price | 0.1x (90% off) |
| Write price | 1.25x on GPT-5.6+ |
| TTL | GPT-5.6+: only 30m. Older models up to 24 hours |
| Real expiry | Usually cleared after 5-10 minutes idle, at most 1 hour |
| Reporting field | cached_tokens in prompt_tokens_details / input_tokens_details |
Don't trust old numbers Many articles online still say OpenAI gives a 50% discount - that was the launch figure in 2024. It is 90% now. When it comes to pricing, always check the date of the source.
How ChatGPT differs from the API
A common question: "On ChatGPT I never resend anything, so how does it still remember?"
Because ChatGPT is an app built on the API, and it handles the resending for you. Their server stores the conversation, reassembles it every turn, and sends it to the model. From the model's point of view, the whole conversation is still resent every turn - exactly like your bot.
| Direct API | ChatGPT | |
|---|---|---|
| Who keeps the history | You | OpenAI |
| Who decides what to trim when it's long | You | OpenAI |
| Does the model see the whole conversation? | Yes, the part you send | Yes, the part they assemble |
| You pay by | Token | Subscription |
| Can you see the token count? | Yes, in usage | No |
When a ChatGPT conversation exceeds the window, they trim or summarize the old part - which is why it sometimes "forgets" what you said at the start of a very long conversation. The exact mechanism you have to write yourself for your own agent.
About the per-plan window ChatGPT Free / Plus / Pro windows are smaller than the model's max window on the API, and the numbers going around online contradict each other - third-party sources, changing constantly, differing by fast-reply versus deep-thinking mode. If you need an exact number, check OpenAI's official help page at that moment.
Mapping to zalo-agent
This piece was originally written to explain token costs for zalo-agent - an open-source (MIT) agent that runs on Zalo. The table below maps the concepts above onto their exact places in the repo, in case you want to see a real implementation.
| Concept in this piece | In the repo |
|---|---|
| Context window | LLM_CONTEXT_WINDOW, default 128,000 |
| Reserve room for the reply | nganSachAnToan() - uses only up to 70% |
| Trim before sending | Drops old images first, then old messages |
| Count tokens, not messages | A single Zalo message can be any length; counting messages can't cap context growth |
| Cap on reloaded images | HISTORY_IMAGE_CONTEXT_LIMIT |
| Enable caching | A stable x-session-id header per thread |
| Session cumulative tokens | The total_tokens column in the agent_turns table |
| Cap on steps per turn | LLM_MAX_STEPS |
Reading the three numbers on the dashboard
| Where you see it | Which kind | Compare to window? |
|---|---|---|
| Sessions: "6 messages - 84,601 tokens" | Whole session | No |
| Trace, turn-card label: "2 steps - 21,653 tokens" | One turn | No |
| Trace, inside a step: "10,600 in / 196 out" | One request | Yes - this is it |
Glossary
| Term | Meaning |
|---|---|
| Token | The unit a model reads and writes. Vietnamese costs more tokens than English for the same idea - don't estimate by word count. |
| Context window | The maximum capacity of one request. Includes the part the model is about to write. |
| Stateless | Keeps no state. The server forgets everything after each call. |
| Step | One API call inside a turn. When the model asks for a tool, the next step begins. |
| Agent loop | The loop: call the model, model asks for a tool, code runs the tool, sends the result, repeat until the model stops asking. |
| Prefix | The front of the prompt. Cache matches here. |
| Breakpoint | A marker for "cache up to here". Claude only. |
| TTL | How long a cache stays alive if nobody reuses it. |
| Cache write / read | Creating the cache the first time (more expensive) and reusing it (much cheaper). |
How this piece is built. Structured after Diataxis - separating the explanation (Parts 1 to 3, read in order) from the reference (Part 4, jump straight in). Presented on Mayer's principles: define the vocabulary before using it, one idea per section, put each diagram right next to the paragraph it illustrates, and use a consistent color for the three kinds of tokens throughout.
Sources. The Claude parts are from official Anthropic docs. The OpenAI parts are from the official prompt-caching guide and the feature announcement. GPT-5.6 prices and windows, and the ChatGPT-plan numbers, are from third-party aggregators in August 2026 and are marked as such in place. Every measurement in the example section is copied directly from a running bot's trace, not a made-up example.