Vũ Văn HảiFull-stack · AI-native
GuidesBlog
Discuss a project

© 2026 Vu Van Hai · Written from real deployment experience.

HomeGuidesBlogRSS
  1. Blog
  2. /Token accounting for AI agents: tokens, context windows, and caching from first principles

Token accounting for AI agents: tokens, context windows, and caching from first principles

A model has no memory - everything that looks like memory is your code resending it. Get that one sentence right and you understand tokens, context windows, and prompt caching, and why an agent's bill balloons.

Published: Sep 23, 202629 min read
LLMTokensPrompt cachingContext windowClaude Code
On this page
  • Five words to learn before you read on
  • One API call: the model remembers nothing
  • Many chat turns: you resend everything, every time
  • What it means for one turn to have many steps
  • Four common questions, answered straight
  • Working one real turn step by step
  • Two things to read out of these numbers
  • The context window: the capacity of a single call
  • Which box is the ring in Claude Code
  • The shape of the two lines says it all
  • Three proofs it is not a cumulative number
  • The biggest item people forget
  • Prompt caching: you pay again, but 10x cheaper
  • What it does
  • Three kinds of tokens, three prices
  • The most important rule: it matches by PREFIX
  • Where the two providers differ
  • Things that break the cache and tell you nothing
  • Three architectural rules, more important than marking the cache
  • Data tables
  • Windows and prices by model
  • Claude minimum cache thresholds
  • Claude cache rules
  • What change breaks which cache (Claude)
  • OpenAI cache rules
  • How ChatGPT differs from the API
  • Mapping to zalo-agent
  • Reading the three numbers on the dashboard
  • Glossary

A model has no memory. Everything that looks like memory is your code resending it. Get that one sentence right and the whole story of tokens, context windows, and caching falls into place.

How to read this First time through, go in order from Part 1 - each part adds exactly one new idea, with a diagram beside it. Coming back to look something up? Jump straight to Part 4 - all tables, no prose.


Part 1 - Foundations

Five words to learn before you read on

Most misunderstandings about tokens come from mixing up the five words below. Four of them nest inside each other like Russian dolls - the outer one holds many of the inner.

SESSION one conversation, from open to clear TURN one user message, one full bot reply STEP one API call. A turn can have many steps. REQUEST the payload sent in that step = total TOKENS must fit the window

Only the innermost box - the request - compares to the context window. The three outer boxes are cumulative totals; they can be far larger than the window and never overflow.

WordIn shortExample
TokenThe smallest unit a model reads and writes. Not a character, not a word."Hello there" ~ 2-3 tokens
RequestOne package of data sent to the API. Holds everything the model sees that time.10,600 tokens
StepOne API call. If the model asks to use a tool, you need another step.This turn has 2 steps
TurnOne full question-and-answer with the user. One or more steps.21,653 tokens
SessionThe whole conversation. Many turns.84,601 tokens

The one sentence for this section The three numbers 10,600 / 21,653 / 84,601 above are all tokens, all from the same real conversation, and none of them is wrong. They measure three different things. Only the first - the request - has anything to do with "will it overflow the window".

One API call: the model remembers nothing

This is the single most important thing in the whole piece. A model has no memory between calls.

Picture each API call as handing a folder to an expert who has never met you before. They read the entire folder, write an answer, then forget everything. Next time you come back, it is a brand-new person - you have to hand over the whole folder again, plus the new part.

YOUR MACHINE REQUEST system prompt tool definitions all messages (old and new) send PROVIDER MACHINE MODEL reads the ENTIRE dossier then writes a reply return THE REPLY + tokens used then FORGETS EVERYTHING keeps nothing for next time

The technical term for this is stateless - it keeps no state. True for both Claude and OpenAI.

The consequence to remember Everything that looks like a bot's "memory" - remembering your name, remembering yesterday, remembering which tool it just called - is not in the model. It lives in your code, and it is resent on every call.

Many chat turns: you resend everything, every time

Because the model forgets everything, if you want it to know the past, your code has to resend it. There is no other way.

EACH BOX = ONE MESSAGE SENT ALONG Turn 1 My name is An 1 msg Turn 2 My name is An Hi An! What's my name? 3 msgs Turn 3 My name is An Hi An! What's my name? An. Thanks Solid box = new message. Faded box = old message, still resent and still billed. The longer the chat, the pricier each turn - even if you only add two words.

This is exactly the problem prompt caching was created to solve. More on that in Part 3.

// What turn 3 actually sends - NOT just "Thanks"
messages: [
  { role: "user",      content: "My name is An" },
  { role: "assistant", content: "Hi An!" },
  { role: "user",      content: "What's my name?" },
  { role: "assistant", content: "An." },
  { role: "user",      content: "Thanks" }        // only this line is new
]

Part 2 - The heart of it

What it means for one turn to have many steps

This is the part people get tangled up in most, so let's go slowly.

A model cannot run tools. It can only say that it wants to call a certain tool with certain arguments. Running the tool is your code's job. And for the model to learn the result, you have to call the API one more time - because the model forgot everything from last time.

So a turn that needs a tool will have at least two API calls. Each of those calls is a step.

YOUR MACHINE PROVIDER MACHINE STEP 1 Assemble dossier: system + tools + history + the user's new message 10,600 tokens Reads all. Decides: "I need the schedule_task tool" writes 196 tokens returns YOUR CODE runs the tool The model can't touch your machine. No tokens spent in this step. STEP 2 Re-assemble ALL of step 1's dossier + the model's tool request + the tool result just produced 10,803 tokens Reads AGAIN from scratch (forgot step 1) "Scheduled for 11:00 on 30/08" writes 54 tokens The model does NOT remember step 1 - you resent it, so it knows.

This loop repeats until the model stops asking for tools. In the example above it stops after 2 steps.

Four common questions, answered straight

Are multiple steps the same API call?

No. Each step is a separate, fully independent API call. A 2-step turn = 2 calls. A 5-step turn = 5 calls.

Is it the same model?

Same model (unless your code deliberately switches), but not the same working session. Think of it as asking the same expert twice, with their memory wiped between the two.

Within that turn, does the model see enough context from the earlier steps?

Yes - but not because the model remembers. It's because your code gathers it and sends it along. Step 2 receives everything step 1 sent, plus two new pieces: the model's tool request, and the tool result. If your code forgets to include them, the model behaves as if it never called a tool at all.

How do tokens grow across steps?

They grow multiplicatively, not by a small addition. Each step pays again for the entire context, and the context keeps swelling because tool results accumulate.

CONTEXT SENT AT EACH STEP Step 1 10,600 Step 2 10,803 Step 3 + web_fetch result Step 4 The darkest block is the original context - resent INTACT at every step. Each lighter block is a tool result added on, then resent forever too.

This is why a tool that returns long content (reading a web page, reading a document) is far more expensive than it looks: you pay for it on every remaining step of the turn.

Images are the most expensive item One image costs on the order of thousands of tokens. Reload 3 old images into every turn, and if the turn runs 8 steps, you pay for those 3 images 8 times. Measured on a real bot: a turn with images ~ 12,600 tokens, a text-only turn ~ 2,600.

Working one real turn step by step

The numbers below are copied straight from the trace of a running Zalo bot - not a made-up example. The user sends "remind me at 11 please", the bot schedules it and replies.

  • Step 1. Your code assembles the folder and sends it: 10,600 tokens in. The model reads it, writes a short line plus a tool call: 196 tokens out. This step costs: 10,600 + 196 = 10,796.
  • Your code runs the tool schedule_task, creating the entry in the database. This step is 0 tokens - it never touches the model.
  • Step 2. Your code re-gathers everything from step 1, plus the tool request and the tool result, and sends it: 10,803 tokens in. The model writes the final answer: 54 tokens out. This step costs: 10,803 + 54 = 10,857.
  • The model asks for no more tools, and the loop stops. Turn total: 21,653 tokens.

Two things to read out of these numbers

One. Step 2 sends 10,803, only 203 tokens more than step 1. Because it resends step 1's 10,600 verbatim, plus two new pieces. The context barely changed - but you paid for it a second time.

Two. The turn total of 21,653 is nearly double the actual context. Not because the folder doubled in size, but because the same folder was billed twice.

THREE NUMBERS, ONE CONVERSATION One request 10,803 comparable to the context window One turn 21,653 just money, not comparable Whole session 84,601 also just money, summed since the chat opened Same scale. These three bars measure three different things.

This bot's default context window is 128,000, and it only uses up to 70% to leave headroom. The 10,803-token request is taking about 12% of the budget - nowhere near overflow.

The most common number-reading mistake You see 84,601 tokens on the dashboard and worry "about to overflow the 128,000 window". No. That is money already spent on the whole conversation, summed across many turns. The number to watch for overflow is 10,803 - the context of a single request.

The tell: the cumulative number only goes up, even when you wipe the context clean. The request number rises and falls with conversation length.

The context window: the capacity of a single call

The context window is the maximum capacity of ONE request. Not of the whole conversation.

CONTEXT WINDOW - ONE REQUEST system prompt tool defs chat history + tool results images room for model to write reply Exceed the window and the API returns an error, it won't trim for you. So every serious agent must trim context BEFORE sending. Easy to forget: what the model is about to WRITE eats the same budget.

This is why a real bot typically only uses up to 70% of the declared window - the reserve is for the reply and for estimation error.

Why a bot "forgets" old things It is not that the model has a poor memory. It is that your code trimmed the history before sending, to make it fit the window. An agent's memory is a decision your code makes - to remember longer, change how you trim, not which model you use.

Which box is the ring in Claude Code

Claude Code shows a context indicator, and it grows after each message. The natural question: so is it counting the whole session?

No. It is the REQUEST box - the innermost box of the Russian-doll diagram at the top.

It grows after each message, but not because it accumulates. It's because that one request keeps swelling: each turn, your code appends the new message to the message list, so the next API call carries more text than the last. A single thing, getting bigger.

The shape of the two lines says it all

context window ceiling 1 2 3 4 5 6 turn # over the ceiling, HARMLESSLY /compact context ring = one REQUEST cumulative tokens (session)

Illustrative scale, not real numbers. What matters is the shape: the cumulative line only climbs and passes the ceiling harmlessly, while the request line drops straight down when context is compacted.

Three proofs it is not a cumulative number

ObservationIf it were cumulativeReality
Fresh session, nothing typed yetshould be 0already 10-20%: system prompt, tool definitions, CLAUDE.md, MCP schema
Run /compact or /clearno change - spent money isn't refundeddrops
After 6 turns of about 10K eachover 60Kmight still be around 12K

The deciding test Whatever can drop is measuring capacity. A cumulative number never drops.

An analogy: the context ring is how full one suitcase is - a fixed capacity; add more and it fills, take some out and it empties. Cumulative tokens are the total weight you carried all day - never decreases, and being many times the suitcase's capacity is perfectly normal. Both rise during an ordinary session. That is exactly why they are so easy to confuse.

The biggest item people forget

On every API call, Claude Code carries: the system prompt, all tool definitions, the contents of CLAUDE.md, the MCP schema, the full conversation history, and the result of every tool it has run.

That last item is the expensive one. A command that reads a 2000-line file stays in the context and gets resent on every subsequent turn - the same accumulation mechanism as the step section, except here it accumulates across the whole session rather than one turn. That is why the ring jumps after a few file reads, even when you only typed a short line.

The practical upshot /clear between two unrelated tasks is a habit worth having: it throws out the pile of old tool results that are being carried back and forth every turn.


Part 3 - Saving money

Prompt caching: you pay again, but 10x cheaper

Parts 1 and 2 leave one obvious problem: the same pile of text is sent over and over - once per step, once per turn. Prompt caching exists to solve exactly that.

What it does

The provider saves the computed state of the front of your prompt. Next time you send a prompt whose front is byte-for-byte identical, they reload the saved state instead of recomputing from scratch. You get a discount and a faster reply.

The most common misconception, stated plainly Caching does not make old tokens free. They are still counted, still billed on every turn - just about 10x cheaper. And they still take up room in the context window as usual. Caching makes things cheaper, not bigger.

Three kinds of tokens, three prices

Token kindPrice factorMeaning
Uncached1.0xFull price. The new part of each request.
Cache write1.25xA little more expensive, paid once at creation.
Cache read0.1x10x cheaper. The reward on every later call.

True for both Claude and OpenAI (GPT-5.6 and later). Claude also offers a 1-hour cache, at a 2.0x write price.

The most important rule: it matches by PREFIX

Cache matches on the front of the prompt, not on separate scattered chunks. Change one byte in the middle and everything from there on loses its cache - even the parts behind it that didn't change.

RIGHT system + tools + old history (unchanged) new question The green part is cacheable; next time you pay only 0.1x for it. WRONG current date-time system + tools + old history + new question The date-time changes every minute and sits at the FRONT, so everything after loses cache. No error is raised. The bill is simply 10x higher than usual.

The shortcut rule: what never changes goes first, what changes every time goes last. True for both providers.

Where the two providers differ

ClaudeOpenAI
How to enableMark it yourself with cache_controlAutomatic, no code change
Choose the cut pointsYes, up to 4No, the system chooses
See the result incache_read_input_tokenscached_tokens

The full details are in the reference part. The practical difference: Claude lets you decide what is worth caching, OpenAI handles it for you - in exchange, Claude can optimize deeper, and can also get it wrong more often.

Things that break the cache and tell you nothing

Every item below fails silently - no error, no warning, just a bigger bill. This is a list to grep for in your prompt-building code.

Code patternWhy it breaks
Date.now() in the system promptThe front changes every request
randomUUID() placed earlyEvery request becomes its own prefix
JSON.stringify() without sorting keysKey order is unstable, bytes differ
Putting a user id in the system promptEveryone gets their own prefix, nothing shared
if (flag) system += "..."Each flag combination is a different prefix
tools = buildTools(user)Tools sit at the very front - break here and all is lost

Three architectural rules, more important than marking the cache

  • Freeze the system prompt. Don't put date, time, mode, or user name in it. If you need to inject dynamic context, inject it at the end of the message list - a message at turn 5 breaks nothing before turn 5.
  • Don't switch tools or model mid-stream. Tools sit at the very front; adding, removing, or reordering a tool wipes the whole cache. Cache is also bound to the model, so switching models loses it.
  • Secondary calls must reuse the exact front of the main call. Summarization, context compaction, sub-agents - if you rebuild the system or tools even slightly differently, you miss the parent call's cache entirely.

A quick way to check Run two back-to-back requests with an identical front, then look at cache_read_input_tokens (Claude) or cached_tokens (OpenAI). Still 0 means something is definitely breaking the prefix - diff the bytes of the two built prompts to find it.


Part 4 - Reference

Data tables

Windows and prices by model

ModelWindowInput $/1MOutput $/1M
Claude Opus 51M$5.00$25.00
Claude Sonnet 51M$3.00$15.00
Claude Haiku 4.5200K$1.00$5.00
GPT-5.6 Sol~1.05M$5.00$30.00
GPT-5.6 Terra~1.05M$2.00$12.00
GPT-5.6 Luna~1.05M$0.20$1.20

Claude prices from official Anthropic docs. GPT-5.6 prices and windows from third-party aggregators (Aug 2026) - re-check the official pricing page before costing anything real. GPT-5.6 also has a long-context tier: past 272K input tokens the whole request is billed at about 2x input, about 1.5x output.

Claude minimum cache thresholds

ModelMinimum prefix
Claude Opus 5, Fable 5, Mythos 5512
Opus 4.8, Sonnet 5, Sonnet 4.6 / 4.5, Opus 4.1 / 4, Sonnet 41,024
Opus 4.7, Mythos Preview, Haiku 3.52,048
Opus 4.6, Opus 4.5, Haiku 4.54,096

The trap most worth remembering in this whole piece This threshold does not rise evenly across model generations. A 3,000-token prompt caches on Opus 5 and Sonnet 4.5, but silently does not cache on Opus 4.6 or Haiku 4.5. Below the threshold there is no error - just cache_creation_input_tokens: 0 and a bill 10x higher.

Claude cache rules

ItemRule
Render ordertools -> system -> messages
Max cut points4 per request
TTL5 minutes (default) or 1 hour
Write price1.25x (5 min) / 2.0x (1 hour)
Read price0.1x
Break-even2 requests (5-min TTL) / 3 requests (1-hour TTL)
Look-back windowUp to 20 content blocks - tool-heavy turns can exceed this and miss silently
Parallel requestsCache is only readable after the first response starts streaming
// The formula to remember when reading Claude's usage
total prompt = input_tokens                  // the uncached part
             + cache_creation_input_tokens   // the part just written to cache
             + cache_read_input_tokens        // the part read from cache

// Looking at input_tokens alone and assuming a small context is a mistake.

What change breaks which cache (Claude)

Changetools cachesystem cachemessages cache
Tool definitions (add/remove/reorder)lostlostlost
Switch modellostlostlost
System prompt contentkeptlostlost
tool_choice, images, toggling thinkingkeptkeptlost
Message contentkeptkeptlost

Which means changing tool_choice or toggling thinking per request does not lose the tools and system cache at all. Only switching tools and switching models forces a full rebuild.

OpenAI cache rules

ItemRule
How to enableAutomatic, no declaration
Minimum threshold1,024 tokens (GPT-5.6+). Older models: 1,024 - 2,048
Prefix incrementOlder models: in multiples of 128 tokens
Read price0.1x (90% off)
Write price1.25x on GPT-5.6+
TTLGPT-5.6+: only 30m. Older models up to 24 hours
Real expiryUsually cleared after 5-10 minutes idle, at most 1 hour
Reporting fieldcached_tokens in prompt_tokens_details / input_tokens_details

Don't trust old numbers Many articles online still say OpenAI gives a 50% discount - that was the launch figure in 2024. It is 90% now. When it comes to pricing, always check the date of the source.

How ChatGPT differs from the API

A common question: "On ChatGPT I never resend anything, so how does it still remember?"

Because ChatGPT is an app built on the API, and it handles the resending for you. Their server stores the conversation, reassembles it every turn, and sends it to the model. From the model's point of view, the whole conversation is still resent every turn - exactly like your bot.

Direct APIChatGPT
Who keeps the historyYouOpenAI
Who decides what to trim when it's longYouOpenAI
Does the model see the whole conversation?Yes, the part you sendYes, the part they assemble
You pay byTokenSubscription
Can you see the token count?Yes, in usageNo

When a ChatGPT conversation exceeds the window, they trim or summarize the old part - which is why it sometimes "forgets" what you said at the start of a very long conversation. The exact mechanism you have to write yourself for your own agent.

About the per-plan window ChatGPT Free / Plus / Pro windows are smaller than the model's max window on the API, and the numbers going around online contradict each other - third-party sources, changing constantly, differing by fast-reply versus deep-thinking mode. If you need an exact number, check OpenAI's official help page at that moment.

Mapping to zalo-agent

This piece was originally written to explain token costs for zalo-agent - an open-source (MIT) agent that runs on Zalo. The table below maps the concepts above onto their exact places in the repo, in case you want to see a real implementation.

Concept in this pieceIn the repo
Context windowLLM_CONTEXT_WINDOW, default 128,000
Reserve room for the replynganSachAnToan() - uses only up to 70%
Trim before sendingDrops old images first, then old messages
Count tokens, not messagesA single Zalo message can be any length; counting messages can't cap context growth
Cap on reloaded imagesHISTORY_IMAGE_CONTEXT_LIMIT
Enable cachingA stable x-session-id header per thread
Session cumulative tokensThe total_tokens column in the agent_turns table
Cap on steps per turnLLM_MAX_STEPS

Reading the three numbers on the dashboard

Where you see itWhich kindCompare to window?
Sessions: "6 messages - 84,601 tokens"Whole sessionNo
Trace, turn-card label: "2 steps - 21,653 tokens"One turnNo
Trace, inside a step: "10,600 in / 196 out"One requestYes - this is it

Glossary

TermMeaning
TokenThe unit a model reads and writes. Vietnamese costs more tokens than English for the same idea - don't estimate by word count.
Context windowThe maximum capacity of one request. Includes the part the model is about to write.
StatelessKeeps no state. The server forgets everything after each call.
StepOne API call inside a turn. When the model asks for a tool, the next step begins.
Agent loopThe loop: call the model, model asks for a tool, code runs the tool, sends the result, repeat until the model stops asking.
PrefixThe front of the prompt. Cache matches here.
BreakpointA marker for "cache up to here". Claude only.
TTLHow long a cache stays alive if nobody reuses it.
Cache write / readCreating the cache the first time (more expensive) and reusing it (much cheaper).

How this piece is built. Structured after Diataxis - separating the explanation (Parts 1 to 3, read in order) from the reference (Part 4, jump straight in). Presented on Mayer's principles: define the vocabulary before using it, one idea per section, put each diagram right next to the paragraph it illustrates, and use a consistent color for the three kinds of tokens throughout.

Sources. The Claude parts are from official Anthropic docs. The OpenAI parts are from the official prompt-caching guide and the feature announcement. GPT-5.6 prices and windows, and the ChatGPT-plan numbers, are from third-party aggregators in August 2026 and are marked as such in place. Every measurement in the example section is copied directly from a running bot's trace, not a made-up example.

Related articles

  • A dedicated Chrome for chrome-devtools-mcp that can sign in

    Set up a per-project Chrome profile for chrome-devtools-mcp (Claude Code, Cursor...) to drive: it can still sign in to Google/Cloudflare and keep the session, run several projects in parallel, and never touch your personal Chrome.

    Dev tools

    Dev tools
  • Fix Claude Code connection errors on Windows (ECONNREFUSED, Bun crash)

    Fully resolve two common Claude Code errors on Windows - can't connect to the API (ECONNREFUSED) and a Bun crash (Internal assertion failure) - by removing the NPM build, clearing config caches, and installing the Native Stable build.

    Dev tools

    Dev tools

Written by Vu Van Hai

I'm Hai, a full-stack developer based in Ho Chi Minh City. These guides come from systems I built and run myself. Need to build or untangle something similar? Get in touch.

Discuss a projectMore posts

Spot a mistake or a command that no longer works? Let me know

On this page

  • Five words to learn before you read on
  • One API call: the model remembers nothing
  • Many chat turns: you resend everything, every time
  • What it means for one turn to have many steps
  • Four common questions, answered straight
  • Working one real turn step by step
  • Two things to read out of these numbers
  • The context window: the capacity of a single call
  • Which box is the ring in Claude Code
  • The shape of the two lines says it all
  • Three proofs it is not a cumulative number
  • The biggest item people forget
  • Prompt caching: you pay again, but 10x cheaper
  • What it does
  • Three kinds of tokens, three prices
  • The most important rule: it matches by PREFIX
  • Where the two providers differ
  • Things that break the cache and tell you nothing
  • Three architectural rules, more important than marking the cache
  • Data tables
  • Windows and prices by model
  • Claude minimum cache thresholds
  • Claude cache rules
  • What change breaks which cache (Claude)
  • OpenAI cache rules
  • How ChatGPT differs from the API
  • Mapping to zalo-agent
  • Reading the three numbers on the dashboard
  • Glossary