The recurrence
You enter three token counts. Stable is the system prompt and the tool schemas. Added is the new tool result and any new instruction on a step. Output is what the model writes on that step, including reasoning tokens if you want the estimate to look like a bill.
Step k, counting from 1, sends this many input tokens:
stable + k × added + (k − 1) × output
Then it produces output tokens, which the next step has to carry.
Try it with small numbers. Stable 10, added 4, output 2, three steps:
- Step 1 sends 14.
- Step 2 sends 20.
- Step 3 sends 26.
Sixty prompt tokens, not 3 × 14. The last step is already almost twice the first. Eight steps with a long tool result is where a flat “per step” quote goes quiet.
A hand check, with toy prices
These prices are not a vendor. They are round numbers so the arithmetic is visible. Input is $1 per token, cache read is $0.10, a 5-minute cache write is $1.25, and output is $2.
Cache off, the run bills all 60 prompt tokens and 6 output tokens: 60 × $1 + 6 × $2 = $72.
Cache on, in the way Looptally models a write premium: step 1 writes all 14 tokens. Each later step reads the previous prompt and writes only the 6 new tokens (the new chunk plus the previous output). Writes are 14 + 6 + 6 = 26. Reads are 14 + 20 = 34. Cost is 26 × $1.25 + 34 × $0.10 + 6 × $2 = $47.90.
One step, cache on, is the other way around. You pay the write premium and never read it back. On these toy prices a single 14-token prompt is $14 with cache off and $17.50 written to a 5-minute cache. The receipt stamps “cache costs extra” when that happens to your numbers.
Two cache shapes, not one
Write premium. Anthropic publishes a 5-minute write, a 1-hour write, and a cache read. Some OpenAI models publish a single cache-write column. Looptally uses the write price on tokens entering the cache and the read price on tokens already there. The break-even line on the receipt uses the short-context standard rates: how many cache reads it takes before the extra write cost is covered. For a normal Anthropic 1.25× write and 0.1× read, that is a fraction of one read. Claude Fable 5.1’s read is $0.25 per million, not 0.1×, so its break-even is computed from $0.25, not from the usual multiplier.
Read discount. Some OpenAI rows list a cached-input price and no write price. Gemini lists a context-cache price plus storage per million tokens per hour. The first step pays the normal input rate. Later steps read the previous prompt at the cache rate and pay the input rate on the new tokens. Storage, when the book has a number, is the largest prompt times the hourly rate times the hours you set. It is not charged when cache is off.
If a batch row has no cache-read cell, Looptally does not invent one. You get a warning, and those tokens stay at the input rate.
Long context
When a model publishes two columns, the whole step uses the long-context rates once that step’s prompt crosses the threshold. Earlier steps can stay on the short rate. OpenAI’s short rate is for context under 272,000 tokens, so 272,000 itself uses the long column. Google’s higher Gemini rate is for prompts over 200,000 tokens, so 200,000 itself stays on the lower rate. Both switches are labeled on the price page.
What this will not match on an invoice
- Retries, failed tool calls, and the wrapper your SDK adds around the prompt.
- Reasoning tokens you did not type into the output field. Providers bill them even when the response hides them.
- A minimum prefix length. If the vendor only caches prompts above a cutoff, a small stable prefix will not get the cache rate. Looptally still applies the cache rate when the switch is on. Turn the switch off to see the uncached bill.
- Fast mode, priority, flex, and US-only data residency. Several of those are a multiplier on the whole request. They are not in this total.
- Grounding, tool-call fees charged per search, image tokens, and audio tokens. Text rates are what the book uses, and the model note says when audio is priced separately.
- Batch cache math, beyond the cells in the book. Anthropic’s batch cache cells are the published cache prices with the stated 50% batch discount stacked, and the note says that. A blank cell is not a guess.
The 30-day figure is this run, times runs per day, times 30. It is not a calendar month and it is not a forecast of your traffic.