Context and cost
Why a session gets more expensive and less capable the longer it runs, and the five moves that fix it — /rewind, /clear, /compact, the handoff pattern, and delegating to subagents.
Two things happen as a session grows, and they’re the same thing seen from two sides: it gets more expensive, and it gets worse.
More expensive, because every turn re-reads the whole conversation. Your thirtieth message isn’t priced like your first — it carries all twenty-nine before it. Cost compounds rather than adds, which is why a long session can spend most of its tokens re-reading its own history rather than doing anything new.
Worse, because attention is finite. A model with a full window has to find what matters inside a lot of noise, and retrieval gets less reliable as the window fills. The symptoms are recognisable: it forgets a decision you made earlier, contradicts itself, edits a file it never read, gets vague.
The practical conclusion is the one people resist: a big context window is insurance, not a target. Room to work is not a reason to fill it.
Know your starting cost
A fresh session is never at zero. The system prompt, CLAUDE.md, every skill, and every connected
MCP server’s tool definitions are all resident before you type anything.
Run /context in a brand new session and look at the number. If it’s already large, that’s a tax
on every conversation you will ever have in this project, and the fix is upstream — trim CLAUDE.md,
drop MCP servers you don’t use, move rarely-needed instructions into files that get loaded on demand.
Prompt caching
The re-reading described above has a discount, and it’s the reason long sessions are affordable at
all: prompt caching. The unchanged prefix of your conversation — system prompt, CLAUDE.md,
tool definitions, every earlier turn — is cached, and a cached re-read costs roughly 10% of fresh
input. It’s automatic in Claude Code; there is nothing to configure. But knowing what breaks it
turns several mysterious cost spikes into avoidable ones.
The clock
A cache entry lives for a limited time after the last request:
| Where you’re running | Cache window |
|---|---|
| Claude Code on a subscription | ~1 hour |
| API keys, and subscription overage (paying per token) | 5 minutes |
Walk away from a session for longer than the window and your next message re-processes the entire conversation at full price — the single most common invisible cost spike. Coming back to a stale session is economically the same as starting over, minus the clean context. If it’s been over an hour, prefer the handoff pattern into a fresh session.
What breaks the cache
The cache is a prefix match: any change near the start of the conversation invalidates everything after it.
| Action | Effect |
|---|---|
Switching model mid-session (/model) |
Full re-cache — each model has its own cache, so the next request reads the whole history with no hits |
| The opus-plan model setting | Resolves to Opus in plan mode and Sonnet in execution, so every plan-mode toggle is a model switch and starts a fresh cache. It can still pay off overall — just know the cost |
| An MCP server connecting or disconnecting mid-session | Tool definitions sit in the prefix, so a stdio server’s process exiting or an HTTP session expiring re-caches — without any action on your part |
| Pausing past the cache window | Everything re-processes at full price |
Editing CLAUDE.md mid-session |
Safe — the edit only applies when the session restarts, so the in-flight cache holds |
If you’re watching a token dashboard: cache create is the one-time cost of writing something into the cache; cache read is the ~10× cheaper reuse. A healthy long session shows large cache reads and modest fresh input.
The five moves after any reply
Once Claude finishes a turn, you always have five options, and picking deliberately is most of the skill:
| Move | When it’s right |
|---|---|
| Continue | The thread is healthy and you’re still on the same problem. The default, and the one that’s overused |
/rewind |
Something went wrong and you want it gone, not just corrected |
/clear |
New task. Nothing before it helps |
/compact |
Same task, long history, and you want a summary to carry forward |
| Delegate to a subagent | The next step is bulky but self-contained — reading a lot, searching, verifying |
Prefer /rewind over “that didn’t work, try again”
This is the habit with the best return. When Claude gets something wrong, the natural reply is “that didn’t work, try X instead” — and it usually does work, so the approach feels fine.
But the failed attempt is still there. The broken code, the wrong turn, the dead end: all of it stays
in the conversation and gets re-read on every subsequent turn, forever, at full price. /rewind (or
Esc twice) jumps back to a chosen message and drops everything after it, so the mistake leaves the
context instead of being buried in it.
The rewind menu also offers Summarize from here, which writes a handoff note before dropping the tail — a message from the session’s future self to its past self saying what was learned.
The handoff pattern
/compact is fine, but it fires on the model’s terms. The alternative gives you control of both the
timing and the contents:
Ask for a handoff
“Give me a full summary of everything we’ve done, the decisions we locked in, the files that matter, and what we’re about to do next.”
Copy the output
Including its pointers to plan files, decision logs, and task lists.
/clear
A genuinely fresh window.
Paste it back
The new session reorients from the summary and keeps going.
You get the reset without feeling like you reset. It only works if the durable state lives on disk — a plan file, a decision log, a task list — rather than in the conversation. Chat history is not storage.
This is also worth turning into a skill so it’s one command rather than a paragraph you retype.
Compact before the cliff, not at it
Auto-compaction fires when you’re nearly out of room — which is precisely when the model is least equipped to decide what’s worth keeping. Compacting deliberately, well before that point, produces a better summary for the same reason packing the night before beats packing five minutes before you leave.
/compact takes instructions, and it should always get them:
/compact keep the schema decisions and the list of remaining tasks
You can also move the automatic trigger yourself:
claude --autocompact 200000 # or `auto` for the default behaviour
Side questions: /btw
/btw opens a quick side conversation that doesn’t enter your main history. Use it for the “wait,
what does this flag do?” questions that would otherwise wedge an unrelated topic into the middle of
a working session.
Feed it the cheapest form of the content
Tokenizers are efficient with plain text and wasteful with layout. A PDF, a .docx, or a page of
HTML carries markup, styling, and structural metadata that the model does not need in order to read
the words.
Converting to Markdown first is usually a large reduction — enough that it’s worth doing routinely for anything long. The exception is when the layout is the content, or you need OCR or vision; then give it the original.
Session chaining
A big project does not have to be one session. Split it along the natural seams and hand off between them:
Discover
Read the codebase and the docs, produce a summary document.
Plan
Read that document, produce a plan.
Build
Read the plan, execute it.
Each session starts fresh and carries only the artifact it needs. The alternative — one session that does all three — pays to re-read the discovery phase on every turn of the build phase.
Habits, in order of payoff
Work in bursts, inside the cache window
A session you keep touching stays cheap — every turn rides the cache. A session you poke once an hour re-pays for its whole history each time. Batch the work; when you’re done, hand off rather than leaving the session to go stale.
Pick a reset number and hold to it
Decide in advance what fraction of the window you’ll work within, and treat crossing it as the signal to hand off. An arbitrary line you actually respect beats a principled one you ignore.
Start in plan mode
Tokens spent getting the plan right are the cheapest tokens in the session. Tokens spent correcting an implementation that misunderstood the goal are the most expensive.
Delegate bulky reading to a subagent on a cheaper model
The subagent absorbs the tokens; your session gets the summary. See tips.
Keep CLAUDE.md small
It loads every session, so its size is a recurring charge rather than a one-off.
When a session feels off, start a new one
Not every bad session is a context problem, but starting fresh is cheap and usually faster than arguing your way back to a good state.