← 전체 글

10 Claude Code Mistakes That Are Using Up Your Tokens

psdjung· 2026년 9월 28일· 15분 소요· Technology
Illustration of a large pile of cash and gold dollar coins streaming toward the Claude app icon

Short answer: Most of what Claude Code spends goes on re-reading context, not writing code. A session can be carrying more than 100,000 tokens before you type a word, and switching models or coming back after a break makes it re-process all of it. Trim what loads at startup, pick your model once per session, compact at natural breaks, filter logs before Claude sees them, and stop repeated retries early.

Context is everything Claude Code sends to the model on each call: the system prompt, your instruction files, the tool and plugin definitions, and the whole conversation so far. Claude Code "sends your full conversation with every request", as Anthropic's cost guide puts it. Prompt caching makes most of that re-reading cheap, but only while the cache holds, and several everyday habits quietly break it.

Why is your Claude Code limit running out faster this month?

From May 13 to September 13, 2026, Anthropic raised Claude Code's weekly usage limits by 50% for Pro, Max and Team plans (Claude support). Since September 14, weekly limits are 25% above where they were before the promotion. If your habits formed during the promotion, you now have noticeably less room each week. The fixes below are about getting that room back without doing less work.

What are the Claude Code mistakes that quietly burn your limit?

1. You're spending 100,000+ tokens before you type anything

Open a new session, don't type anything, and run /context. It lists everything loaded before your first message: system prompt, tool definitions, MCP servers, plugins, skills, memory and instruction files.

On one of our developers' machines, the first call of an empty session sent 126,683 tokens. Most of it had nothing to do with the code: plugins for legal, sales, finance, email and calendar, installed once for other work and loaded into every session since. With those plugins turned off for the coding project, the first call dropped to 70,117 tokens, 45% less. That saving applies to every call in every session.

How to fix it:

  1. Run /context in an empty session and note the total.
  2. Switch off the plugins this project doesn't use in the /plugin menu, and the MCP servers it doesn't use in /mcp.
  3. Make these changes at the start of a session. Enabling or disabling a plugin that brings MCP servers mid-session can make the next request re-read the whole conversation.
  4. To see the difference quickly, start a session with only the project's own settings: claude --setting-sources project,local. Then run /context again and compare.

2. Your 19-line CLAUDE.md is really 95 KB

Run wc -l CLAUDE.md ~/.claude/CLAUDE.md and it looks harmless. Now look inside for lines starting with @: each one imports another file, and those files load too. On the same machine, a 19-line global CLAUDE.md imported eight framework files totalling about 95 KB, roughly 24,000 tokens, sent with every request in every project.

Anthropic recommends keeping CLAUDE.md under 200 lines and moving workflow-specific instructions into skills, which load only when they're needed.

How to fix it:

  1. List the imports: grep '^@' CLAUDE.md ~/.claude/CLAUDE.md.
  2. Check their real size with wc -c on each imported file. Roughly 4 bytes is one token.
  3. Keep only what every session needs in CLAUDE.md. Move workflow-specific instructions into skills, and point to longer documents with a plain link instead of an @ import.
  4. Move global imports that only one project needs into that project.

3. You came back from lunch and paid to re-process the whole conversation

Prompt caching expires: after about an hour on a subscription, and after 5 minutes by default on the API or once you're drawing on usage credits. After a break, your first message "reprocesses your full context". Leave a 600,000-token session for lunch, type "ok, continue", and that one message processes all 600,000 tokens again.

How to fix it:

  1. Before a break of an hour or more, run /compact with a note on what to keep, such as /compact keep the migration plan and open questions. Compacting while the cache is still warm costs a fraction of compacting after it expires.
  2. If the task is finished, run /clear instead. It costs nothing.
  3. On Pro and Max, when you return to a large session after a long break, accept Claude Code's offer to resume from a summary.
  4. On the API, where the cache lasts 5 minutes by default, consider setting promptCacheTtl to 1h if you often step away mid-task.

4. Your retries cost more than your code

When we studied our agent runs, cost tracked the number of attempts, not the size of the change. The same one-line change cost $0.96 in one run and $7.35 in another. One goal spent $110, almost all of it re-reading cached context across repeated attempts.

The pattern is familiar: a test fails, you say "try again", it tries something similar, and every attempt re-reads a context that keeps growing.

How to fix it:

  1. Set yourself a rule: two failures for the same reason means stop.
  2. Run /rewind to go back to before the first attempt. Rewinding drops the failed attempts and lands on a prefix that is already cached.
  3. Restate the problem with what the failures taught you, such as the exact error and what didn't work.
  4. If the session is already large, /clear and start the fix fresh with that summary.

5. You switched models in the middle of a session

This one surprises almost everyone. "Each model has its own cache," Anthropic's prompt caching guide explains, so after a /model switch "the next request reads the entire conversation history with no cache hits, even though the content is identical." Switch from Opus to Sonnet deep into a 500,000-token session to save money, and the first Sonnet message pays to process all 500,000 tokens from scratch.

It happens in less obvious ways too. The opusplan setting uses Opus in plan mode and Sonnet for execution, so every plan-mode toggle is a model switch. On most models, changing the effort level mid-session also starts a fresh cache. A skill whose settings name a different model triggers a switch for that turn.

How to fix it:

  1. Choose your model and effort level at the start of a session, and leave them there for the task.
  2. When a new task needs a different model, start a new session or /clear first, rather than switching in the middle.
  3. If Claude Code asks you to confirm a model switch, that means the cache is still warm and the switch will cost a full re-read. Say no unless the task really needs it.
  4. Use opusplan for tasks with one planning pass and a long execution, not ones where you go in and out of plan mode repeatedly.

6. You're paying for thinking you didn't need

Thinking tokens are billed as output tokens, the most expensive kind, and the default budget can reach "tens of thousands of tokens per request". Renaming a variable or updating a README doesn't need deep reasoning.

How to fix it:

  1. Before you start, check the model with /model. Anthropic's guidance is that "Sonnet handles most coding tasks well and costs less than Opus."
  2. Set a lower effort with /effort for sessions of routine work, at the start of the session (see mistake 5).
  3. For subagents that only search or summarise, set model: haiku in the subagent's configuration. Subagents keep their own cache, so this doesn't disturb your main session's.
ModelInputCache readOutput
Opus 5.5$4$0.20$20
Sonnet 5$2$0.20$10
Haiku 4.5$1$0.10$5

API list prices per million tokens, from Anthropic's pricing page.

7. You pasted the whole test log

You run the test suite, 4,000 lines scroll by, and you paste them in with "fix this". Those 4,000 lines are now re-sent on every call for the rest of the session.

How to fix it:

  1. Paste only the failures: npm test 2>&1 | grep -E "FAIL|Error" | head -50. Anthropic notes this kind of filtering can cut tens of thousands of tokens to a few hundred.
  2. Better, ask Claude to run the tests itself with a filtered command, so the full log never enters the conversation.
  3. For a permanent fix, add a hook that trims test and build output before Claude sees it. Anthropic's cost guide includes a ready-made example.
  4. Prefer command-line tools such as gh over an equivalent MCP server. Anthropic says they are more context-efficient.

8. You're misusing subagents

Subagents are great for noisy lookups, such as "find every place we call the billing API", because only a summary comes back to your main conversation. A subagent that carries out an entire task builds up its own large context instead. Agent teams use "approximately 7x more tokens than standard sessions" when teammates run in plan mode, according to Anthropic.

How to fix it:

  1. Give subagents questions with short answers: "find every call to the billing API and list the files".
  2. Keep the actual change in your main session, where you can see and steer it.
  3. Use agent teams only when the work truly splits into independent pieces, and check /usage afterwards to see what they cost.

9. You fell into the 1M-context trap

A 1M-token window feels like room to spare. On those models, auto-compact triggers at roughly 967,000 tokens by default (model configuration), and until then everything is re-sent on every call. In our own longest sessions, context averaged 600,000 to 760,000 tokens per call near the end, and single calls passed 930,000.

How to fix it:

  1. Watch your context as you work: /context shows it, and you can add it to your status line.
  2. Run /compact at natural breaks in a long task, such as after the plan is agreed or after tests pass.
  3. Lower the automatic threshold with /autocompact, for example /autocompact 300k, so it never drifts toward 1M.
  4. Don't overdo it: compaction is itself a request, so once per phase beats constantly.

10. You use one session for everything

It feels like the session remembers everything for you, and it does, at a price: moving to a new task in the same session means every new request carries every old one. This is the mistake everyone has heard of, and it is still the biggest. In one of our long sessions, the model re-read 503 million cached tokens and wrote about 1 million. "/clear costs nothing," as Anthropic puts it.

How to fix it:

  1. One task, one session. Run /clear or open a new session when you switch to unrelated work.
  2. Before clearing, run /rename so you can find the session again with /resume, and ask Claude for a short summary of decisions and open items to carry forward.
  3. Keep long-lived knowledge in CLAUDE.md or project docs, not in a conversation you're afraid to close.

What does a 5-minute token checkup look like?

  1. Run /context in an empty session and note the starting size.
  2. Switch off plugins and MCP servers the project doesn't use, then run /context again.
  3. Run wc -l on your project and global CLAUDE.md, and list their @ imports.
  4. Decide your default /model and /effort for the work you do most, and set them before you start.
  5. Run /usage at the end of the day. Its prompt cache line shows your cache hit rate and misses, and on a plan it breaks usage down by skill, subagent, plugin and MCP server.

What else can you do to cut token use?

Fixing habits gets you most of the way. These go further, from safest to riskiest.

  1. Route work between Claude models, per session. Sonnet or Haiku for routine work, Opus only when the problem needs it. Decide at the start of a session, not in the middle (mistake 5).
  2. Let a tool pick the model. If you also use GitHub Copilot, its Auto model selection chooses a model for each request and gives paid plans a 10% discount on model costs.
  3. Try an open-source tool built around small context. Aider, a terminal coding assistant, sends a ranked "repo map" of only the most relevant parts of your code, with a default budget of 1,000 tokens. Its architect mode has a stronger model plan and a cheaper editor model write the edits. It works with Claude and other providers.
  4. Point Claude Code at other models, with care. Proxies such as LiteLLM or the community project OpenCodex let Claude Code send requests to other providers through ANTHROPIC_BASE_URL. Before you do, know that:
    • It isn't supported. Anthropic "doesn't support routing Claude Code to non-Claude models through any gateway" (LLM gateway docs), so some features may not work.
    • It doesn't stretch your subscription. With a gateway credential set, "the subscription's usage limits don't apply": you pay per token to whoever owns that credential.
    • You mustn't share your Pro or Max login through it. Anthropic's terms say developers may not "route requests through Free, Pro, or Max plan credentials on behalf of their users" (legal and compliance).
    • Automated account checks can misfire. In August 2026, a developer who ran Claude Code against an OpenAI model through a small proxy had his account suspended. Claude Code lead Boris Cherny said Anthropic doesn't ban people for using its harness with other models, and the account was restored (AI Times).
    • A proxy sees all your code and prompts. Use a well-known project, pin its version and keep secrets out of its reach.

Don't want to worry about all this? Let Interactor Build take care of it

Every fix above depends on you staying disciplined, in every session, every day. That's the part Interactor Build takes off your plate. You describe the feature you want, and Build plans the work, writes the code on Claude Code, reviews it and ships it as a single pull request, with the habits above built into how every run works:

  • No 100,000-token start (mistakes 1 and 2). Every run starts in a clean container: no personal instruction files, plugins or memory, just your repo's own instructions and Build's own connection.
  • No lunch-break penalty (mistake 3). A run ends as soon as its step is done, so nothing sits idle holding a large context. Context that stays the same across a goal is sent as an identical prompt prefix, so a goal's steps can share the prompt cache while it's warm.
  • No runaway retries (mistake 4). Each goal has a run budget, 15 runs by default, and circuit breakers stop it when it keeps failing the same way or passes a spending ceiling.
  • No mid-session model switches (mistakes 5 and 6). The model is chosen when each run starts, so changing models never throws away a warm cache. Error analysis, feedback review and discussion tasks default to Haiku, and you can set the model for each task type and phase.
  • No session that grows forever (mistakes 9 and 10). Every phase of every task (investigation, execution and review) is a fresh run. What carries over is written down, such as the plan, the commits and earlier review findings, not a growing conversation. When a change goes back for another review, only the new commits are checked.
  • No paying while you wait on CI. CI, the merge queue and re-review run on the server, and goals waiting their turn sit untouched until they reach the front of the queue.

The result is small steps. Across 4,714 Build runs in the two weeks to September 27, 2026, the median context per model turn was about 81,000 tokens in investigation, 102,000 in execution and 85,000 in review. Our own long Claude Code sessions, by comparison, were carrying 200,000 to 760,000 tokens per call near their end.

What does Interactor Build cost per goal?

For goals Build completed in the 30 days to September 27, 2026, the median cost was about $63 per goal at Claude API list prices, and 90% of goals cost less than $268.92. Smaller steps don't automatically make every goal cheaper, because Build also spends tokens on work a single session usually skips: a planning step, an independent review of every change and rework when CI fails. We don't yet have a like-for-like figure for the same goals done by hand in Claude Code, so we won't claim a percentage saving. What changes is where your tokens go: into planning, review and verification instead of into re-reading an old conversation.

If you'd rather spend your attention on the feature than on managing context, you can run a pilot with your own goals and see where your tokens go.

Frequently asked questions

Does switching models in Claude Code use more tokens?

Yes. Each model has its own prompt cache, so the first request after a /model switch re-processes the entire conversation with no cache hits. Pick your model at the start of a session, or /clear before switching.

Does /compact cost tokens?

Yes. /compact sends the conversation to be summarised, so it is a request of its own. While the cache is warm it costs a fraction of the context size; after a long break it re-processes the full history. /clear costs nothing.

Why does my first message after a break use so much?

The prompt cache expired. After about an hour on a subscription, or 5 minutes on the API, your next message re-processes the whole context. Compact or clear before a long break.


Figures come from Claude Code transcripts on our own development machines and from Build's run records, measured on September 27, 2026. Costs are Claude API list prices reported by the Claude CLI, not what was billed.

AIClaude CodeToken UsageAI CostInteractor Build

댓글