← Effective AI Code Development

See Claude Code Token Usage Without Spending Tokens

Add two local scripts to Claude Code that show a status line and a Stop hook with tokens, context size and approximate cost. Reporting costs zero model tokens because it reads local session data.

·6 min read·
#Claude Code#Hack#Hooks#Token Usage#Developer Productivity

TL;DR

  • This setup is a Claude expert user hack that can help you watch your token at the end of each response run.
  • Because Claude Code only shows token usage when you run /context or /cost, so most of us may not run it during session runs.
  • So reporting a status line and a Stop hook can show the exact numbers all the time.
  • Both are small local scripts, so they use no model tokens and the numbers are exact.
  • Setup is calling two scripts and one change to your settings file

Why do we need this hack?

Claude Code does not show you how many tokens you have used unless you ask by running /context or /cost. When you are busy debugging or building something, you forget to run them, so you only notice the problem after the context has already grown too large.

The first idea most people have is to add an instruction such as "print the token usage at the end of every reply". I avoided this for two reasons. First, the model has to write that summary, so you pay output tokens on every single turn just to be told how many tokens you used. Second, the model is not given an exact count of the tokens used in each request, so whatever it prints would be its own estimate and not the real number.

The replies have become slow, the answers have become less accurate, or a usage warning has appeared. By that point the context is already too large and the tokens have already been spent. I wanted to see the numbers while I was working, not after the problem had happened.

The Claude Code program on your machine does have the exact numbers, because every request and response passes through it, and it writes them to a local session file. It can also run a script of your choice after each reply, which is what a hook is. The script reads those exact numbers and shows them to you. The model is never called, so the reporting costs no tokens and is accurate.

To do this I added a status line and a hook to Claude Code. Together they show the exact token usage all the time, and they run as local scripts, so they use no model tokens at all.

The two parts of the setup

The first part is a status line. This is a script that Claude Code runs to draw one line at the bottom of the terminal. My script reads the session file and shows the model name, the current context size, the total output tokens, the number of model calls and the cost:

Sonnet 5.5 | ctx 146.5k | out 7.3k | calls 15 | $1.01

The second part is a Stop hook. This is a script that runs each time Claude finishes a reply. My script adds up the usage since your last prompt and prints one line with the result:

tokens this request: in 6 | out 1742 | cache-read 190957 | cache-write 88736 | api-calls 2 | context now ~140871

Both scripts are short and use jq. Claude Code passes them the path to the session file. They read the token usage recorded for each reply and add it up, counting each reply only once.

How to set it up

The setup uses two scripts, statusline.sh for the status line and token-report.sh for the Stop hook. Both are in this gist, together with a short install guide. Download them, copy them into ~/.claude/hooks/ and make them executable. Then add the following to ~/.claude/settings.json:

{
  "statusLine": { "type": "command", "command": "~/.claude/hooks/statusline.sh" },
  "hooks": {
    "Stop": [
      { "hooks": [ { "type": "command", "command": "~/.claude/hooks/token-report.sh", "timeout": 10 } ] }
    ]
  }
}

Add these entries to your existing settings and do not replace the whole file. If the settings file contains a mistake, Claude Code silently ignores every setting in it, so check the file with jq -e after you edit it. I also tested each script by feeding it sample input before I relied on it, including empty input and a machine where jq was missing.

What the numbers mean

  • ctx is the size of the most recent prompt, including content that was read from the cache. It grows as the conversation gets longer, and it drops after you run /clear or the conversation is compacted. It tells you when it is time to start fresh.
  • out is the total number of output tokens in the session so far.
  • cache-read is usually the largest number, because the whole conversation is read from the prompt cache on every call. Cached tokens are billed at a lower rate than new input.
  • The dollar amount comes from Claude Code and is an estimate. It is not an invoice.

Why every daily user should have it

If you use Claude Code every day, I think you should have this. The built-in commands only help when you remember to run them, and in the middle of a debugging session nobody does. A number that is always on the screen helps in four ways:

  • You clear the context at the right time. When you see ctx growing, you can run /clear or compact the conversation before the answers get worse, instead of after.
  • You see what each habit costs. A long exploratory prompt, a large file read or a heavy MCP server causes a visible jump in the numbers, so you learn which patterns are expensive.
  • You are not surprised. Whether you pay per token or work within a plan limit, you know where you stand before a limit or a bill tells you.
  • You can tune your setup with real data. It is easy to justify trimming a large CLAUDE.md file or removing an unused MCP server when you can watch the context number go down.

It takes a day or two to get used to seeing a dollar amount on the screen. Think of it as the fuel gauge in a car. It helps you make decisions, and it is not a reason to stop driving. If it distracts you, keep the status line and remove the Stop hook, because the status line is the quieter of the two.

Limitations

  • The scripts read the session file, which is not a documented or stable interface. A future Claude Code update could change the field names and break the numbers.
  • The status line reads the whole session file every time it refreshes, so it may feel slower in a very long session. Saving the running totals between refreshes would fix this.
  • The per-request totals may not include calls made by subagents.
  • I have only tested this on macOS.

Before you rely on any field, check the official Claude Code documentation on status lines and hooks for the current input format.

Conclusion

  • Asking Claude to report its own token usage costs tokens and gives estimates. But adding a status line and a Stop hook give you the exact numbers for free, because they are local scripts that read what Claude Code already records.
  • Seeing the token usage and approximate cost for each request helps you clear the context before quality drops, find the prompt patterns or habits that cost the most, and avoid surprises.
  • If you use Claude Code every day, it is a small change that is worth making.

The scripts and a short install guide are in this gist.

Public profile lookup

Ask AI About the Author

Open this query in ChatGPT, Claude, or Perplexity.

Comments

Comments are open to confirmed email subscribers. Use the email you subscribed with. To edit a comment, delete it and post a new one.

0/2000
Verify:

    Subscribe to get the new blogs.

    Field notes from someone who ships before they write about it. Sovereign AI, AI-SDLC, DevOps, and what 59 production deployments teach you. No spam. Unsubscribe anytime.

    Related field notes