A coding agent spends tokens on more than reasoning. It spends them moving information: listing files, reading modules, parsing test logs, comparing manifests, and repeating context across turns. That work is necessary, but the raw output often is not. A thousand-line test run may contain one failing assertion. Five files may be opened to find one relevant function.
Token efficiency starts by treating the context window as a finite engineering resource. The objective is not to minimize usage at any cost. It is to keep high-signal evidence near the model that makes the difficult decisions, while filtering routine noise and moving bounded work to appropriately sized agents.
Spend tokens where they improve the decision. Compress the rest.
Token volume is a trace, not a performance metric
Nate's token burn dashboard makes a useful distinction: token usage becomes meaningful only when it is connected to outcomes. A large number can mean that agents handled more real work. It can also mean that the workflow repeatedly loaded irrelevant context. The number alone cannot tell you which one happened.
This is especially important across tools. Codex, Claude, and ChatGPT do not necessarily account for tokens in directly comparable ways. A dashboard should therefore behave like an observability surface, not a leaderboard. Track what task ran, which model handled it, whether the output was accepted, how much rework followed, and what engineering result changed.
The same principle applies to teams. Ranking engineers by token burn rewards volume without proving value. A better review asks whether a successful one-off run became a repeatable workflow. A short weekly review is enough to find those patterns and decide what should be automated next.
The context window is an architectural concern
Every file, command result, requirement, and previous decision competes for attention inside the context window. Long before the hard limit is reached, the signal-to-noise ratio can deteriorate. The model must repeatedly recover the relevant facts from a larger transcript, which adds latency and increases the chance that an earlier constraint is missed.
This makes context engineering part of system design. Prompts matter, but so do data selection, tool output, delegation boundaries, and return formats. In practice, I use three controls together:
- filter noisy command output before it reaches the model;
- give agents a structural map before they start reading files;
- route bounded work to the least expensive model that can perform it reliably.
First control: filter tool output with RTK
RTK is a command-line proxy that sits between an agent and common development commands. It keeps the part the model needs and collapses routine output. A failing test remains visible, while passing-test scaffolding becomes a count. A large git status becomes a compact state summary. Repeated container logs can be deduplicated.
# Raw commands still exist, but the agent can prefer compact variants
rtk git status
rtk pytest
rtk tsc
rtk docker logs api
rtk aws sts get-caller-identity
For Codex, rtk init -g --codex adds instruction-based integration so the agent prefers RTK for supported commands. This is different from hook-based integrations that rewrite shell calls automatically. The operational goal is the same: reduce low-value text before it enters the conversation history.
AutoScout24 reports a 60% to 90% reduction in development-loop output from this approach. RTK itself is careful about what that means. The percentage describes compressed shell output, not a 90% reduction in the total AI bill. Prompts, system instructions, conversation history, model output, and reasoning still consume tokens. RTK also estimates token counts from bytes, so its absolute token totals are approximate.
The useful commands are rtk gain, which shows estimated cumulative savings, and rtk discover, which finds noisy commands that still bypass the proxy. That turns output compression into something observable instead of a rule that nobody knows is working.
Second control: give the agent a map
Blind repository exploration is expensive. An agent searches for auth, opens several files, follows imports, reads tests, and slowly reconstructs a dependency graph that the codebase already contains.
Think of the repository as a large city. Searching with rg and opening every match is like exploring it street by street. A structural index is the map: it already knows which functions call each other, where a symbol is defined, which tests cover it, and which modules may be affected by a change.
Tree-sitter and language-server-based tools build this map from the code's syntax and references. MCP gives the agent a standard way to ask questions such as "who calls this function?" or "what is the impact radius?" Semantic search helps when the agent knows the intent but not the symbol name. Repo packers provide a cleaned snapshot when the whole repository is genuinely needed.
In my Codex setup, code-review-graph provides this layer. It uses Tree-sitter to build a persistent graph of functions, classes, imports, calls, tests, and dependencies, then exposes targeted graph queries through MCP. Incremental updates reparse changed files instead of rebuilding the entire repository. I keep cloud embeddings disabled and expose only the analysis tools needed for discovery, review context, and impact analysis.
The map does not replace source inspection. It narrows the search. The efficient order is to query the graph first, use rg to confirm the relevant locations, and read the exact code before editing it. Structured navigation has setup and maintenance cost, so it pays off most in large or tightly coupled repositories. Small codebases may not need it.
Third control: route work to the right model
Model routing classifies a subtask by ambiguity, volume, and the cost of a mistake. It is not a contest between a "strong" model and a "weak" one. In Codex, I use three levels:
| Owner | Workload | Why |
|---|---|---|
| gpt-5.6-terra low · read-only | Read-heavy scans, large files, cross-repository correlation, and document processing. | Explores broad context and returns distilled evidence. |
| gpt-5.6-luna low | Searches, counts, manifests, boilerplate, localized changes, and predictable validation. | Executes narrow, clear, repetitive tasks quickly. |
| Primary agent | Architecture, complex debugging, security, conflict resolution, integration, and final conclusions. | Preserves end-to-end coherence and accountability. |
This prevents two expensive mistakes. The first is using the primary model for repetitive work. The second is allowing a cheaper worker to make decisions that require system-level judgment. Reasoning effort belongs in the routing decision too. A narrow scan rarely needs deep reasoning, even if it uses the same model family.
Fan-out, reduction, and synthesis
For broad analysis, the primary agent can use a small map-reduce pattern:
user request
│
├── broad_reader ──► topology + relevant files
├── mechanical_worker ──► counts + manifests + checks
└── primary agent ──► security + architecture + final synthesis
During fan-out, independent subtasks run in parallel. During reduction, each worker returns structured conclusions: files, line numbers, validation, and risks. The primary agent correlates those results and inspects the decisive sections.
The output contract matters. If a subagent returns long logs or complete files, the savings disappear. Its output must be deliberately small, precise, and auditable.
AGENTS.md as a control plane
The AGENTS.md file acts as a declarative policy layer. A global rule can define when to delegate, which profile to use, and which decisions must never leave the primary agent. Repository-specific files can add constraints closer to the code.
A condensed example:
## Default delegation and model routing
- Use broad_reader for repository exploration, large-file review,
and cross-file correlation.
- Use mechanical_worker for narrow, repeatable work.
- Keep architecture, security, complex debugging, and final
integration in the primary agent.
- Never assign overlapping files to concurrent editing agents.
This creates a readable, versioned policy independent of each task's prompt. Instructions do not replace technical isolation. Agent profiles complete the policy.
TOML profiles: capability and least privilege
Custom agents can define their model, reasoning effort, sandbox, and dedicated instructions. For broad exploration, a read-only profile reduces the chance that investigation produces accidental changes:
name = "broad_reader"
model = "gpt-5.6-terra"
model_reasoning_effort = "low"
sandbox_mode = "read-only"
developer_instructions = """
Correlate evidence and return concise findings with exact paths
and line numbers. Never edit files or make final decisions.
"""
The mechanical worker can receive workspace-write, provided that its scope is narrow, its files do not overlap with another worker’s, and the primary agent reviews the diff. This is a direct application of least privilege: each agent receives only the authority necessary to complete its part.
What should not be delegated
Model routing is not an excuse to fragment accountability. Some tasks depend on deep context, judgment, and the ability to evaluate indirect consequences:
- architectural decisions and changes to system boundaries;
- security analysis, IAM, and data exposure;
- debugging concurrency, race conditions, and intermittent failures;
- destructive or difficult-to-reverse migrations;
- resolving contradictory results from different agents;
- final review, risk-proportionate testing, and approval for delivery.
Subagents can collect evidence in these domains, but final interpretation should remain with the agent that owns the end-to-end view.
Delegation has a break-even point
Delegation adds orchestration latency, repeats some context, and consumes tokens in the worker. For a small file or a task that takes only a few minutes, the overhead may exceed the savings. Delegate when the job has enough volume, real parallelism, or a stable output contract to justify coordination.
Spotify reports mean savings of roughly 90% in its Claude-side bulk-read scenarios when Portal workers return summaries. AutoScout24 reports 60% to 90% less development-loop output with RTK. RTK advertises up to 90% compression for supported shell output. None of these claims means a universal 90% reduction in total cost. They measure different boundaries.
Measure efficiency against outcomes
A useful dashboard combines consumption, workflow, and delivery data. I would start with these measures:
- input and output tokens by tool, model, task type, and repository;
- RTK's estimated shell-output savings and the commands that still produce noise;
- cache-hit ratio, latency, and cost per completed or merged AI-assisted task;
- acceptance, rework, test failures, rollbacks, and review findings;
- the proportion of successful one-off runs converted into reusable workflows.
Cross-tool totals should remain separate unless their accounting methods are normalized. Even then, the dashboard needs labels that distinguish measured tokens from estimates. Cost per accepted outcome is more useful than raw burn, but quality checks must stay beside the cost number. A cheap result that creates rework is not efficient.
Accountability remains with the primary agent
The best implementation is not the one that creates the most agents or produces the lowest token chart. It keeps the primary context focused without losing traceability. The primary agent validates evidence, inspects consequential changes, runs appropriate tests, and rejects weak conclusions.
This is Platform Engineering applied to coding agents: give each workload a paved path, filter routine noise at the boundary, isolate capabilities, and measure the result. RTK reduces information movement. Structural indexes reduce blind exploration. Model routing assigns work by ambiguity and risk. The primary agent still owns the final decision.
References
- Spotify Engineering — Portal by Spotify cut my Claude Code token usage by 90%
- Nate's Newsletter — Build a Token Burn Dashboard to Track What Your AI Actually Does
- AutoScout24 TechBlog — 3 Simple Techniques to Reduce Token Consumption in Claude Code and Codex
- RTK — CLI proxy for compact development-tool output
- Code Review Graph — Local-first code intelligence graph for MCP and CLI
- OpenAI Docs — Subagents
- OpenAI Docs — Custom instructions with AGENTS.md