All posts
Sep 12, 20268 min read

You Can't Just Caveman Your Way Into Savings

By Quotient Labs

There's a viral, and appealingly simple, way to lower the cost of a coding agent:

Talk like caveman. Use less token same quality. Save money.

At the outset, it makes sense. Articles, prepositions and transition words don't contribute much to a model semantically. Why pay for a whole sentence when 3 nouns and a semicolon will do?

The viral Caveman response skill is a clever version of this: it makes an agent less chatty, and shorter responses are obviously cheaper than longer ones. But, as we will discuss in the rest of this post, a coding agent is not merely a chat window. Its bill is generated by an entire loop of searches, file reads, tool arguments, command output, edits, cache writes, and cache reads. The prose you see is only one part of the loop, and a surprisingly small part at that.

What's Actually in a Coding Session?

We did an analysis of the SALT-NLP/SWE-chat dataset, a corpus of real coding-agent sessions over public repositories. We used the Claude Code subset of the fixed April 29, 2026 snapshot (f66cca95b14caaa4177f7ed5eaa424608dadcffa): 4,852 sessions and 2,632,125 logged interactions. Unlike a chat benchmark, SWE-chat includes the tool calls and tool results between the human and assistant messages, so we can do a comprehensive analysis of what Claude Code data distribution actually looks like.

We deduplicated transcript events and excluded progress notifications, file snapshots, queue operations, and other client metadata that is not part of the model-visible conversation. We then counted the characters in each remaining content block.

Content typeCharactersShare
Tool results693,448,27567.41%
Tool calls and arguments211,693,64220.58%
User and system input94,642,9679.20%
Assistant prose28,549,8572.78%
Assistant thinking301,3780.03%
Compaction summaries42,607<0.01%

Tool I/O accounts for 87.99% of the model-visible transcript. Assistant prose accounts for 2.78%. That is, 97 of every 100 visible context characters are not narrative assistant prose at all.

Moreover, the raw size of the transcript actually still understates the effect. Tool results don't enter context once and disappear: they are carried into later model calls and repeatedly billed as cache re-reads until the context is compacted or cleared (and can become even more expensive after cache misses or expiries). We computed an API-call-weighted estimate, giving earlier blocks more weight according to how many subsequent calls carry them as context and billing baggage. Under such an estimate, tool I/O rises to 88.85% of context, while assistant prose is only 3.08%.

The Upper Bound of Prose-Only Compression

The transcript distribution allows us to estimate actual costs as well. For the 4,806 Claude Code sessions with nonzero provider token counters, we repriced base input (1x), five-minute cache writes (1.25x), cache reads (0.1x), and output (5x) for the Claude models represented in the corpus. We held the task trajectory fixed (identical tools, turns, and files), and reduced assistant prose output by 65%, which was the original headline associated with Caveman-style output. Everything else remained unchanged. We accounted for compounding these savings over multiple cache reads over the session trajectory as well.

But take a look at the numbers:

CounterfactualEstimated direct bill reduction
Assistant prose reduced by 65%2.47%
Every assistant-prose token deleted3.80%

This is the headline result: even if we reduce the length of the average prose output by 2/3, and compound the savings, assistant prose makes up so small a share of the cost distribution that the actual savings are marginal. And that's even before you take into account the fact that the reduced context may have unintended consequences on the agent trajectory (it could lead to earlier stopping, but conversely may also lead to more model turns if the agent feels it's missing context or substance).

We Aren't the Only Ones To Find This

JetBrains forced Caveman on across roughly 240 billed SkillsBench trials. The measured output-token reduction was 8.5%, from 592,000 to 542,000 tokens — which is useful, but nowhere near 65%. (Not to mention that even this number is just with output tokens as a denominator, most of your actual bill is input tokens).

And they flagged the same fact that agent output is dominated by code, diffs, tool invocations, and exact strings that a terseness prompt leaves alone.

Nathan Cavaglione ran 92 paired coding tasks on SWE-bench Pro and measured a 14% output-token reduction and a 12% total-cost reduction. His transcript analysis found that only 13.3% of baseline output was prose. Moreover, even SWE-bench is still far less tool-heavy than real workflows. As Verdent's technical report puts it: "This suggests a potential bias in current public benchmarks, as the minimal toolset required to succeed on these benchmarks stands in stark contrast to the complexity of real-world software engineering."

Caveman also used more output tokens on 34% of the paired tasks, illustrating how a style instruction can also change the path the agent takes rather than merely shortening an otherwise identical response. It's certainly not guaranteed that the resultant trajectory will be shorter, either.

Savings Are Fundamentally a System Problem

None of this means that prose compression is irrelevant (We do it in Fermat too, though with a compression model rather than a system prompt). But optimizing one small region can't be a substitute for optimizing the rest of the agent loop.

Fermat has four separate attack arms:

  1. Tool-result compression. Large reads, search results, and shell output are reduced before irrelevant bytes can accumulate and compound through later turns.
  2. More efficient tools. Fermat's Search and Edit interfaces batch related operations and return focused evidence, reducing both the number of calls and the amount of I/O produced by each call.
  3. Prose compression. Eligible narrative content is compressed with a dedicated model, rather than depending solely on an instruction that changes how Claude communicates and may alter its trajectory. This model runs after Claude produces its outputs, but before those outputs make it into the prompt cache, so you also don't have to suffer from the readability issues that come with overly terse outputs.
  4. Idle context optimization. After the prompt-cache TTL has elapsed (by default every 5 minutes!), Fermat compresses older context while preserving the freshest working state, avoiding an unnecessary cache invalidation while reducing the next cache write and subsequent reads.

Things like Caveman are fun, and brevity is often good. But if you want to materially reduce a coding-agent bill, you have to optimize over the distribution of tokens an agent actually spends.

Or you can install Fermat, and let us handle the optimization for you :)