Quotient Labs
The Margin
Notes on context compression, benchmarking, and building Fermat.
You Can't Just Caveman Your Way Into Savings
A terse-output prompt attacks the visible assistant output prose. But the overwhelming majority of context and cost lives elsewhere: tool calls, results, and the same I/O re-read (and re-billed) at every subsequent turn.
Sep 12, 20268 min readBenchmarking Context Optimization
How Fermat's idle compression works — and how we verify it preserves the working state an agent actually needs.
Aug 27, 20268 min readCompressing Prose
Fermat's prose compression model trims low-value tokens from text-heavy turns at write time. Here, we detail how we benchmark it.
Aug 14, 20265 min read