Do You Cap Your Engineers' Token Budget, or Not?
A flat monthly cap feels like control. It sails over the people who barely use the tools and lands entirely on the ones inventing how everyone works next year. The real choice was never cap or no cap.
At some point, if you are running engineering at any scale, someone shows you a number. The monthly spend on AI coding agents. It has a comma in it that was not there last quarter, and finance wants to know what the plan is.
The plan, almost always, is a cap. A per-engineer monthly token budget. It is the obvious move, and I understand exactly why people reach for it. You do not want the tap wide open and then discover, three months in, that you spent tens of millions nobody signed off on. A cap feels like control.
I have spent a good part of this year inside this question. Not as a thought experiment, as an actual decision I have to make and defend. And the honest answer is that both sides of it are right, which is the part most of the debate misses.
The case for the cap is real
Start with the strongest version of it, because it deserves better than a strawman.
The sharpest reply I have seen to the loudest “never cap your engineers” posts is always some version of the same question: how many engineers do you have? At a few thousand, an uncapped tap is not brave, it is ten to twenty million a month that nobody budgeted. Uncapped is a luxury of small teams. At scale it is just unaffordable, and you cannot run a business where a single line item is unbounded and nobody is watching it.
So the anti-cap purists are wrong when they pretend cost does not matter. It does. Token spend with nobody watching it is a genuine problem, and pretending otherwise is how you end up explaining a surprise to your CFO.
That is the ceiling argument, and it holds.
But a cap is a blunt instrument, and it hits the wrong people
Here is where it gets interesting.
If you watch how engineers actually use these tools, they sort themselves into roughly three groups.
The first group fights the tool. They blame the model, they distrust the output, they finish most of their work by hand anyway. Their spend is near zero. A cap does not touch them.
The second group is careful. They approve every step, they watch each action, they use the agent like a slightly smarter autocomplete. Their spend is small change. A cap does not touch them either.
The third group is small. These are the people running agents overnight, rewriting their own tooling on a Sunday, figuring out how everyone in the building is going to work eighteen months from now. Their spend is high, because they are doing the most.
Now look at what a flat monthly ceiling actually does. It sails clean over the first two groups. It lands entirely on the third. You have built a control that leaves the low-value users untouched and squeezes precisely the people you are betting the future on.
Worse, it hands them a clear signal. Capped, your most productive people do one of two things. They leave for somewhere that pays for experiments, or they quietly rejoin the comfortable middle and stop pushing. Either way you lose the exact people you brought the tools in for.
You might object that “high spend equals high value” is just a flattering story the power users tell about themselves. The data says otherwise. a16z recently pulled together spend figures from Ramp showing the top 10% of firms spend roughly fifty times more per employee on AI than the median. Not fifty percent, fifty times. And a BCG analysis of 107 public companies found the top two quintiles of token usage posted meaningfully faster revenue growth than the rest of the field. The distribution is not an accident. The heavy spenders are, disproportionately, the ones extracting value, which is the whole reason cutting them off first is such a strange thing to do.
The part nobody costs: the cap makes savers, not explorers
There is a second effect, quieter and slower, and I think it is the more dangerous one.
Give someone a productivity mandate and then hand them a tight budget, and you have set up a contradiction they will resolve in the most rational way available. They will spend their allocation to deliver, never to discover. Exploration is the first thing cut when the tap is tight, because exploration is optional, unmeasured, and has no ticket number attached to it.
This does not just hit the power users. It reshapes everyone. A flat cap quietly turns your whole engineering org into savers. They protect their allocation. They use it on what was asked. They stop poking at the thing that might, in three weekends, change how the team works.
So the cap has two costs, not one. It drives out the small group inventing the future, and it manufactures caution in the middle that was your best chance of a second such group emerging. You do not just lose your explorers. You make the middle even more middle.
What the people who have done this at scale actually do
I recently read a piece from Databricks that put numbers on something I had only been able to argue from conviction. They synthesized how the earliest large-scale adopters, cross-checked with Stripe, Coinbase, Uber and Ramp, actually manage this spend once agentic coding is deployed broadly.
The finding that matters here: hard budgets, where usage is cut off entirely at a threshold, are used only as a last resort in every company they spoke with. Two reasons. One, cutting a developer off mid-task is simply debilitating, nobody wins. Two, and this is the sharp one, some of the highest spenders are the ones who have achieved the biggest efficiency gains. The blunt cap punishes the people who cracked it. The exact conclusion, from companies operating at a scale most of us are not.
What they converged on instead is not “no limits.” It is a ladder of friction that rises with spend rather than a wall that stops it.
First, visibility. Near-instant feedback to each person on their own spend, across all their tools, often with a nudge toward a cheaper model for the task. People cannot self-regulate what they cannot see, and it turns out most will, once they can.
Second, spend gates that escalate. The lightest is self-clearing: a warning you acknowledge and move past, which catches accidental runaway spend without blocking anyone. Think of a policy that pops up mid-task, “this task has already cost you ten dollars, keep going?”, and lets you carry on with one click. It does nothing to the person doing real work, but it stops the runaway loop nobody meant to launch. Heavier gates ask for an actual approval up the chain.
Third, downshifting. When you hit a gate, you get moved to a cheaper model, not suspended. The cheapest models are dramatically less expensive than the frontier ones, so work continues at a fraction of the cost instead of stopping.
Fourth, suspension, kept as the limit case. And even then treated as the start of a conversation about how to work efficiently, not as a punishment.
The principle underneath is one line: preserve access, escalate friction. The system informs and nudges long before it ever blocks, and full blocking is a deliberate exception, not the default setting.
That is not a softer cap. It is a different instrument entirely. The cap asks “have you spent too much.” This asks “are you spending well, and does everyone around you know what good looks like.”
The reframe
So the question in the title is a bit of a trap, and I set it deliberately.
“Cap or no cap” is a false binary. The uncapped purists are right that a flat ceiling teaches your best people the wrong lesson. The cap defenders are right that unbounded spend with nobody watching is not a strategy. Both of those are true at the same time, and the useful answer lives between them.
The real choice is not flat cap versus no cap. It is flat cap versus visibility plus differentiated budgets. Meter everything through one point so you can actually see the spend. Then let friction rise with cost rather than dropping a wall at a number. And stop pretending every engineer should have the same budget, because they are not doing the same work. The middle who will never notice the ceiling and the handful who would hit it in a week are not the same problem, and one number cannot govern both.
One more thing worth knowing before you set any number at all, because it changes where you should even be looking. That same Databricks piece points out that when someone types “investigate and fix this bug,” the instruction is a handful of tokens, and by the time the model actually runs, that instruction is a negligible fraction of what it is processing. The cost is dominated by everything the agent pulled in on its own: the files it read, the tools it called, the context it assembled. They report that simply tuning the harness and cache settings cut token spend by nearly half, with no drop in quality, touching neither the model nor the prompt.
Which means a lot of the spend you are about to cap is not your engineers being wasteful. It is the machinery around them being untuned. Before you ration the people, it is worth asking whether the cheaper win is in the plumbing.
I will keep the same honesty I try to keep on all of this: Databricks is selling the products it credits, and the savings numbers are self-reported and directional, not a benchmark. But the pattern is sound, it is corroborated across several companies, and it matches what I have seen firsthand.
So, do you cap your engineers’ token budget?
If the question is whether spend should be unwatched, no, and anyone who tells you otherwise has not seen the number with the comma in it yet. But if the instrument is a flat monthly ceiling applied equally to everyone, then you are aiming a blunt tool at the wrong people, and quietly teaching your best engineers that curiosity is expensive. There are better instruments. They are just less satisfying to put in a spreadsheet.