Ground Truth

Ground Truth

Home
Notes
Archive
About

The Cost of Being Busy

Most companies are paying for the work twice and changing nothing underneath

Craig Hepburn's avatar
Craig Hepburn
May 28, 2026
Cross-posted by Ground Truth
"A long read but explains something about what I am seeing... AI in service of enabling us to focus on the more important work requires offloading the less important work."
- Greg Petroff

Last week, Microsoft began cancelling internal Claude Code licences for the thousands of employees it had given access to six months earlier.

This is the company that owns up to five percent of Anthropic, has a thirty billion dollar Azure compute commitment from them, sells their product to its own customers through Microsoft 365 Copilot, and whose CEO said last year that twenty to thirty percent of code in some Microsoft repositories is now written by AI. The reporting from The Verge and Fortune is consistent. Adoption caught on faster than expected, the token bill outran the budget model, and the workforce is being redirected to GitHub Copilot CLI.

The story being told about this is that Microsoft is being inconsistent. I do not think that is right. Microsoft is the first large company to publicly run into a problem that every other large enterprise will hit over the next eighteen months. The problem is the difference between using AI as a tool and building agentic systems, and the people running the budget have not yet realised they are on the wrong side of it.

The leaderboards

The visible response across the rest of the market is the leaderboard.

Meta built one in early April and called it Claudeonomics. Top two hundred and fifty token users, ranked, with titles like Token Legend and Cache Wizard. It was taken down forty-eight hours after it leaked. Disney has an AI Adoption Dashboard with one engineer reported using Claude fifty-one thousand times in a single day. Visa consumed roughly 1.9 trillion tokens in March, double its February figure. JPMorgan is running internal dashboards tracking Copilot and Anthropic usage per user.

The pitch behind all of this is that token consumption equals adoption equals productivity equals value. Jensen Huang said it cleanly in March. If a five hundred thousand dollar engineer did not consume at least two hundred and fifty thousand dollars worth of tokens, he would be deeply alarmed. Cristina Cordova, COO of Linear, put the other view just as cleanly on X. Ranking engineers by token spend is like ranking a marketing team by who spent the most money. Don’t mistake a high burn rate for a high success rate. Two senior people in the same month taking completely opposite positions on what the tokens are actually measuring.

The leaderboard is what happens when an organisation has been told to adopt AI but does not know what to adopt it for. Tokens are countable in real time. Transformation is not. So tokens become the proxy, the proxy gets a leaderboard, and the leaderboard rewards the behaviour that produces the most of the proxy. The behaviour that creates value is invisible, because the work to make it visible has not been done.

What Uber actually admitted

Uber is the another version of the same problem.

In December last year, Uber gave its five thousand engineers access to Claude Code. By March, adoption had moved from thirty-two percent to eighty-four percent. Around eleven percent of pull requests at Uber are now opened by AI agents. Monthly per-engineer token spend ranged from a hundred and fifty dollars at the low end to two thousand at the high end. The CTO, Praveen Neppalli Naga, said he personally burned through twelve hundred dollars in a two-hour demo session.

Then in April, Naga told The Information something no other large company CTO had said publicly. “I’m back to the drawing board because the budget I thought I would need is blown away already.” Uber’s R&D rose nine percent year-on-year to 3.4 billion dollars, with AI cited as a primary driver.

Uber did not waste this money. The productivity is real. The problem is that the budget model was wrong, and the budget model was wrong because the company was running an AI-tools strategy while expecting agentic-systems returns. Those are not the same thing. They look identical on the way in. They are completely different by the time the bill arrives.

Tools versus systems

The distinction between an AI tool and an agentic system matters, because most of the confusion in the market comes from treating them as the same thing.

An AI tool sits on a person’s desk. Claude in a browser tab. Copilot in an IDE. ChatGPT in a paid subscription. The person opens the tool when they need it, asks it something, copies the output, pastes it where they need it, and goes back to their job. The tool speeds up specific tasks inside that job. The job has not changed. The workflow has not changed. The tool is used in the moment and put down when the user logs off.

An agentic system is infrastructure. It runs continuously, sits inside the workflow rather than on top of it, and replaces categories of work rather than speeding up individual tasks. A triage agent reads the inbound queue, classifies each item, generates responses for the routine cases, escalates the ambiguous ones, and writes the outcome to the CRM. It does this without a human opening a browser tab. A reporting agent pulls performance data across systems on a schedule, writes the first draft of the commentary, flags anomalies, and surfaces what needs a human decision. A research agent runs overnight against a brief, produces a structured output by morning, and waits for direction on what to do next.

The cost model is different. Tools are billed per seat. Systems are billed per token, per operation, per outcome. Tools scale with headcount. Systems scale with workload. A tool’s value is measured by how much faster the user feels. A system’s value is measured by what the organisation no longer has to do.

The organisational implication is different. Tools are deployed by procurement and used by employees. Systems are designed, built, evaluated, and maintained by a team that does not yet exist in most organisations. Tools require a licence agreement. Systems require workflow redesign, evaluation frameworks, model routing, agent architecture, and a continuous cycle of experimentation.

Most enterprise AI today is the tools version. Microsoft gave its workforce Claude Code as a tool. Disney’s dashboard tracks how often each employee opens the tool. Meta’s leaderboard ranks the people who use the tool most. The tokens are being consumed at the desk, by individuals, doing their existing jobs slightly faster. The cost arrives without the structural shift, because the structural shift requires the systems version of the conversation, which most companies have not started.

The maths of the tools strategy

The maths of the tools version is worth working through, because it is the bit most people are not doing.

Take an engineer in London earning a hundred and twenty thousand pounds a year. Add roughly thirty percent of on-costs for tax, pension, and benefits. The total wage cost is around a hundred and sixty thousand. The company gives the engineer a Claude Code licence. The engineer spends, at the middle of Uber’s reported range, around a thousand dollars a month in tokens. That is roughly ten thousand pounds a year, on top of the wage.

The company now pays a hundred and seventy thousand for the same engineer, doing the same job, in the same role. The engineer reports feeling twenty to thirty percent faster on their own desk. The company sees almost none of that at the P&L line, because the engineer’s reclaimed three or four hours a week get absorbed into more code, more meetings, more pull requests that need more review.

This is the double-spend problem. The company is paying the person to do the work, and paying the tokens to do the same work, and the work itself has not changed.

Multiply that across five thousand engineers. That is Uber. Multiply it across a hundred thousand employees. That is what Microsoft is sitting on. The token line is now a structural item on the P&L. The wage line has not moved. The two are stacked on top of each other, and the productivity that should be paying for both is showing up in feelings on the desk and not in revenue per head at the company level.

Two companies, same market

To make this concrete, picture two companies of the same size in the same market.

Company A is a fifty-person professional services firm running the tools strategy. They give every employee a Claude or Copilot licence at around two hundred pounds a month per seat. They tell people to use AI in their work. The licence bill is roughly a hundred and twenty thousand a year. The token consumption on top of that, for the dozen heavier users, adds another sixty to eighty thousand. Total year-one AI cost, somewhere around two hundred thousand pounds. The fifty-person wage bill is unchanged at, say, six million pounds. The output of the firm is up at the margin. Some client work is delivered slightly faster. But nobody is doing materially different work, and the bill of two hundred thousand has produced no structural shift in how the firm operates.

Company B is also fifty people in the same market. They have spent the past six months mapping the work in two functions, marketing and operations, into two categories. The bounded, repeatable work has been built into an agent stack with proper routing, evaluation, and stop conditions. The build cost was around a hundred and fifty thousand pounds, paid to a specialist team over four months. Ongoing operational cost is around forty thousand a year in tokens and infrastructure, plus another forty thousand to keep the system maintained and refined. Total year-one AI cost, around two hundred and thirty thousand pounds, slightly higher than Company A.

Here is where the numbers diverge. Company B’s marketing function, which had three people in it, has been redesigned. One of the three has moved into a client-facing growth role the firm has been wanting to fill for two years. The other two now supervise the agent stack and do the strategic work the function never had time for. Operations has gone the same way. Two roles that would have been hired into next year are no longer needed; the hires that are happening are senior client roles instead. The wage bill is flat. The output of the firm has roughly doubled in the two functions that were redesigned.

By year three, Company A has spent six hundred thousand on AI and produced no structural change. Company B has spent roughly six hundred thousand on AI, plus the build cost, and has redirected four roles into growth work that has measurably increased the firm’s revenue capacity. Independent research from Decipher and Riseup puts payback for well-scoped agentic systems at six to twelve months. Three-year total cost of ownership for an agent stack typically runs at one and a half to two times the build cost. The economics work, but only if the systems strategy is actually executed. Company A is paying just as much and getting none of it.

This is the gap that compounds over the next five years, and it exists at every scale. Two fifty-person firms. Two five-thousand-person firms. Two hundred-thousand-person firms. The numbers change. The shape does not.

Where the work actually divides

To see what an agentic system looks like in practice, it helps to think about how the work inside any function actually breaks down.

The work in any knowledge function divides into two categories. The first is bounded and repeatable. It has clear inputs, clear outputs, and rules that can be specified well enough to be evaluated. Inbound triage. CRM updates. Report generation. Data wrangling. Document formatting. Compliance logging. Calendar coordination. First-draft creation against a brief. Status summaries.

The second category is unbounded and contextual. Strategy. Positioning. Customer relationships. Judgment calls that require reading the room. Creative direction. The decision about which campaign to bet on, which deal to walk from, which product to kill, which market to enter. The work that has no specifiable correct answer because it depends on a hundred soft signals that change weekly.

In most companies today, one person does both categories. A marketer triages their own inbound, updates their own CRM, writes their own reports, drafts their own copy, and also tries to set the strategy and run the customer relationships. An engineer writes the boilerplate, runs the build, fixes the small bugs, formats the documentation, and also tries to design the architecture and make the difficult technical calls. The same person, the same week, the same job, doing both categories of work and serving neither well.

Giving that person a Claude licence speeds up specific bits of the first category. They still do the rest of the first category by hand. They still do all of the second. The token bill arrives. The wage line does not move. The output improves at the margin and the company carries on roughly as it always has.

In a company that has built systems, the shape is different. The first category sits inside an agent stack. Inbound flows through a triage agent that handles routine cases and escalates the rest. A reporting agent runs the weekly numbers and writes the first draft of the commentary. A monitoring agent watches the things that need watching. A creative agent generates first drafts that the human then directs and shapes. These systems run continuously, in the background, against measurable evaluation criteria, with cost controls and stop conditions built in. The human is supervising the line, not running it.

The time that returns to the human is then deployed against the second category. The marketer spends more time on positioning and the campaign decisions and the customer conversations that move the needle. The engineer spends more time on the architectural calls and the difficult technical decisions and the work that compounds. The agents handle the operational drag. The people handle the work that creates value.

What you end up with is the agent stack absorbing the first category of work, while the wage line does not move because the people are still there. The people are now doing the work the business has been understaffing for years. The output of the function multiplies. The token bill is real but it is bounded by the operational volume, not by how many prompts the user feels like typing.

A few large companies have started disclosing what this looks like in practice. Klarna built an AI customer service agent that, by their own published numbers in early 2024, did the work of seven hundred full-time agents, dropped average resolution time from eleven minutes to two, and contributed to a firm-wide hiring freeze outside engineering. They have since walked some of this back, hiring humans again for the higher-touch cases the agent could not handle, which is the most useful data point in the story. The redesign was real and measurable. It was also harder to sustain than the initial announcement suggested. Morgan Stanley reports that AI has saved its engineers two hundred and eighty thousand hours this year on legacy code conversion. Goldman Sachs has said internally that AI can now produce ninety-five percent of an IPO prospectus, work that previously took six people two weeks, in minutes. The direction across all three is the same. The transformation works when the work is properly redesigned. It does not work when AI is bolted onto unchanged roles.

What the operational drag actually is

The real opportunity sits in the operational drag, and this is the bit most of the industry keeps getting wrong.

Most of the public conversation about AI is about the high-end stuff. Can the model reason? Can it replace strategists? Can it write a novel? Those are the wrong questions, and they are why a lot of people are confused about what to actually do with this. The opportunity is in the operational work humans have been doing because nobody else could, not because it actually required them.

Information moving from one system to another. Documents formatted for different audiences. Data cleaned and joined. Meetings summarised. Inbound classified. Reports compiled. Compliance logged. Receipts categorised. Drafts produced. Approvals routed.

Asana’s Anatomy of Work Index puts knowledge workers at sixty percent of time on this kind of work. A separate HP study published last year put it at fifty-one percent of the average office worker’s day. The numbers vary by methodology but the direction is consistent. More than half of most knowledge workers’ weeks goes to operational work that is not the work they were hired for. It is the work that has accumulated around the work they were hired for, because organisations have always needed someone to do it and that someone has always been the human in the seat.

The agent stack is the first technology that can absorb this layer of work without losing the context humans bring to it. The opportunity is in the wrangling, not in the high-end thinking. Most of the industry has been pointing at the wrong thing.

Every business has a list of things it has been meaning to do that it has not had time for. The customer conversations that would have grown the account. The market analysis that would have surfaced the opportunity. The product decisions that should have been made six months ago. The growth bets that have been deferred. The strategic thinking that has been crowded out. The agent stack returns the time. The redirected people do the work.

Replacement versus redirection

There is another place where the industry narrative is making the problem worse, and it sits inside the workforce itself.

The dominant framing right now is replacement. AI will replace jobs. AI will replace knowledge workers. AI will replace whole functions. Pick a CEO with a microphone and you can find a quote. The framing is wrong on the substance, and it produces exactly the wrong behaviour in the workforce.

If you tell an employee their job is being replaced, they will, rationally, play the game to protect their existing role. They will demonstrate use of the tool. They will run more prompts than the next person. They will produce more output, even if the output is slop, because output is the evidence they will be asked for. They will not show you the parts of their job that could be automated, because doing so accelerates their own redundancy. They become an obstacle to the transformation, not a participant in it.

If you tell the same employee that the operational drag is going to be absorbed by systems and that they are being redirected into the higher-value work the company has been short on for years, the incentives flip. The employee has a reason to identify what in their role can be automated. They become an active partner in the redesign. They want to move up the stack because the move up is rewarded. The transformation accelerates because the people closest to the work are pulling for it instead of resisting it.

The companies that frame this as redirection rather than replacement will move faster. The companies that frame it as replacement will spend the next three years fighting their own workforce. Both companies will eventually arrive at agent stacks. One will arrive with its best people still in the building. The other will arrive after losing them and replacing them at higher cost.

Out of every operational decision a leadership team is making about AI right now, this one matters most, and a lot of them are making it the wrong way because they are listening to the industry narrative instead of thinking about how their own people will actually respond to it.

You cannot cut your way to growth

There is one more thing worth saying directly. Pure cost-cutting plays with AI do not work.

The companies running a tools strategy as an efficiency programme are optimising yesterday’s work. The savings are real but they are bounded by the size of the legacy cost base, and they shrink the company while the competition is being rebuilt around something different. The companies running a systems strategy are using the agent stack to absorb the operational drag and to do work that was not being done at all. The growth comes from the second of those, not the first.

You cannot reach a new cost structure by trimming the old one. You have to build the new one, and build it before the old one stops being affordable. The companies treating AI as a cost programme will produce slightly leaner versions of what they already are. The companies treating it as a growth programme will be operating in a different category by the time the cost-cutters notice.

None of this is easy

None of this is easy work. Anyone telling you it is, is either selling you something or has not actually done it.

Building an agent stack that absorbs real operational work requires expertise across a stack of decisions that all have to be made well together. Which model handles which task, when a reasoning model is justified and when a small fast model is better, how to bound an agent so it does not loop forever, how to write evaluations that catch failure modes rather than only measuring throughput, how to route work between models to optimise cost without losing quality, how to build the infrastructure that determines whether the maths works at all. A wrong model routing decision can move costs by thirty to fifty percent. A poorly designed evaluation framework can ship a system that fails silently on edge cases nobody thought to test. An unbounded agent can burn a month’s budget in a day. The expertise is real and it is rare.

Which is also why smaller and newer companies have a higher chance of getting this right than incumbents. Not because they are smarter. Because they do not have the legacy operational structure that needs to be disassembled before the new one can be built. They do not have ten thousand employees whose roles will need to be redesigned. They do not have a CFO trying to defend the headcount line against an unproven agent stack. They do not have a procurement process that knows how to buy seats but does not know how to buy systems. They do not have a culture that has spent twenty years rewarding the people who do the operational drag and will now have to tell those people their work is moving. They start from the redesigned shape because that is the only shape they have ever had.

What the Industrial Revolution actually did

The Industrial Revolution did not give individual weavers more tools. That part of the history gets misremembered constantly, and it is the bit that matters now.

Before industrialisation, cloth was made by one person who did every step. They spun the yarn, wove the cloth, dyed it, finished it. They were skilled across the whole process. The industrial leap did not arrive because someone built a better hand loom. It arrived because someone built a factory: a system that absorbed the bounded steps of cloth production into specialised stations, and redirected the skilled people into supervising the system, designing what got made, and managing the quality. The skilled all-rounder did not disappear. Their role changed. They moved up the stack.

The factories that thrived in the nineteenth century were the ones that built this system early. The owners who refused, who insisted that the right answer was to buy slightly better hand looms and put them on each worker’s bench, were out of business within twenty years. They could not compete with the factory on cost, volume, or consistency. They were no longer in the same business.

The agent stack is the factory. The redirected employee is the weaver who became a supervisor. Giving every employee a Claude licence is the hand loom. The companies that have rebuilt the function around an agent stack, with the people moved up the stack into the work that compounds, are operating on a different cost structure from the companies that have not. The companies that have not are paying twice and producing roughly what they always produced, at slightly lower marginal cost, until a competitor that has redesigned puts them out of business.

Where companies sit right now

Most companies will not make this shift cleanly. They will keep running the tools strategy for two to four more years. They will spend more on tokens, see productivity gains at the individual level, fail to convert them into margin at the organisational level, and eventually be forced into restructuring under cost pressure. They will do it badly because they will do it under duress, and they will frame it as replacement because by then the cost pressure will leave them no room to frame it any other way.

A smaller number of companies are starting now. They are taking specific functions, mapping the work into the two categories, building the systems for the first category, and redirecting the people whose work has moved into growth, strategy, and the work the business has been short of. Block is the most public example, and I have written about Jack Dorsey’s restructuring in detail before. Shopify is doing something different but pointed in the same direction. Tobi Lütke’s April 2025 memo told every Shopify team they had to demonstrate why a job could not be done by AI before they could ask for new headcount. AI use is now part of performance reviews. Neither company is laying people off as the headline strategy. Block redirected; Shopify is gating growth in hiring through proof of AI exhaustion. The shape is the same in both cases: the headcount they were going to add is being put into the work that compounds, not into work the agent stack can absorb.

A third group is being built from scratch on this cost structure. Cursor reached one hundred million dollars in annualised revenue in roughly a year. By February this year the figure was two billion, with a team of around fifty people. That makes it the fastest company in B2B software history to a billion in ARR, and by current estimates the highest revenue-per-employee ratio of any software company that has ever existed. Lovable hit a hundred million in eight months with forty-five people. Midjourney is above two hundred million with around forty. Anthropic itself runs with about five thousand employees against revenue that puts it somewhere between four and six million dollars per head.

These companies are not better at directing tokens. They never built the additive cost structure in the first place. They built systems from day one, redirected every employee into the work that compounds, and let the agent stack absorb the operational drag. The wage line that does not exist is the largest hidden saving in software economics. A company that hires linearly with revenue is competing, in the same market, with a company that does not. The incumbent pays the salary of the person who delivers each additional dollar of revenue. The AI-native competitor pays the marginal cost of tokens. The gap between those two numbers is the margin difference that will play out over the next five years.

The question for any operator reading this is which of those three companies is closer to the one they currently run.

Not in five years. Today. The licences they have handed out, the leaderboards they may have built, the tools sitting on every employee’s desk with no systems sitting behind the workflow, the people who are still doing the operational drag by hand while their token bills arrive, the roles they have not yet questioned, the procurement process that knows how to buy seats but has never bought an agent stack, and the workforce that has been told their jobs are being replaced and is, rationally, fighting back.

The tokens, the spending, the leaderboards are not the problem. The organisation around them is, and the organisation is the thing almost nobody wants to redesign. Microsoft cancelling Claude Code is what that looks like at the largest scale: a company hitting the point where the tools strategy is no longer affordable and the systems strategy has not yet been built. They pulled the licences because the unit-cost maths stopped working. It is also a quiet admission that the redesign has not been done.

A leaderboard tells you who used the most tokens. It tells you nothing about whether those tokens absorbed the work people should never have been doing, or whether those people are now doing the work the business has been short of for years. Almost nobody is measuring that second thing yet, because it is harder, slower, and it does not fit on a dashboard.

That measurement is the whole game, and it is the part almost no one has started.


If this landed, subscribe. The next anchor piece looks at where value accumulates when the cost structure inverts: which application layer businesses capture the margin, and which ones pay for it without realising they are.

Craig Hepburn is an AI strategist and Perplexity Fellow. Twenty years building at the frontier of digital, from Microsoft and Nokia to Art Basel and UEFA. Now building at the frontier of agentic intelligence.

No posts

© 2026 Craig Hepburn · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture