I came downstairs on a Saturday morning to find my two sons on the living room floor, surrounded by loads of Lego figures. Bags, boxes, loose pieces everywhere. They were photographing them, sorting them, looking up prices on their phones.
I assumed they were tidying. They were not.
They were liquidating assets.
The figures had appreciated. Some of them significantly. And my sons, aged fifteen and seventeen, had independently concluded that the most productive thing they could do with that stored value was convert it into something more useful to them than plastic.
Tokens.
Not to spend on games. Not on food or clothes or anything you would expect teenagers to want. Tokens to feed into Claude, into Codex, into the models they run through OpenRouter. Tokens to build with. To code with. To keep making the games and tools and applications they have been building in their bedrooms for the past year.
I stood there and watched them work.
And I thought: they already understand something that most of the boardrooms I walk into have not yet grasped.
From energy to intelligence
Every economy in history has been built on a primary input. An industrial economy ran on coal and oil. A digital economy ran on electricity and bandwidth. You could not participate without access to the underlying resource, and the nations and companies that controlled it set the terms.
The intelligence economy runs on compute. And compute is now being manufactured into a specific, tradeable, priceable unit: the token.
The chain runs like this.
A GPU running at capacity has a measurable token output rate. Huang’s core metric is tokens per watt per second. When Microsoft announced its hundred billion dollar infrastructure investment, it was not buying storage. It was buying token manufacturing capacity. The physical infrastructure of the intelligence economy is being built right now and its output is denominated in tokens. Compute generates tokens.
A single generative prompt consumes roughly thirty tokens. A reasoning task consumes hundreds of times more. An agentic workflow running autonomously across tools, data sources, and decision points consumes millions. The more tokens you direct at a problem, the more reasoning cycles the system completes, the higher the quality of the output. Token consumption is not usage in the way a software subscription is usage. It is the compute equivalent of thinking time. You are purchasing cognition. Tokens generate intelligence.
That intelligence has to connect to something real. A contract drafted. A codebase built. A market understood well enough to move faster than a competitor. A decision made with better information than the one made the week before. The organisations that can draw a direct line from their token spend to their output quality, and from output quality to business results, are running a fundamentally different operation from those treating AI as a productivity add-on. Intelligence generates output.
And output, at scale, compounds into something larger. Jensen Huang put it directly at the Morgan Stanley Technology conference: “I am certain compute equals revenues. I am certain also that compute equals GDP. Therefore, every country will have it because not one country in the future will say, ‘Guess what, we will not participate in GDP.’” The nations building sovereign compute infrastructure are not doing it for technological prestige. They understand that the production function of their economy is changing. Output generates revenue and GDP.
Every link in that chain is now measurable, priceable, and allocatable. That is what is new. Not the idea that intelligence creates value. The fact that you can now buy it by the unit, direct it at specific problems, and account for what it produced.
The raw material of the last economy was what a person knew and how long they worked. The raw material of this one is cognitive compute, and access to it is already diverging fast between the organisations that understand this and those that do not.
The factory, not the warehouse
Huang’s argument is that computers have evolved from unprofitable storage warehouses, where humans pre-recorded information and machines retrieved it, into revenue-generating factories. Those factories now manufacture tokens at scale, segmented and priced by quality. Low-end tokens run at around a dollar per million. Mid-tier at three to six dollars. High-end engineering-grade tokens at forty-five dollars and above.
In the warehouse model, compute was overhead. A cost you minimised. The machine stored what humans created and retrieved it when asked. The value was in the human knowledge. The machine was the filing cabinet.
In the factory model, compute is production. The data centre is not storing intelligence: it is manufacturing it, continuously, at a cost falling by an order of magnitude every year.
This reframe changes the question organisations need to be asking. Not “how much does compute cost?” but “how much intelligence are we manufacturing per pound spent, and how well are we directing it?”
Most finance departments, most HR teams, and most boards are not yet having that conversation. They are still treating AI spend as a software cost rather than a production input. That is the equivalent of an industrial company treating its factory floor as an overhead line.
The scale shift most organisations have not priced in
Agentic AI consumes one million times more tokens than a standard generative prompt.
Not a hundred times. Not a thousand times. One million times.
When people are not just asking an AI a question but running agents that plan, execute, revise, and iterate across complex workflows, token consumption moves into a completely different order of magnitude. One researcher put their personal weekly consumption at between one billion and ten billion tokens. An agent running full time on a coding task can burn through seven hundred million tokens in a week.
At platform level the numbers are already visible. OpenRouter grew from handling roughly ten trillion tokens per year to more than one hundred trillion in eighteen months, processing more than one trillion tokens every single day as of late 2025. Average session length has more than tripled in twenty months, driven almost entirely by agentic and coding workflows.
Most organisations are budgeting for the generative AI era. They are about to be running in the agentic one. The token consumption those two worlds require is not comparable.
The skill nobody is talking about yet
Here is the irony inside all of this.
The companies with the largest AI budgets and enterprise agreements are learning the least about how to use tokens well. The bill goes to IT. Nobody feels the constraint. Tokens get burned on tasks that did not need them, through models far more powerful than the job required, in workflows that loop expensively because nobody thought carefully about the architecture.
Meanwhile my sons, working with a budget scraped together from Lego sales, are developing exactly the optimisation instincts that will matter most as agentic workloads scale.
They have to choose the right model for the right task. A complex reasoning problem needs a frontier model. A simple classification or formatting task does not. Using the most powerful model for everything is the equivalent of chartering a transatlantic flight to get to the next town. The cost difference is real. The output difference is negligible. Knowing which model fits which job is itself a compounding skill.
They cache aggressively. Prompt caching means that context which does not change between calls, system instructions, background documents, established workflows, does not need to be re-processed at full cost each time. For an agent running thousands of iterations, the difference between a cached and uncached architecture is not marginal. It is the difference between a viable product and an uneconomic one.
They think about workflow architecture before they start spending. How a task is broken down determines how many tokens it consumes. An agent that loops back through the full context window on every step because the workflow was not designed carefully will burn ten times the tokens of one that was. The intelligence of the system is partly in the model. A significant part of it is in how the workflow is constructed around the model.
This is not a technical skill in the narrow sense. It is a resource management skill applied to a new kind of resource. And it is one that most organisations, and most professionals, have not started developing because they have not yet felt the pressure to.
They will. As agentic workloads scale and token costs become a visible line on the P&L rather than a rounding error on an IT invoice, the ability to architect for token efficiency will be as commercially important as the ability to architect for any other production cost. The organisations that develop it early will carry a structural advantage. The ones that wait will find themselves paying a premium for outcomes they could have achieved at a fraction of the cost.
Token budgets are the new capital allocation question
Huang described a future where every engineer at Nvidia would receive an annual token budget alongside their salary, giving engineers tokens worth roughly half their base pay on top of it so they could be amplified ten times. Speaking on the All-In Podcast, he described a thought experiment: a software engineer paid five hundred thousand dollars a year who consumed only five thousand dollars worth of tokens. His response was that he would, in his own words, “go ape something else.” If that engineer did not consume at least two hundred and fifty thousand dollars worth of tokens, he would be deeply alarmed.
The analogy he reached for: a chip designer who insists on working with paper and pencil rather than CAD tools. Not just slower. Someone who has misunderstood what their role now requires of them.
Huang at GTC 2026: “In Silicon Valley today, tokens are a recruiting tool. People are asking: how many tokens come with my job?”
That question will reach London, Edinburgh, Paris, and Frankfurt within eighteen months. Probably sooner.
But the more important implication is not recruitment. It is how organisations think about capital allocation itself.
Token budgets are not a technology expense. They are a human capital investment. Two people with identical talent will produce radically different results if one has a meaningful token budget and the other does not. The amplification is not marginal. It is an order of magnitude.
Which means the right question is not “what is our AI budget?” It is “which roles require the highest quality judgement, and are those roles resourced with the token capacity to act on it?”
Most organisations have it backwards. The senior people who carry the best judgement are the least likely to be deep inside the tools. The junior people closest to the tools often lack the context to deploy them with maximum effect. The token budget, when you look at how it is actually distributed, tells you exactly how seriously an organisation is thinking about where intelligence is applied.
And the harder question nobody is asking yet: who decides how tokens are allocated across a business, and on what basis? This is an operating model question wearing the clothes of an IT procurement decision. The organisations that see the difference will move differently.
What this means for how you value yourself
The old measure of professional value was legible: what can you do, how well can you do it, how long have you been doing it. Experience as accumulated capability.
The new measure has an additional dimension: how well do you direct cognitive compute toward problems, and how much useful output can you generate from a given resource allocation?
That is a different skill from expertise. It is closer to the skill of a good editor, or a good director. You are not the one doing all the work. You are the one who understands the problem clearly enough to direct the work well, and who can tell the difference between good output and expensive noise.
Sam Altman has proposed a version of this at the level of whole societies. He has described a future called Universal Basic Compute: everyone receives an allocation of AI compute they can use, sell, or donate. Not money. Productive capacity. The right to manufacture intelligence.
Whether that vision materialises or not, the underlying principle is already operational. Access to cognitive compute shapes what you can produce. The people who understand the economics of that, who treat their token allocation as a resource to be directed rather than a tool to be occasionally used, are already in a different category from those who do not.
The conversation has already changed in the places where this is most visible. It used to be “what are you building?” Now it is “how many agents do you have running?” That shift will move through every industry.
The honest part
None of this means token consumption equals value creation.
Satya Nadella said something at Davos worth holding onto. His argument: if tokens are not improving health outcomes, education outcomes, and business competitiveness across the board, then the social permission to use the energy required to generate them will not last.
There is already a term in engineering circles for the bad version: tokenmaxxing. Burning tokens to look productive, generating output that never ships. An organisation that measures token spend without measuring outcomes has confused the raw material with the product.
The equation is judgement multiplied by tokens, directed at problems that matter. Remove the judgement and you have an expensive machine running in circles. The token is not the point. What you manufacture with it is.
Back to the bedroom floor
I have spent years working with boards and executive teams on why organisations struggle to realise value from technology investments. The answer is almost always the same. The technology moves faster than the mental model people use to think about it.
Two teenagers, no management theory, no enterprise experience, had independently arrived at a clean understanding of the new economics. They knew what the resource was. They knew what it unlocked. They knew what they were willing to sell to get more of it.
That is the mental model organisations need to build. Not the technical understanding of what tokens are. The economic understanding of cognitive compute as a factor of production, one that sits in the chain between energy and output, and one that needs to be allocated with the same seriousness as any other resource a business depends on.
Compute equals intelligence. Intelligence equals output. Output equals value.
My sons already know this.
They sold the Lego.
If this piece resonated, subscribe. The next one goes inside what it actually looks like to architect an organisation around intelligence allocation rather than headcount. The difference is more radical than most leaders expect.
Craig Hepburn is an AI strategist and Perplexity Fellow who builds advanced agentic systems. He spent years leading digital transformation at Art Basel and UEFA. Now he works on the harder question: not whether organisations adopt AI, but how they govern it when they do.



such a great post on so many levels...