Every time an AI model reads or writes, it consumes tokens. These are the small units of text it works in, and each one represents roughly three-quarters of a word. You pay for how many you use. A chatbot exchange burns a few hundred. A simple agent consumes up to 15,000 tokens per task; a complex, multi-agent system uses from 200,000 to more than a million.
This is the new tokenomics (read: token + economics) of AI, and it’s no surprise academia is on the case. Stanford researchers found agentic coding uniquely token-dense, using around 1,000 times more tokens than ordinary code chat. The finance world is also on alert. Goldman Sachs expects total token consumption to multiply roughly 24 times by 2030, to 120 quadrillion a month. It’s the volume of work an agent does, more than the price of a token, that drives the bill.
So, what are P&L-minded IT teams — specialists who did not expect to also operate as economists — to make of this?
Bharat Patel, a solution architect at Dell Technologies Customer Solution Center, puts the consequence plainly: Where you run AI matters as much as which AI you run.
Cheaper tokens, bigger bills
Between mid-2023 and early 2026, token prices fell by around 80%, but companies didn’t bank the savings. In the same period, enterprise AI spending increased by roughly 320%. Companies deployed agents consuming far more tokens, then ran more of them across more workflows. A lower price per token met dramatically higher consumption, and the total climbed.
“That shift,” Patel said, “means the entire supporting system — the hardware, security, budget model, and management tools — has to shift as well.”
For businesses, what matters is cost per outcome: the business result that spending delivered. An agentic workflow does not answer once and stop. It reads files, forms a plan, checks the output, revises, and then loops until the task is done. Each step is a separate call that resends the entire accumulated context.
Agentic AI shifts enterprise spending from the predictable cost of software licenses to variable, consumption-based compute, where the token bill is only the visible part of a wider shift in infrastructure, governance, and regulation. Usage and costs become board-level questions.
The answer, Patel said, is to run the everyday tasks on hardware you already own. That way, the marginal cost of a token falls toward the price of the electricity to run it. Metered cloud use for the largest frontier models should be called in like a specialist, not used as a workhorse.
From desk to data center
This is the gap that Dell Deskside Agentic AI fills. Launched in May, it lets a work group run production-ready agents at the desk, on Dell workstations with the Nvidia NemoClaw open-source stack. It handles validated workflows for coding, research, and private assistants on compact, efficient models with 30 billion training variables, up to the largest cutting-edge trillion-parameter models.
The agents run inside an OpenShell environment that enforces policy-based privacy and security rules and logs each agent’s actions, putting governance in place from the start. Using OpenShell means that once a prototype needs larger infrastructure, it can move to Dell’s PowerEdge servers in the data center, without revising the architecture.
It’s aimed at those who cannot send work to the public cloud: engineers running coding agents that must keep source code in-house, or researchers analyzing pre-publication and patient data in compliance with privacy regulations.
With this system in place, the tokenomic indicators are strong. Analysis by Signal65 and Futurum, commissioned by Dell, puts the savings at up to 87% on token spend over two years compared to public-cloud APIs, with break-even in as little as three months.
Patel summarizes his advice to leaders in six words: Start local, govern early, scale smart. The cost of agentic AI is now a design decision, taken before the first agent ships. If you decide where each workload runs and put governance around it early, the bill becomes something you manage rather than something that happens to you.
These implications mean a token strategy should transcend IT. “It’s no longer just an IT conversation,” Patel said. “It’s a whole company conversation.”
Find out more about Dell Deskside Agentic AI and reduce your AI token costs.
This sponsored post was created by BI Studios with Dell AI Factory with Nvidia.

