top of page

A Declaration of Independence from AI Token Waste and Excess Charges

  • Jul 7
  • 6 min read

Last week, the United States marked two hundred and fifty years since its founding, and it is worth remembering what actually provoked the founders. The colonists who boarded ships in Boston Harbor in December 1773 were not protesting a large tax. The duty on tea was three pence a pound, and the Tea Act had in fact made tea cheaper. What they refused was the arrangement itself: a small, metered charge on every pound, levied by a distant monopoly, with no voice in the terms, arriving on top of a decade of stamp duties, customs fees, and navigation restrictions. Small unit charges, compounding at scale, imposed without representation. That was one of many grievances worth a revolution.


Artificial intelligence (AI) has brought us astonishing tools, and I have spent the last three years teaching business leaders how to use them. But the current AI economics have a familiar shape. Enterprises and API users now pay for intelligence by the token, a small metered charge on every unit of work, set by a handful of distant providers, on terms no customer negotiates, and at a scale where fractions of a cent compound into budget lines that boards are beginning to question. And these tokens are charged on throughput, regardless of whether the output is useful.


The frustration is no longer confined to private conversation. On Wednesday, July 1, 2026, Palantir CEO Alex Karp told CNBC's Squawk Box that something in the token model "has gone completely wrong," and that "every single enterprise I deal with is livid," with customers telling him "I am paying for tokens that create no value." I hear the same sentiment, in nearly the same words, in one-on-one conversations with business executives at all enterprise levels. The public outrage and the private grumbling have converged.


Some of the grievance, though, is self-inflicted, and this is the part we can control. Most people use AI sloppily, not from carelessness of character but because they never before needed to think about the cost of a word. The habit formed in an unmetered world, and it persists in a metered one. Field reports from operators who have audited their production stacks suggest that 40 to 60 percent of token budgets in production applications is pure waste: money paid for capability never used, or for inefficiencies nobody priced at the design stage.


Here is the good news. If tokens are the measured medium, then the medium can be managed, and managed dramatically well. I will add one candid caution: it is possible that as enterprises grow efficient, providers will devise new ways of charging for what amounts to oxygen for AI agents. That prospect is an argument for the discipline, not against it, because the organizations that measure their consumption are the ones that notice when the terms change, and the ones with the standing to push back. Karp's remedy is to own the means of production. For most organizations, the nearer remedy is to govern the consumption.

That is what my new book is about: identifying and stopping token waste, and implementing the procedures that turn AI spend into a strategic advantage. Today I published TokenOps:


From Token Waste to Competitive Advantage, and its argument begins with a pattern we have seen before.


The Pattern We Have Seen Before

In 2012, the average enterprise cloud bill was a curiosity item on the desk of the chief financial officer. By 2018, it had become a top-five line in many technology budgets, and a discipline called FinOps had emerged to govern it. The companies that built FinOps competence early found themselves operating cloud infrastructure at materially lower unit costs than their less disciplined competitors, and the gap widened year on year.


That sequence was not an accident of the cloud era. It is the characteristic pattern by which new operational disciplines come into being. A previously unmetered or flatly priced resource becomes commercially metered. Organizations adopt the resource faster than they can measure it. Costs accumulate silently and unevenly across teams. Executive attention eventually arrives, and a discipline emerges to bring order, supported by tooling, certification, and dedicated roles.


Tokens are now traversing the same trajectory. They flow through every prompt, every retrieval-augmented query, every agentic workflow, and every multi-agent system. In 2024, most organizations paid for tokens without measuring them. In 2026, the leading practitioners measure them carefully. By the end of the decade, I expect token governance to be a recognized operational function with its own tooling, its own role definitions, and its own seat at the executive table.


The scale of the opportunity warrants the attention. Industry estimates place enterprise spending on large language models at approximately $8.4 billion in 2025, more than double the previous year's figure, with consensus forecasts suggesting another doubling is plausible by the end of 2026. A Deloitte analysis published in early 2026 identifies artificial intelligence as the fastest-growing expense in corporate technology budgets, with some firms reporting that it consumes up to half of their information technology spend. Set against those figures, the waste number bears repeating: 40 to 60 percent of production token budgets is buying nothing at all.


The book is a foundation document for the emerging discipline, and it follows the naming precedent that FinOps established: treat tokens as the primary currency of artificial intelligence work and the primary object of measurement and optimization.


A Gauge with Five Bands

The book organizes the journey as a gauge whose needle sweeps across five bands. Organizations begin in Waste, where consumption is unmeasured. They move to Efficiency as individual habits eliminate the obvious losses, and then to Optimization as routing, caching, and prompt structure become deliberate engineering choices. Leverage follows, when context architecture and agentic design allow the same token to do compounding work across systems. The final band is Advantage, where token discipline becomes an institutional capability that competitors cannot quickly replicate. The subtitle of the book traces the needle's full sweep.


The Orchard Principle

I frame the practice with what I call the Orchard Principle. Left unattended, a young tree grows enthusiastically in every direction, producing branches, leaves, and complexity, but not necessarily fruit. A master gardener prunes without hesitation, removing healthy branches so the tree directs its energy toward what matters most. Artificial intelligence behaves much the same way. Left unmanaged, prompts become verbose, context accumulates, agents multiply, models are overpowered for simple tasks, and token consumption expands almost invisibly. TokenOps is the discipline of pruning, and the objective is not to consume fewer tokens but to cultivate an ecosystem that produces more insight, better decisions, and greater advantage from every token invested. A wild tree grows; an orchard is designed.


What Is Inside

The playbook contains sixty-five strategies across seven categories, moving from input and output discipline through format strategy, prompt engineering, model routing, context and caching, and the architecture of agentic systems. The strategies are model-agnostic; they apply whether your teams work in Claude, ChatGPT, Gemini, Grok, or open-source models such as Llama and Qwen, because the underlying mechanics of tokenization, the asymmetric pricing of input and output, and the compounding behavior of context are stable across providers.


The book is also built for different readers. An executive can extract the strategic frame and the highest-leverage governance levers in roughly twenty minutes. A technology leader or solution architect will find the structural material on routing, caching, and agentic architecture in the later sections. A hands-on practitioner can begin with ten habits that take fifteen minutes to learn and compound daily thereafter.


What the Discipline Returns

The evidence base for the effort is substantial. Multiple independent practitioner surveys and production case studies place the realistic savings range for disciplined token usage at 60 to 80 percent of pre-optimization spend, without compromise to output quality. The most disciplined teams, those that layer caching, routing, compression, and retrieval so that the techniques compound, report savings in the range of 80 to 90 percent. The gain is not incremental; it is the difference between AI economics that scale and AI economics that quietly consume the budget innovation needed.


Every era of enterprise technology eventually acquires its discipline of stewardship. Cloud found FinOps. Artificial intelligence is finding TokenOps now, and the organizations that build the competence early will hold the same widening advantage that the early FinOps adopters enjoyed. Two hundred and fifty years on, the lesson of 1773 still holds: the remedy for a metered charge you cannot see is not resentment. It is representation, measurement, and a discipline of your own.


Optimize every input. Amplify every outcome.

And the news gets better. I think most, if not all, of the strategies inside my book TokenOps can be automated. The question of what that looks like for your business or operation is one only you can frame. So I invite you to enjoy your weekend, ponder the significance of Independence Day, and on Monday, dive into this book with the intentional curiosity that leads to action.


TokenOps: From Token Waste to Competitive Advantage is available now on Amazon in Kindle, with paperback to follow: https://a.co/d/02NzXwd4


Copyright © 2026 by Severin Sorensen. All rights reserved.

Comments


bottom of page