Sep 10, 2026
The Token Bill Is Coming. Most Companies Aren’t Ready.
Everyone is integrating AI agents. Almost nobody is estimating what they’ll cost. There’s a conversation happening in engineering teams right now that didn’t exist two years ago. It used to go like this: “How long will this take?” The answer came in hours, which translated into cost through a relatively simple formula most clients understood. Now the conversation goes differently. AI agents enter the project. The …
Everyone is integrating AI agents. Almost nobody is estimating what they’ll cost.

There’s a conversation happening in engineering teams right now that didn’t exist two years ago.
It used to go like this: “How long will this take?” The answer came in hours, which translated into cost through a relatively simple formula most clients understood.
Now the conversation goes differently. AI agents enter the project. The work gets done faster. And then someone asks: “How much did that cost in tokens?” And the room goes quiet because most teams don’t have a good answer.
Token-based costs are the new black box of software development. And until companies develop a structured way to estimate them, every AI-assisted project carries a financial risk that nobody is explicitly accounting for.
Why Hours No Longer Tell the Whole Story
The traditional model of estimating software work was built around human time. A developer takes X hours to build Y feature. Multiply by rate. Add buffer. Deliver estimate.
AI agents break this model in both directions simultaneously.
On one side: an agent can do in minutes what a developer spent hours on. The throughput is real — a single senior developer working with AI agents can close work that previously required a larger team, and close it faster.
On the other side: agents consume tokens in ways that don’t map cleanly onto human hours. A request with a large context window costs more than a request with a small one, regardless of how “difficult” the task seems. A poorly specified instruction can cause an agent to make multiple attempts at the same task, each consuming tokens. A workflow with many sequential agent calls can accumulate costs that aren’t visible until the invoice arrives.
Hours measure human time. Tokens measure computational consumption. They’re related but they’re not the same thing and treating them as equivalent leads to estimates that are either significantly over or significantly under the actual cost.
The Black Box Problem
The most common response to token costs right now is avoidance. Teams integrate AI agents, observe that the work gets done faster, and don’t look too closely at what’s happening on the cost side until something forces them to.
What forces them to is usually a bill that’s higher than expected.
This isn’t a technical failure. It’s a planning failure. Token costs are entirely predictable but only if you define the right inputs before you start, not after you’ve already deployed.
The problem is that “approximately $50 in tokens” is as meaningless to a non-technical founder as “approximately 40 hours of development” was before that became a normalized unit. A number without a framework behind it doesn’t reduce uncertainty. It just delays the moment when the uncertainty becomes visible.
The Framework: Criteria First, Tokens Second
The solution isn’t to avoid token-based estimation. It’s to approach it with the same structural discipline that good hour-based estimation requires.
What will the agent actually do? “Build a customer support agent” is not a specification. “Receive a customer message, retrieve relevant documentation from a knowledge base with a maximum context of X tokens, generate a response, and flag messages that require human review” is a specification. The difference in token consumption between these two descriptions is not marginal.
What data will the agent receive and generate? The size of the context window is one of the primary drivers of token cost. A support agent that has access to the full history of every customer interaction is dramatically more expensive to run than one that accesses only the current conversation and relevant documentation. Both might produce acceptable results. Only one of them has been costed correctly.
How many times will the agent be called? Token costs scale with usage. Many token cost surprises come from actual usage volume being higher than assumed usage volume in the original estimate.
What is the cost per token for the model being used? Different models have different pricing. A workflow that uses a frontier model for every step costs more than one that uses a frontier model for complex reasoning and a cheaper model for routine tasks. Model selection is a cost lever that’s often treated as a technical decision when it’s equally a financial one.
From Tokens to Limits
Once the criteria are defined, the translation to a predictable budget is arithmetic.
Average context size × average number of agent calls × cost per token = cost per unit of work. Scale by expected usage volume. Add buffer for edge cases. The result is a number that can be explained, defended, and adjusted.
This is what token estimation looks like when it’s done correctly:
This customer support workflow will make approximately 500 agent calls per day. Each call has an average context of 2,000 tokens for input and generates approximately 500 tokens of output. At current pricing for the model we’re using, that’s roughly $X per day, or $Y per month at expected usage levels. If usage doubles, the cost scales proportionally. Here’s where we’d optimize first if we needed to reduce costs.
Compare that to: “Token costs will be approximately $50 per month.”
The first version gives a client or investor something they can evaluate. The second gives them a number they have to trust blindly.
Why This Matters for Non-Technical Founders
Token costs don’t respond to the same intuitions as hour-based costs. A workflow that seems simple can be expensive if the documents are long, if the summarization quality requires a frontier model, and if the workflow runs at scale. A workflow that seems complex can be cheap if it’s well-specified and the agent rarely needs to retry.
The founders who navigate this well are the ones who’ve learned to ask the right questions before the work starts. Not “how much will this cost” but “what are the inputs that determine the cost, and how do we bound them.”
The EU AI Act’s transparency requirements, now in force for AI systems deployed in Europe, add a compliance dimension: organizations need to be able to account for how their AI systems operate, including their cost structure. Token estimation done well isn’t just good financial practice — it’s part of the governance framework that regulators are increasingly expecting.
The Logic Worth Remembering
Token estimation isn’t harder than hour estimation. It just requires a different set of inputs and a willingness to define them explicitly before the work starts.
Criteria → token estimate → limit → predictable budget.
The teams that develop this discipline early will be able to make commitments about AI-assisted project costs with the same confidence that experienced development teams make commitments about hour-based costs. The teams that don’t will keep getting surprised by their bills.
Wamisoftware has been building software products since 2014. We work with senior engineers and AI agents on projects where the cost of what gets built is as important to understand as the quality of what gets built.


