Technology

Stop Measuring AI by Tokens: Build Unit Economics Before the Bill Arrives

Server racks representing the infrastructure behind AI operating costs
Photo by Kevin Ache on Unsplash (Unsplash License)

AI costs rarely fail all at once. They drift: a longer context here, an extra agent loop there, a premium model used for work a smaller model could handle. The invoice grows while the dashboard still celebrates more tokens, more calls and more users.

That is why this week’s most useful AI story is not another benchmark. On 11 September, IBM argued that enterprise AI investment is outpacing the ability to connect cost with business value. Its guidance says most organisations can see token and cloud charges, yet still struggle to tie total cost of ownership to outcomes. Four days later, Anthropic highlighted model defaults, per-user visibility and usage reporting as practical ways to make consumption pricing more predictable.

The signal is clear: AI has moved from experimentation into financial operations. Founders and technical teams now need to manage it as a product with unit economics, not as an API bill.

A token is a meter, not a unit of value

Cost per token is useful for comparing model rates. It does not tell you whether the system created value. A cheap response that is wrong, rejected or escalated can cost more than an expensive response that completes the work correctly.

The FinOps Foundation’s unit-economics guidance separates resource measures, such as cost per token, from business measures, such as cost per transaction or case resolved. For AI products, the second category should drive the first.

Choose one outcome that reflects the job:

  • customer support: cost per case resolved without reopening;
  • sales: cost per qualified lead accepted by the sales team;
  • software delivery: cost per AI-assisted change that passes review and production checks;
  • document operations: cost per document extracted above a defined accuracy threshold;
  • security: cost per valid incident triaged within the service target.

The quality condition matters. Without it, teams can make the metric look better by producing more low-value output.

Capture the fully loaded cost

The model invoice is only one line. A useful cost record also includes retrieval and vector storage, data pipelines, tool and search calls, orchestration, retries, guardrails, evaluation, observability, human review and the infrastructure that serves the experience. For an agent, one user request may trigger several model calls and external actions before it returns an answer.

Tag every production request with a product, environment, team, use case and outcome identifier. Record model, input and output tokens, latency, tool calls, retries, cache hits and review status. Then join that trace to billing data. If a shared service cannot allocate an exact charge, use a documented rule and improve it over time; an honest estimate is more useful than an unowned total.

This also reveals where optimisation belongs. A rising model bill may be healthy if completed work grows faster. A stable bill may be unhealthy if quality falls and people quietly redo the work.

Design a cost envelope into every workflow

Budgets should exist inside the runtime, not only in a monthly finance report. Define a maximum cost, latency and number of steps for each workflow. When a run reaches a boundary, it should stop, downgrade, ask for clarification or hand the task to a person.

Practical controls include:

  • Model routing: use the least expensive model that can meet the quality target, and escalate only when confidence or task complexity requires it.
  • Context discipline: retrieve narrowly, summarise durable history and prevent entire documents or conversations from being resent on every turn.
  • Loop limits: cap agent steps, tool calls and retries; require approval before an expensive branch.
  • Caching and batching: reuse stable answers, embeddings and repeated preprocessing where freshness rules allow.
  • Workload scheduling: pause non-urgent processing when demand or pricing makes it inefficient. Google Cloud’s 14 September Dataflow update, for example, made pause/resume generally available for long-running batch jobs so failed or interrupted work need not always restart from zero.

Automate cost control in stages

AI can investigate anomalies and recommend savings, but it should not receive broad power to delete or resize production resources on day one. AWS’s 10 September FinOps guidance describes a sensible progression: read-only insights, human-approved changes, rule-based automation, then autonomy within explicit boundaries.

That ladder is provider-neutral. Start by asking automation to explain a cost spike and show its evidence. Next, allow a reversible action after approval. Only automate recurring changes once exclusion rules, permission scopes, maximum financial impact, rollback and audit logs have been tested.

Cost optimisation can damage reliability if teams chase savings without operational context. Every recommendation should display expected savings beside latency, capacity and service-risk effects.

A four-week operating plan

Week one: pick the outcome. Select one production AI workflow and name its owner. Define success, failure and the unit that the business actually values.

Week two: instrument the path. Add trace identifiers and capture model, infrastructure, tool and human-review costs. Build a daily view by use case rather than by cloud account alone.

Week three: set the envelope. Establish thresholds for cost per successful outcome, agent steps, latency and error rate. Add alerts before the monthly budget is exhausted.

Week four: run two experiments. Test model routing, shorter context, caching or batching. Compare total cost and accepted quality against the same baseline. Keep only changes that improve the complete unit, not just token spend.

The Qomra Tech view

The teams that scale AI well will not necessarily buy the cheapest tokens. They will know what one reliable outcome costs, which technical choices move that number, and when automation is allowed to act.

For founders, this is a margin and product decision. For engineers, it is an observability and architecture decision. For finance, it is an allocation and forecasting decision. Put those views in one operating loop now, before usage grows faster than accountability.

Let's talk

Tell us about your project.

We'll come back within one business day with the right person to talk to.

Trusted by founders across healthcare, hospitality and professional services. London HQ · Bilingual EN/AR delivery · NDA-friendly