FuturumBecome a Client
Research Report

The Token Cost Reckoning Arrives in the CFO’s Office

Brendan Burke, Daniel Newman

Updated

The short answer

In our latest thought leadership brief, The Token Cost Reckoning Arrives in the CFO’s Office, completed in partnership with Google Cloud, Futurum Research examines how enterprises can track cost per successful agent run and cut token spend before their next budget review.

Futurum's Brendan Burke and Daniel Newman,

Cite this

Brendan Burke and Daniel Newman, The Futurum Group, "The Token Cost Reckoning Arrives in the CFO’s Office," September 9, 2026. https://trial.futurumgroup.com/research-reports/the-token-cost-reckoning-arrives-in-the-cfos-office/

The Token Cost Reckoning Arrives in the CFO's Office

Enterprise agents have started calling other agents, and the token volume behind that traffic looks nothing like the earlier wave of employees typing into a chatbot. Vendors priced tokens below cost for the past two years to build share, and those subsidies are ending as bills catch up to real usage at scale. The habit of pointing every task at the largest available model, regardless of cost, is losing ground as finance teams open the invoice and ask what they can actually afford.

Finance already tracks cost per transaction on the cloud bill, and AI spending needs that same discipline applied to tokens. Only 36% of AI Platform decision-makers track cost per request today, and just 15% track token efficiency, according to Futurum Research’s 1H2026 AI Platforms Decision Maker survey. Cutting costs before switching hardware starts with four moves: caching the prompt prefix, routing routine calls to a cheaper model, batching non-urgent work, and capping retries on failed runs.

In our latest thought leadership brief, The Token Cost Reckoning Arrives in the CFO’s Office, completed in partnership with Google Cloud, Futurum Research examines the token economics reshaping enterprise AI budgets and lays out the architecture questions that determine what an agent workload actually costs to run.

In this report, you will learn:

  • How to calculate cost per successful run (CPS), the metric that connects agent spend to what customers actually accept
  • Four practical moves that cut token costs before a hardware switch: prompt caching, model routing, batching, and retry caps
  • Why compute silicon choice, not just model choice, drives the blended cost of every agent transaction
  • How software portability lets workloads move across accelerators without paying twice for idle capacity
  • What questions finance should bring to the CIO at the next budget review

If you are interested in learning more, be sure to download your copy of The Token Cost Reckoning Arrives in the CFO’s Office today.

Published by Futurum.

More from Brendan Burke