Loading...

THE HIDDEN COST OF AI

AI may be getting cheaper by the token, but enterprise AI bills are rising as we ask AI to do far more than simply answer questions.

The alarm bells rang when reports of an organisation spending an astonishing $500 million on Claude in a single month made everyone sit up. The startling numbers revealed a hidden story about the economics of AI at scale.

As enterprises move from experimenting with chatbots to deploying AI across coding, customer service, analytics and autonomous agents a paradox is playing out: while the cost of tokens is falling, the explosion in AI adoption within the organization is sharply pushing up the cost of AI.

Reuters recently reported that companies are already confronting unexpectedly high AI bills, with some exhausting budgets months ahead of schedule. Savvy organizations are recalibrating the AI strategy and increasingly choosing smaller models for routine workloads, allocating expensive frontier models for complex tasks to justify the cost.

Dig a little deeper and a new reality is revealed in that AI is no longer just answering questions, but it is actually performing tasks as virtual employees. Intelligent agents may reason through a problem, retrieve information, call multiple tools, generate code, evaluate its output, retry when something fails and persist until the job is done. Each step triggers additional model calls and token consumption. This is what changes the economics of the AI deployment from what it was before.

Gartner’s latest forecast captures this emerging ‘inference paradox.’ While the underlying cost of AI inference is falling, Gartner predicts that inference costs per agentic workflow could increase more than fivefold through 2028, driven by the growing complexity and token consumption of agentic AI.

Crucially, the economics for enterprise AI is rapidly shifting from measuring the cost of the model to assessing the cost required to complete a business task.

Google Cloud India Managing Director Sashikumar Sreedharan has also pointed to the broader economics behind AI. The cost equation extends beyond tokens to the infrastructure underneath them—compute, power, cooling and data centres.

And those costs are under pressure too. According to a Reuters report, Nvidia has told some customers that AI-server prices could rise by more than 15% because of soaring memory costs, signalling the pressure on the physical infrastructure powering the AI boom.

This calls for a reality check for enterprises. According to IT heads of leading organizations, the smartest strategy may not be to pick the most powerful model but to match the right model with the right task. For example, a simple query does not need a frontier reasoning model, while a complex business process such as financial analysis or risk assessment requires complex models. AI platforms such as OpenRouter are consequently moving towards model routing—automatically sending workloads to different models based on performance, latency, security and cost.

Today, AI FinOps is becoming as important as cloud FinOps, because the real cost of AI is not just the price of tokens but its uncontrolled consumption. As AI becomes more useful and autonomous, enterprises will delegate more tasks to it, which will consume more inference in the process. The challenge, therefore, is not just to track how much AI is being used, but whether that usage is delivering measurable business value.

The winners in the next phase of enterprise AI will not be those that use the most AI, but those that know where AI delivers the maximum value.

About The Author