Risk Transfer in the Generative AI Economy

How Usage-Based Pricing Is Reshaping the Allocation of Performance, Cost, and Operational Risk

8/16/202637 min read

a close up of a one dollar bill
a close up of a one dollar bill

How Usage-Based Pricing Is Reshaping the Allocation of Performance, Cost, and Operational Risk

Independent Research Report | August 2026

Executive Summary

Generative artificial intelligence is changing not only the economics of software but the contractual and operational relationship between technology suppliers and their customers. The first phase of commercialization was dominated by access to models: enterprises paid for subscriptions, application-programming-interface calls, tokens, generated images, or dedicated inference capacity. The next phase is increasingly defined by systems that reason across longer contexts, call external tools, conduct searches, execute code, use computer interfaces, coordinate subtasks, and operate with varying degrees of autonomy. The economic significance of this transition is greater than a simple increase in model capability. In traditional software, the customer typically buys a relatively deterministic capability and decides how intensively to use it. In generative and agentic AI, the system's own behavior can increasingly determine how much computation is required to complete the customer's objective. This is creating a widening distinction between the technical unit for which the provider charges and the business outcome for which the customer believes it is paying.

The market is moving into this problem at considerable scale. Gartner forecast worldwide AI spending of approximately $2.59 trillion in 2026, up 47 percent year on year, although that broad figure includes infrastructure, devices, software, and services rather than model inference alone. Gartner separately projected direct end-user spending on generative-AI models of approximately $14.2 billion in 2025, with the model segment expected to grow materially as enterprise adoption expands. McKinsey's 2025 global AI survey found that almost nine in ten respondents reported regular organizational AI use, while 62 percent said their organizations were at least experimenting with AI agents; 23 percent reported that agentic systems were already being scaled in at least one business function. Yet the same survey found that nearly two-thirds of respondents had not begun scaling AI across the enterprise and only 39 percent reported enterprise-level EBIT impact from AI. The commercial environment is therefore characterized by a striking combination of high adoption, accelerating spending, and incomplete translation from technical use into enterprise-wide financial value. (Gartner)

The pricing architecture of the leading model ecosystem reinforces this transition. OpenAI prices GPT-5.4 at $2.50 per million input tokens and $15 per million output tokens under standard API pricing, with additional pricing structures for cached input, longer contexts, batch processing, and priority service. Anthropic's published model pricing similarly distinguishes input, output, cache-write, cache-read, batch, and long-context usage. Google's Gemini API charges separately for input, output, context caching, storage, grounding, and several service tiers, and its agent pricing explicitly states that model inference—including intermediate reasoning and input generated during agentic loops—is charged at standard model rates while tool fees can apply separately. Amazon Bedrock likewise meters input, output, cache-read, and cache-write tokens and provides separate pricing arrangements across models, inference modes, and other services. The industry has therefore converged substantially on metering computational activity rather than measuring whether the resulting business task was accepted, deployed, or economically successful. (OpenAI Developers)

This report argues that the resulting issue is best understood as risk allocation rather than simply pricing. Usage-based pricing is economically rational in many circumstances. Providers incur real infrastructure costs when models are invoked, whether or not a particular response ultimately satisfies the customer. Consumption pricing lowers entry barriers, lets customers scale incrementally, and creates a transparent relationship between compute usage and supplier revenue. The structural problem emerges when the amount of computation necessary to achieve an acceptable outcome is itself materially influenced by uncertain model performance. An incorrect answer may require another generation; incomplete reasoning may require a larger context; a failed tool call may require recovery; an agent may use additional model invocations to verify its own output; and a high-consequence workflow may require a second model, retrieval layer, evaluator, or human reviewer. In each case, some of the cost of resolving uncertainty can become additional consumption.

The FinOps community is beginning to formalize this problem. The FinOps Foundation describes AI as a distinct cost-management domain because of granular consumption, spend unpredictability, rapid pricing changes, multi-provider purchasing, and the growing need to connect consumption to business value. Its 2025 State of FinOps research found that 97 percent of respondents investing in AI planned investment across multiple infrastructure areas, while the most important early AI FinOps activities centered on understanding usage and cost and quantifying business value. Subsequent FinOps guidance argues explicitly that advertised per-token pricing can be misleading because total cost depends on context growth, modalities, caching, routing, orchestration, and other operational factors; its 2026 work increasingly refers to the need to move from token counting toward use-case economics and “true cost ownership.” This development mirrors the earlier rise of cloud FinOps but with one important difference: generative-AI workload intensity can be partly endogenous to the behavior of the model or agent rather than determined entirely by deterministic application demand. (data.finops.org)

The source thesis motivating this report describes the most severe form of this mechanism as risk laundering: technical uncertainty generated upstream is converted into downstream financial, operational, legal, or validation exposure while the provider remains paid for the execution that created or attempted to resolve it. That framing is analytically useful but should not be overstated. Current public evidence does not demonstrate that major model providers deliberately engineer unnecessary retries, degrade performance to increase token consumption, or systematically maximize revenue through model failure. In fact, competitive pressure creates substantial incentives in the opposite direction. Models that solve tasks more reliably can win workloads; lower token usage can improve customer economics; better latency can increase adoption; and providers themselves bear infrastructure costs for additional execution. The more defensible conclusion is structural: execution-based pricing does not inherently distinguish consumption generated by additional customer value from consumption generated by the cost of obtaining the value already sought.

This distinction will become increasingly important as agentic systems scale. Google already states explicitly that inference generated during agentic loops—including intermediate reasoning tokens—is billable under its agent pricing model. FinOps guidance similarly identifies multi-agent and autonomous workloads as a major source of nonlinear cost growth and warns that uncontrolled agents can generate rapidly compounding bills. McKinsey's finding that 62 percent of surveyed organizations were at least experimenting with agents in 2025 suggests that this is moving from an architectural edge case toward an enterprise-management issue. As autonomous systems gain permission to initiate searches, execute code, use paid tools, retrieve additional context, and invoke other agents, organizations face a growing mismatch between execution velocity and governance velocity. (Google AI for Developers)

The central management implication is that price per token, model call, or inference will become an increasingly incomplete measure of AI economics. Enterprises need to know the cost per accepted outcome, the amount of consumption attributable to retries and remediation, the human-verification burden associated with model output, the distribution of execution cost rather than only its average, and the maximum financial exposure that a single authorized autonomous workflow can create. Vendors, in turn, will increasingly need to demonstrate that rising consumption reflects rising customer value rather than uncontrolled execution complexity. The commercial winners are unlikely to be determined simply by the cheapest tokens or the highest benchmark scores. They are more likely to be determined by the ability to combine model quality, predictable task economics, expenditure containment, observability, and credible allocation of downside.

The report therefore does not recommend universal outcome-based pricing. Outcome definitions can be contested, customers control important parts of deployment, providers incur costs irrespective of customer acceptance, and outcome guarantees could create new adverse incentives. The more plausible market direction is a hybrid economic architecture: usage pricing will remain foundational, but enterprise products will increasingly add task-level budget controls, bounded retries, deterministic termination, quality commitments, selective service credits, risk-sharing provisions, and outcome-oriented metrics capable of translating computational consumption into business economics. The deeper strategic shift is from selling intelligence as a metered technical resource toward governing intelligence as a variable-cost business input.

1. AI adoption has moved faster than enterprise value realization, increasing pressure on the economics of deployment

The generative-AI market has moved in three years from novelty to near-universal strategic relevance, but organizations remain much less mature in converting adoption into enterprise economics. McKinsey's early-2024 survey found that 65 percent of respondents said their organizations regularly used generative AI in at least one business function, almost double the rate reported only ten months earlier; overall AI adoption rose to 72 percent. By the 2025 survey, nearly nine in ten respondents reported regular AI use, but most organizations remained in experimentation or pilot phases rather than scaled transformation. McKinsey found that almost two-thirds had not begun scaling AI across the enterprise, while only 39 percent reported an enterprise-level EBIT impact despite substantially broader adoption. These results do not imply that AI investments are failing; enterprise technology transformations often require significant time before financial benefits emerge. They do indicate that adoption metrics are substantially ahead of value-realization metrics and that the economic discipline surrounding AI deployment is becoming more important as experimentation converts into recurring operational expenditure. (McKinsey & Company)

The expenditure environment is expanding simultaneously. Gartner forecast worldwide generative-AI spending of $644 billion in 2025, although this aggregate includes hardware and AI-enabled devices and should not be confused with direct inference spend. Its narrower estimate of end-user expenditure on generative-AI models was $14.2 billion for 2025. In May 2026, Gartner forecast overall AI spending of $2.59 trillion in 2026, 47 percent above the prior year, with infrastructure accounting for a large share. IDC has separately described AI investment as entering a second wave in which enterprise spending increasingly shifts toward applications and agentic workflows rather than infrastructure alone. Forecast methodologies differ substantially across Gartner, IDC, and other research firms, so their headline estimates should not be added or treated as directly comparable. Their directional convergence is nevertheless clear: a growing share of enterprise technology expenditure is being exposed to AI-specific operating economics rather than one-time experimentation budgets. (Gartner)

This matters because the financial-management question changes once AI becomes embedded in recurring workflows. During experimentation, an enterprise can tolerate weak cost attribution because the primary objective is learning. At production scale, finance organizations require an answer to a more demanding question: what unit of business output is being purchased by each incremental dollar of AI consumption? Traditional enterprise software provides several familiar denominators—cost per seat, transaction, order, ticket, user, server, or business process. Model APIs instead expose technical measures such as input tokens, output tokens, cached context, image generation, search grounding, tool use, or compute time. These are legitimate measures of consumption, but they sit one or more layers below the outcome CFOs ultimately need to evaluate.

The resulting management gap is increasingly visible in FinOps research. The FinOps Foundation's 2025 survey found that the most important early capabilities for managing AI were not sophisticated optimization techniques but the basic disciplines of cost allocation, data ingestion, reporting, anomaly detection, planning, forecasting, and quantification of business value. This is consistent with an immature but fast-scaling cost category: organizations first need to understand where money is being spent before they can optimize it. The importance of this point is easy to underestimate. Generative AI is frequently introduced through decentralized developer tools, embedded SaaS capabilities, multiple cloud providers, direct model APIs, and product-team experiments. AI consumption can therefore spread across organizational budgets before a company has established a common economic model for evaluating it. (data.finops.org)

For enterprise leaders, the first conclusion is therefore not that AI is too expensive or that token pricing is inappropriate. It is that AI value accounting is structurally less mature than AI adoption. Organizations that scale agents, copilots, automated workflows, or generative customer experiences before connecting their technical usage data to business outcomes risk creating a large variable-cost category with weak economic observability. The experience of public cloud suggests that cost management will eventually mature; the distinction is that AI introduces probabilistic work generation on top of variable infrastructure consumption, making the mapping between usage and value substantially more complex.

2. The market has converged on execution-based pricing because computation is measurable, scarce, and costly

The widespread use of token and invocation pricing is not arbitrary. Foundation models consume expensive infrastructure, and the amount of infrastructure consumed varies with model size, input length, output length, modality, context, processing tier, caching, and tool use. Consumption therefore provides an administratively efficient method for allocating cost. OpenAI's GPT-5.4, for example, is priced at $2.50 per million input tokens and $15 per million output tokens under standard pricing, with cached input priced lower and longer-context sessions subject to different economics after specified thresholds. Anthropic's public pricing similarly distinguishes base input, output, caching, batch processing, and premium long-context use. Google's Gemini API differentiates standard, batch, flex, and priority tiers, context caching, grounding, modality, and agentic use. AWS Bedrock extends the model further across a multi-provider platform with separate rates for model families, on-demand or batch processing, caching, provisioning, and associated services. (OpenAI Developers)

These structures reveal an important economic fact: generative AI is not one product with one cost driver. The supplier's underlying expense is influenced by several forms of consumption, and token accounting provides a relatively portable abstraction for allocating them. Even within token pricing, however, not all tokens have equivalent economics. Output tokens generally command substantially higher prices than input tokens. Cached input can be much cheaper. Long contexts can trigger premium rates. Batch processing can reduce price where latency is less important. Tool calls and external grounding can create separate charges. FinOps guidance therefore warns against comparing providers or models simply on a headline “price per token,” since different workloads can have materially different cost compositions. (finops.org)

The execution model also provides benefits to customers. A start-up can access frontier models without acquiring GPUs or making long-term infrastructure commitments. A business can scale consumption with workload demand. Teams can choose cheaper models for simple tasks and more expensive models for harder problems. Caching, batching, model routing, and context optimization create opportunities to lower costs through engineering. Usage pricing therefore supports a highly flexible market and has played an important role in the speed at which generative AI has diffused.

The potential misalignment begins only when consumption becomes an unreliable proxy for economically useful work. This distinction is critical because the same billing model can be highly efficient in one application and poorly aligned in another. A classification workflow with consistent inputs and well-measured accuracy may have predictable token consumption and a clear cost per transaction. An open-ended research agent may autonomously retrieve documents, summarize them, call search tools, revise hypotheses, expand context, and run verification loops before completing a task. Both can be billed through tokens and calls, but the second has substantially greater variance between the technical unit consumed and the business result obtained.

Execution-based pricing is therefore best understood as cost-reflective infrastructure pricing, not outcome pricing. That characterization is not criticism; it clarifies what the pricing model does and does not promise. It meters use of the provider's resource. It does not by itself establish the economic productivity of that use.

3. Generative AI introduces a new distinction between the cost of producing an answer and the cost of resolving uncertainty

Traditional software failures frequently create visible error states. A database transaction fails, a web server returns an error, a calculation violates a validation rule, or a program crashes. Generative-AI failure is often less discrete. An answer can be fluent but wrong, partially correct but incomplete, logically plausible but unsupported, or appropriate in general but unsuitable for the customer's specific context. The system may also correctly identify that uncertainty exists and invoke additional reasoning or retrieval. As a result, the cost of producing an output and the cost of determining whether the output is good enough are separate economic categories.

This distinction becomes particularly important in professional and high-consequence work. AI-generated software must often be compiled, tested, scanned, reviewed, and integrated. Legal output may require professional verification. Financial or strategic analysis may require evidence review. Customer-facing content may require policy or brand checks. Healthcare applications may demand clinical validation. Autonomous enterprise actions may require authorization and audit trails. In these settings, inference cost can represent only a small part of total cost of ownership. The customer pays not just for generation but for the assurance infrastructure required to convert probabilistic output into operationally acceptable work.

NIST's Generative AI Profile reinforces this point from a risk perspective. The framework treats generative-AI risk management as a lifecycle activity involving design, development, deployment, evaluation, and use rather than a narrow model-selection problem. Its broader AI RMF organizes governance around Govern, Map, Measure, and Manage, explicitly recognizing that trustworthiness depends on organizational processes surrounding the model. Although NIST is not a pricing framework, its architecture implies an important economic conclusion: the cost of trustworthy AI includes the cost of measurement, evaluation, monitoring, and controls required to manage model risk. (NIST)

This creates a new denominator for enterprise economics. Cost per inference is useful for infrastructure engineering; cost per accepted outcome is closer to the measure required for business evaluation. FinOps already recognizes cost per inference as a core KPI and increasingly advocates use-case economics rather than token counting alone. The next logical step is to separate attempts from successful resolution. For a customer-service agent, the economically relevant outcome may be a resolved ticket without inappropriate escalation. For a coding agent, it may be a merged change that passes testing. For a research workflow, it may be an accepted deliverable with verified sources. For an enterprise process, it may be completion without human remediation. (finops.org)

The distinction does not imply that vendors should automatically refund every unsuccessful attempt. The provider cannot control all factors influencing outcome quality, including prompts, proprietary customer data, system integration, task ambiguity, and deployment context. It does imply that buyers who evaluate providers only on token price may optimize the wrong layer of the cost structure. A nominally more expensive model can be economically superior if it completes the task with fewer iterations, shorter context, less validation, and less human intervention. Conversely, a low-cost model can become expensive when it creates significant downstream remediation.

4. Failure-related consumption creates a structural risk-transfer mechanism even without any vendor intent to monetize failure

The most controversial version of the thesis is that AI providers profit from failure because failed responses generate retries. Public evidence does not support treating that as a generalized claim of deliberate vendor strategy. The economics are more nuanced. Additional model calls can create additional revenue, but they also consume infrastructure. Persistent failure reduces customer trust, increases competitive vulnerability, and can reduce long-term workload share. Model providers invest heavily in accuracy, efficiency, latency, reasoning capability, and tool performance precisely because customers value successful outcomes. A claim that vendors deliberately optimize for avoidable failure would therefore require strong evidence concerning internal objectives or product decisions that is not available from pricing structures alone.

The structural mechanism does not depend on that claim. Consider two AI systems solving the same task. One produces an accepted result on the first execution. The other requires three attempts, a retrieval step, and an evaluator model before reaching the same outcome. If all execution is billable, the second workflow generates more consumption despite delivering no more final value. The vendor may strongly prefer to improve the second system; the customer nevertheless bears the immediate cost variance while that improvement remains incomplete. The economic misalignment is therefore local rather than necessarily strategic: marginal non-performance can produce marginal customer cost even when the provider's long-run incentive is to improve performance.

This is what makes “risk transfer” a more precise concept than “failure monetization.” The relevant question is not whether the supplier wants failure. It is whether the economic consequences of uncertainty fall disproportionately on the customer relative to the customer's ability to control their source. Providers control model training, serving architecture, context mechanics, tool design, and substantial parts of agent orchestration. Customers control prompts, workflows, proprietary data, integration choices, and deployment context. The efficient allocation of risk should therefore vary by failure type rather than assigning all uncertainty to one party.

The same logic is familiar in other markets. Cloud providers charge for compute that runs inefficient customer code, because the customer controls the application. Professional firms may bill time even when the final recommendation is not adopted. Payment processors charge for transaction attempts even when a downstream business objective is not achieved. Generative AI is not unique in pricing inputs rather than outcomes. Its distinctive feature is that the supplier-provided cognitive component can itself create variable amounts of additional input consumption while attempting to achieve the same user objective.

This suggests that the market will increasingly need to distinguish customer-induced consumption from system-induced consumption. The distinction will never be perfect, but it can become economically useful. An agent that invokes a tool because the task objectively requires it is different from one that repeats the same failing action ten times. A user who requests five alternative creative concepts is intentionally purchasing exploration; a coding agent that repeatedly breaks and repairs its own changes is consuming remediation. Both may remain billable, but enterprise buyers will want visibility into the difference.

5. Agentic AI increases the problem because autonomous systems can determine their own consumption path

The shift from chat interfaces to agents materially changes the financial-control problem. In a standard conversational workflow, a human typically observes one response before deciding whether to continue. The user therefore acts as a natural pacing mechanism. In an agentic workflow, the AI system can decompose the problem, retrieve additional context, call another model, search the web, execute code, use a computer interface, evaluate an intermediate answer, retry a failed action, or delegate to another agent without waiting for a human decision after each step.

Google's current Gemini pricing documentation makes this dynamic explicit. Its agent pricing notes that model inference is charged at standard rates, including input, output, intermediate input and reasoning tokens generated during agentic loops, while tool fees can apply according to their own pricing structures. The significance is not that Google is unusual; it is that a leading provider is explicitly pricing the internal execution path of an autonomous workflow, not only the final user-facing answer. (Google AI for Developers)

Enterprise interest in this model is already substantial. McKinsey's 2025 survey found that 23 percent of respondents said their organizations were scaling at least one agentic AI system, while another 39 percent were experimenting. Agent adoption remained concentrated in a limited number of functions, with IT and knowledge management among the more common areas, but the survey makes clear that autonomous execution is moving into the mainstream enterprise experimentation portfolio. (McKinsey & Company)

The financial consequence is that workload demand becomes partially endogenous. Traditional infrastructure consumption is usually driven by external transactions, users, or scheduled workloads. An agent can create additional workload internally because its policy determines how many reasoning and tool steps to take. The FinOps Foundation describes agentic and multi-agent workloads as a key source of nonlinear token growth and notes that poorly bounded agents can create rapidly compounding cost. Its 2026 guidance on AI tooling similarly warns that an agent that loops excessively can generate significant financial exposure in a short period. (finops.org)

This creates a new management requirement: authorization needs an economic dimension. Giving an agent permission to act without specifying a maximum resource envelope is analogous to providing purchasing authority without a spending limit. Rate limits are helpful but are not equivalent to task budgets. Monitoring is helpful but is not equivalent to deterministic containment. A system can remain within a provider's technical rate limit while still consuming far more than the enterprise intended for one business objective.

The most mature agentic architectures are therefore likely to treat cost as part of the execution policy itself. Expensive model escalation, large context expansion, external search, paid tools, or branching into subordinate agents should increasingly be governed by task importance and remaining budget. The relevant design pattern is not “agent can act” but agent can act within a declared economic envelope and must escalate before exceeding it.

6. Average cost will become less useful than cost distributions as autonomous workflows expand

One of the most important consequences of agentic execution is that average cost can become an incomplete risk measure. A workflow that usually completes for $0.05 but occasionally consumes $50 because of long context, repeated searches, or a retry loop may look inexpensive under simple averages while creating unacceptable exposure at high volume. Enterprise finance therefore needs to understand not only the mean but the variance and tail of cost per completed task.

The issue is familiar in reliability engineering and finance but relatively new in AI procurement. Public model-price pages emphasize deterministic rates per unit: price per million tokens, per image, per search request, or per processing tier. Those prices are transparent at the unit level. What remains uncertain is the number and mix of units consumed by an open-ended task. Google's current pricing is an instructive example: standard inference, priority inference, context caching, storage, Google Search grounding, Google Maps grounding, and agentic loops can each contribute to the final cost structure. The individual rates are published, but task-level spend is an emergent result of model behavior and application design. (Google AI for Developers)

OpenAI and Anthropic expose a similar pattern through differentiated input/output pricing, caching, batch discounts, long-context premiums, and model tiers. OpenAI notes that GPT-5.4 prompts above a defined context threshold receive higher pricing across the session, while Anthropic applies premium pricing for long-context Sonnet requests above its stated threshold. These mechanisms are rational reflections of infrastructure cost but illustrate why model selection and prompt length alone are insufficient for forecasting autonomous workloads: once an agent decides to accumulate additional context, the economic regime itself can change. (OpenAI Developers)

FinOps guidance increasingly calls this problem “context window creep” and warns that apparently small application changes can significantly increase total token consumption as conversation history, retrieved documents, tool results, system prompts, and multimodal data accumulate. It recommends evaluating use-case economics rather than selecting models primarily on headline price. (finops.org)

For finance teams, the relevant operating metric should therefore evolve from unit rate × expected volume toward a probability distribution of task-level expenditure. High-volume deterministic use cases may remain straightforward. Research, coding, automated operations, and other agentic tasks require stronger tail-risk analysis. In practical terms, enterprises need to know what a normal task costs, what a 95th- or 99th-percentile task costs, what conditions cause the tail, and whether the system can prevent an abnormal branch from exceeding a hard threshold. That is a materially different discipline from negotiating a lower token price.

7. The cost of validation may become as important as the cost of inference

The economics of AI are often discussed as though model usage were the central cost category. For some workloads that is true. For others, the larger expense lies in the assurance system around the model. Generated software requires testing and review; high-stakes analysis requires evidence verification; automated actions require authorization and logging; customer-service agents require quality monitoring; regulated workflows require compliance; and AI systems may require security controls, evaluation data, red-team testing, and incident-management capability.

AWS's Bedrock pricing illustrates the point indirectly. Model evaluation itself consumes inference, and human-based evaluation can incur an additional fee per completed human task. RAG evaluation or LLM-as-a-judge also generates model usage charged at standard rates. Evaluation is therefore not simply an abstract governance activity; it can be another computational workload with direct marginal cost. (Amazon Web Services, Inc.)

The economic consequence is important because quality improvements can create second-order savings beyond cheaper inference. A more reliable model can reduce human review, decrease retry frequency, shorten context, lower evaluator use, and reduce downstream remediation. Conversely, a superficially cheap model can be expensive once assurance is included. This is why direct model-price comparisons frequently fail to capture total cost of ownership.

The same issue applies to internal AI platforms. Organizations increasingly build gateways, model routers, prompt registries, observability systems, test suites, policy controls, cost monitors, and evaluation harnesses around external models. These layers can be strategically valuable because they allow enterprises to govern models consistently and switch providers. They also create cost and organizational complexity. When assessing the economics of a particular AI use case, finance teams should therefore distinguish model cost, orchestration cost, tool cost, validation cost, human-review cost, compliance cost, and failure-remediation cost rather than treating the API bill as total AI expenditure.

FinOps research is moving in this direction. Its AI framework explicitly distinguishes cost per inference, token consumption, hardware utilization, API cost, and other categories while emphasizing the need to quantify business value. Its practitioner guidance notes that cost visibility and allocation remain foundational challenges. (finops.org)

This reframes the purchasing question. The economically superior AI model is not necessarily the one with the lowest unit rate. It is the one that minimizes risk-adjusted cost per acceptable outcome within the quality, latency, security, and governance requirements of the use case.

8. Current pricing transfers some uncertainty downstream, but providers also retain substantial economic risk

A balanced analysis requires recognizing that AI providers are not insulated from the consequences of poor performance. Model development is capital intensive. Inference requires significant infrastructure. Providers compete on benchmark performance, latency, reliability, developer experience, and price. Additional token consumption is not pure margin because serving additional tokens consumes real resources. Poor quality can lead customers to switch providers, use cheaper models, self-host open-weight alternatives, or redesign applications to reduce dependency on a vendor.

The current market also demonstrates active price competition. OpenAI, Anthropic, Google, AWS, and other providers offer multiple price/performance tiers, caching, batching, discounted processing modes, and model-routing options. OpenAI's launch materials for GPT-5.4 explicitly state that although its per-token rate is higher than GPT-5.2, greater token efficiency can reduce the total number of tokens required for many tasks. That commercial framing itself demonstrates that vendors recognize task-level token efficiency as part of customer value, not only unit price. (OpenAI)

Anthropic similarly provides model tiers and caching mechanisms designed to reduce cost for repeated context, while Google offers batch and flex pricing below standard rates and multiple model families for different workload requirements. AWS enables customers to select among competing model providers within Bedrock. These mechanisms create pressure on providers to improve the cost-quality frontier rather than simply maximize gross token consumption. (Claude Platform Docs)

The correct interpretation is therefore a partial risk transfer, not complete vendor insulation. Vendors retain technology risk, capital-expenditure risk, infrastructure-utilization risk, competitive risk, reputation risk, and some contractual service risk. Customers retain application risk, integration risk, validation risk, context-specific liability, and much of the uncertainty concerning whether generated output is useful in their environment.

The strategic question is whether the current allocation is efficient as AI systems become more autonomous. During an API's early adoption phase, usage pricing provides simplicity and flexibility. At enterprise scale, buyers may increasingly demand contracts that differentiate infrastructure consumption from quality obligations. The emergence of such arrangements would not indicate that usage pricing “failed”; it would indicate that the product category matured enough to support finer-grained risk allocation.

9. Cloud FinOps is the closest historical analogy, but agentic AI adds an important new layer

The public-cloud market provides the strongest historical analogy. Cloud replaced fixed-capacity infrastructure with variable consumption. Customers gained elasticity, speed, and lower entry barriers but also acquired the risk of uncontrolled consumption, weak cost attribution, architectural inefficiency, and large unexpected bills. FinOps emerged to connect engineering, finance, and business teams around cloud unit economics and accountability.

Generative AI has already begun to produce an analogous discipline. The FinOps Foundation now treats AI as a distinct scope because of its cost complexity, fast development cycles, spend unpredictability, and need to align consumption with business value. Its 2025 State of FinOps research found broad planned AI investment across multiple infrastructure categories, and its 2026 activities included the announcement of an intended Tokenomics Foundation focused on open practices and standards around AI billing. (finops.org)

The analogy should not be pushed too far. Traditional cloud workloads are variable, but the application's logic is generally deterministic. A server does not decide autonomously to deploy five more servers because it is uncertain whether its previous computation was correct unless the application developer explicitly programmed such behavior. In agentic AI, the policy controlling additional consumption may itself be probabilistic. A model can decide to retrieve more information, retry, call a stronger model, generate a longer reasoning trace, or invoke another agent.

This creates a new form of cost endogeneity. FinOps traditionally asks whether engineers provisioned and used infrastructure efficiently. AI FinOps must also ask whether the cognitive strategy chosen by the system was economically efficient. A workflow that repeatedly invokes a premium model where a cheaper model would suffice is not just infrastructure waste; it is reasoning-policy waste. A research agent that continues searching after sufficient evidence has been obtained creates cost through stopping-policy design. A customer-service agent that escalates to expensive reasoning unnecessarily creates cost through routing policy.

The maturation of AI economics will therefore require closer integration between model evaluation and cost engineering. Quality, latency, and expenditure cannot be optimized independently. The relevant objective is likely to become minimum sufficient compute for the required quality and risk level, rather than either maximum model capability or minimum token cost.

10. Enterprise procurement will shift from model-price comparison to use-case unit economics

Early model procurement often resembles commodity comparison: price per million tokens, context length, throughput, benchmark scores, and latency. Those measures remain relevant but become increasingly insufficient as enterprises deploy complex systems. A procurement team evaluating an AI coding platform, for example, ultimately cares about engineering output, quality, security, developer productivity, and total cost. Whether the underlying system used 50,000 or 500,000 tokens matters only because it affects those outcomes.

The FinOps Foundation increasingly advocates precisely this transition. Its recent guidance states that the “real unit of measure” is the use case and argues that the advertised per-token price can be misleading when context, caching, modality, routing, and provider architecture differ. (finops.org)

For buyers, this implies a hierarchy of economic metrics. Token and call metrics remain useful for engineering optimization. Above them sit workflow measures such as model attempts, tool calls, retries, context expansion, evaluator calls, and human interventions. Above those sit outcome measures such as successfully resolved tickets, accepted code changes, validated reports, completed transactions, or analyst hours saved. Procurement should increasingly connect all three levels.

The importance of this approach is demonstrated by the gap between AI adoption and EBIT impact in McKinsey's 2025 survey. Organizations are experimenting broadly, including with agents, but enterprise value remains concentrated among a smaller set of organizations that have redesigned workflows and embedded AI more deeply. The implication is not that pricing explains the value gap; workflow redesign, organizational change, data, skills, and governance are major factors. It does indicate that technical consumption metrics alone provide very little evidence of business impact. (McKinsey & Company)

As the market matures, enterprise RFPs are therefore likely to evolve. Instead of asking only for unit prices and model benchmarks, sophisticated buyers will seek evidence about task completion, quality-adjusted cost, tail spend, routing efficiency, failure handling, observability, spend controls, enterprise SLAs, and contract treatment of significant service deficiencies.

The more AI becomes an operating expense rather than a research budget, the more procurement will ask a conventional question in a new form: what economically useful output does this contract buy?

11. Outcome-based pricing is attractive conceptually but difficult to implement as a universal model

If usage-based pricing creates risk transfer, outcome-based pricing appears to provide an obvious solution: charge only when the customer receives value. In practice, this model is difficult to generalize.

The first challenge is outcome definition. A customer-service case can sometimes be classified as resolved, but even this can be ambiguous if the customer later reopens the issue. A software change can be merged but later create a defect. A marketing asset can be approved without producing sales. A strategic analysis can influence a decision without having a directly measurable independent value. Generative AI operates across many tasks where quality is continuous and context-dependent rather than binary.

The second problem is causal attribution. The model may generate an excellent output that the customer deploys poorly. Conversely, business performance may improve for reasons unrelated to the AI. Suppliers cannot reasonably guarantee outcomes determined substantially by customer operations or external market conditions.

The third is adverse selection and gaming. If customers pay only for accepted outputs, they may set strategically strict acceptance criteria. If vendors are paid for a specific performance metric, they may optimize the metric rather than the underlying business objective. Outcome pricing moves the measurement problem rather than eliminating it.

The fourth is supplier cost reality. Model providers incur inference expenses whether or not the customer's subjective evaluation is positive. A pure outcome model could make pricing uneconomic for open-ended workloads unless prices include a significant risk premium.

These constraints explain why hybrid structures are more plausible. A provider may continue charging for compute while offering service credits for defined failures, performance commitments for narrow workflows, fixed-price task packages, enterprise capacity contracts, or shared-savings models in domains with measurable outcomes. The appropriate model will depend on how much control each party has over the result.

The strategic point is therefore not that outcome pricing should replace execution pricing. It is that enterprise AI economics will become progressively more differentiated according to measurability and controllability of outcomes.

12. The highest-value near-term control is not a new pricing model but stronger bounded execution

Because outcome attribution is difficult, the fastest path to better risk allocation is technical and contractual containment. Autonomous systems should increasingly operate with predefined economic limits.

The principle is straightforward. An enterprise should know the maximum cost that a single authorized task can generate before the task begins, or at least be able to set such a limit. The implementation can vary: maximum tokens, maximum tool spend, maximum model calls, maximum elapsed compute, maximum branching depth, model-tier restrictions, or a total currency-denominated budget.

Rate limits alone do not provide this control. They regulate consumption speed, not total task expenditure. Provider quotas are useful for infrastructure stability, but an agent can remain below a per-minute quota while consuming far more total resources than the customer intended. AWS, for example, documents token-per-minute, request-per-minute, and in some cases daily quotas for Bedrock, but these are service-level capacity controls rather than outcome-specific financial authorizations. (AWS Documentation)

The same distinction applies to monitoring. AWS provides detailed model-invocation logging, including input and output token counts, model IDs, principals, and request metadata. Those capabilities are important for attribution and post hoc analysis. But logging identifies what happened; it does not necessarily prevent a workflow from exceeding its intended budget. (AWS Documentation)

Enterprise-grade agent governance therefore needs both observability and enforcement. The system should know how much has been consumed, but it should also stop or escalate when defined limits are reached. High-value workflows may use progressive authorization: an agent can spend a small amount autonomously, request additional budget when evidence suggests further computation is valuable, and require human approval before crossing a consequential threshold.

This approach reduces the severity of risk transfer even when usage pricing remains unchanged. The customer continues paying for execution but regains control over the maximum downside. In financial terms, bounded execution converts an open-ended exposure into a capped one.

13. Retry economics should become a first-class management metric

Retries are technically normal in probabilistic systems, but their economic role is under-measured. A retry can represent legitimate exploration, robustness, or remediation. Those categories should not be treated as equivalent.

For creative work, multiple samples may be the intended product. A designer requesting eight visual concepts is not experiencing seven failures; the portfolio of alternatives creates value. In code generation, however, repeated attempts to repair defects introduced by previous attempts may represent quality cost. In an agentic operational workflow, retrying the same failed tool call may be pure waste. In research, additional search can be valuable until the marginal evidence no longer changes the conclusion.

The useful enterprise metric is therefore not a simplistic “retry rate.” It is retry purpose and yield. Organizations should understand what proportion of total AI spend is associated with deliberate exploration, validation, recovery from external failures, and remediation of model-generated problems.

This distinction can influence product design. A system that detects repeated failure should not necessarily continue attempting indefinitely. It may switch models, change strategy, ask for clarification, escalate to a human, or terminate. The economically optimal stopping policy will depend on the value of the task, probability that another attempt succeeds, cost of the attempt, and consequence of failure.

The source thesis behind this report is strongest at precisely this point. If a system repeatedly consumes billable resources while attempting to resolve problems created by its own prior executions, then the customer's expenditure contains a failure-remediation component that should be visible even if it remains contractually billable. Hiding that component inside aggregate token usage makes it difficult for customers to determine whether rising consumption reflects adoption or poor convergence.

The industry is likely to develop a vocabulary around this distinction as AI FinOps matures. Cloud practitioners eventually learned to distinguish productive capacity from idle or overprovisioned resources. AI practitioners will need to distinguish productive reasoning from redundant, failed, or excessively repeated reasoning.

14. Better model quality can reduce cost even when list price rises

One of the most important implications for model selection is that per-token price and total task cost can move in opposite directions. A more expensive model may require fewer tokens, fewer iterations, less retrieval, less human review, or fewer evaluator calls. OpenAI explicitly makes this argument in its GPT-5.4 launch material, noting that the model is priced above GPT-5.2 per token but can reduce total token use for many tasks through greater token efficiency. (OpenAI)

The broader principle is familiar in industrial economics: unit input cost should not be optimized independently of productivity. A cheaper worker who takes twice as long is not automatically less expensive per unit of output. A lower-priced AI model is similarly not necessarily the economically optimal choice if additional execution is required to achieve the required quality.

This creates a strong case for model routing rather than one-model standardization. Simple classification, extraction, rewriting, or summarization may be handled by smaller models. High-consequence reasoning may justify a more capable model. Difficult cases can escalate dynamically. The economically efficient architecture assigns compute according to task difficulty and value rather than applying premium inference universally.

The risk is that dynamic routing itself becomes opaque. If an agent can escalate to expensive models autonomously, cost governance requires visibility into when and why escalation occurred. The optimal architecture therefore combines model routing with policy boundaries and outcome measurement.

For vendors, improved task efficiency can become a competitive differentiator even when list prices appear higher. For customers, benchmarking should increasingly include fixed-task comparisons: total tokens, attempts, elapsed time, human corrections, and cost required to reach an acceptance threshold. That approach is much closer to the economic question the buyer actually needs to answer.

15. Liability and contractual allocation will matter more as AI moves into consequential workflows

Pricing is only one channel through which AI risk is allocated. Contracts determine another. Model providers typically define service availability, acceptable use, data handling, intellectual-property provisions, and limitations of liability. Customers, meanwhile, remain responsible for how model outputs are incorporated into downstream products and decisions.

As AI moves into increasingly consequential settings, this allocation will face pressure. A consumer using a general-purpose assistant and an enterprise delegating operational authority to an agent present fundamentally different risk profiles. The latter may involve financial transactions, system changes, customer communications, regulated decisions, or interaction with third-party services.

The key distinction is control over the failure mechanism. If a customer configures an agent poorly, uses low-quality proprietary data, or deploys the model outside stated conditions, the customer may reasonably bear substantial responsibility. If the provider's system materially fails against a defined enterprise service commitment, the case for provider-side remediation is stronger.

The current market is therefore likely to segment. Commodity access to general-purpose models may remain predominantly execution-based with limited outcome responsibility. Higher-value managed solutions may include stronger performance commitments, workflow-level guarantees, enterprise support, or shared responsibility for defined failure classes.

This is analogous to the maturation of other technology markets. Infrastructure providers do not guarantee the customer's business outcome, but they do accept responsibility for increasingly well-defined properties of the service they control. AI contracts are likely to evolve similarly as model behavior becomes more measurable and buyers gain experience distinguishing controllable from context-specific uncertainty.

The relevant governance principle is not to transfer all risk upstream. It is to allocate each risk to the party with the greatest ability to observe, prevent, mitigate, or insure it.

16. Risk-transfer concerns are strongest where autonomy, opacity, and consequence are all high

Not every AI application warrants equal scrutiny. The risk-transfer mechanism is modest where individual calls are cheap, a human controls every invocation, outcomes are easily checked, and failures have minimal consequence. It becomes significantly more important when three factors increase simultaneously.

The first is autonomy. Systems capable of initiating their own actions and consumption reduce the frequency of human checkpoints.

The second is economic opacity. Long context, multiple models, paid tools, retrieval, evaluators, and agent loops can make task-level spend difficult to predict from headline prices.

The third is consequence. Where an incorrect or uncontrolled output creates legal, financial, security, reputational, or operational cost, the gap between inference price and total risk becomes larger.

These factors suggest a differentiated governance model. A low-stakes summarization assistant may require simple budget monitoring. An autonomous coding system with repository write access requires stronger execution and rollback controls. A financial or operational agent able to initiate paid external actions requires hard authorization and economic limits.

The same logic applies commercially. Enterprises should not spend as much governance effort on a low-cost brainstorming assistant as on an agent capable of initiating thousands of dollars in services or making consequential changes to production systems. Risk management should be proportionate to expected downside.

This perspective is consistent with NIST's risk-based approach. The AI RMF is explicitly designed to adapt to use case and context rather than impose one universal control model. (NIST)

17. AI FinOps is likely to evolve into a broader discipline of “intelligence unit economics”

The rise of cloud FinOps created a shared language for engineering and finance around variable infrastructure. Generative AI now requires an analogous discipline, but one that incorporates model behavior and business outcomes as well as resource consumption.

The FinOps Foundation already frames AI cost management around allocation, forecasting, anomaly detection, unit economics, and value. Its 2026 materials increasingly emphasize token economics and the need to move beyond raw token counts. One recent practitioner resource argues explicitly that a token is a unit of computation rather than ownership or value and identifies agentic workloads as a major driver of unpredictable cost expansion. (finops.org)

The next stage is likely to integrate three layers of measurement. Infrastructure economics will track tokens, model calls, cache use, tool charges, GPUs, and platform costs. Cognitive-workflow economics will track attempts, context growth, routing, retries, verification, escalation, and human intervention. Business economics will track resolved cases, accepted outputs, productivity, revenue, error reduction, cycle time, and risk outcomes.

The strategic advantage of this architecture is that it prevents optimization at one layer from degrading another. A team can reduce token use but harm quality. It can increase model quality but make the workflow economically unscalable. It can automate large volumes while increasing downstream correction. A full unit-economics system shows these trade-offs.

This will become particularly important as AI providers package more functionality into subscriptions where the underlying token meter is obscured. FinOps has already identified opaque SaaS-model token economics as a significant emerging management challenge. (finops.org)

Enterprise buyers therefore need a stable internal value unit even when provider billing units differ. The most useful measure will vary by function, but the principle is consistent: normalize AI expenditure against a business outcome the organization actually values.

18. The likely market endpoint is hybrid pricing with stronger execution governance, not the disappearance of token pricing

The current token model is deeply embedded and economically rational enough that it is unlikely to disappear. The more probable evolution is layering.

Commodity model APIs will continue to price input, output, caching, search, and tool usage. Enterprise customers will negotiate volume discounts and commitments. Application vendors will bundle model cost into per-seat or consumption packages. Agent platforms will add budgets, quotas, and governance. High-value workflow products will increasingly experiment with task, resolution, or outcome pricing where the result is measurable.

The market will also increasingly differentiate provider pricing from enterprise accounting. A company can pay a vendor per token while internally allocating cost per resolved customer case. There is no requirement that the billing unit and management unit be identical. Cloud providers bill per compute resource while enterprises calculate cost per order or customer; AI will follow the same pattern.

What is different is that better internal accounting alone cannot solve uncontrolled autonomous execution. The technical layer must also provide reliable spend controls. The strongest enterprise products will therefore combine transparent metering with bounded execution and high-quality telemetry.

Vendors that help customers answer “what did this AI spending accomplish?” may gain commercial advantage even if their unit rates are not the lowest. This is particularly likely as the market moves from experimentation toward CFO scrutiny. Gartner's 2025 generative-AI forecast explicitly noted that enterprises were becoming more skeptical of early proof-of-concept results and were seeking more predictable implementation and business value. (Gartner)

The commercial shift is therefore from cheap intelligence toward predictable intelligence economics.

19. Strategic implications for AI vendors

For foundation-model providers, the risk-transfer thesis should not be interpreted solely as a threat to usage pricing. It identifies an opportunity to differentiate on economic reliability. Providers can help customers reduce uncertainty by exposing richer task-level usage, improving cost forecasting, enabling budgets and stop conditions, reducing unnecessary token generation, making model routing transparent, and publishing performance measures tied to real workloads rather than benchmark accuracy alone.

There is also a strategic case for selective risk sharing. Enterprise customers may be willing to pay premium rates for stronger assurances on defined workloads. Service credits, fixed-price task packages, managed agents, quality guarantees, or contractual remediation can convert provider confidence into customer trust. These models will not apply to all tasks, but they can be powerful where outcomes are sufficiently measurable.

Vendors also benefit from reducing failure-related consumption where the long-term effect is stronger adoption. An agent that completes a business process reliably with fewer steps may generate fewer tokens per task but increase the number of tasks customers are willing to delegate. The economically relevant optimization for the provider is therefore not necessarily tokens per task; it may be total trusted workload share.

This suggests that vendor economics and customer economics can converge even under usage pricing. The provider wins when lower cost per successful outcome expands demand enough to compensate for fewer tokens per individual task. The cloud market has repeatedly demonstrated similar effects as improvements in price/performance stimulated larger workloads.

20. Strategic implications for enterprise CFOs, CIOs, and procurement leaders

For enterprise buyers, the most important shift is to stop treating AI as a simple software line item. It is a variable-cost production input with quality variance.

CFOs should demand cost attribution at the use-case level and distinguish model expenditure from total AI operating cost. CIOs should ensure that observability captures model calls, context growth, tool invocation, retries, and routing. Procurement leaders should evaluate not only unit price but cost distributions, contract structure, enterprise support, and economic controls.

The governance question becomes particularly important for agents. Every autonomous workflow should have a defined economic owner, maximum exposure, escalation path, and acceptance measure. If a company cannot identify who owns the budget of an autonomous agent, that is itself a governance gap.

The relevant management discipline resembles credit authorization. Enterprises do not normally give employees unlimited purchasing authority simply because each individual purchase is permitted. They define thresholds and escalation. AI agents capable of consuming paid services require an equivalent structure.

The objective is not to eliminate experimentation. Hard controls can coexist with rapid innovation. In fact, bounded downside can make experimentation easier to authorize because management knows the maximum exposure.

21. Outlook to 2030: economic control is likely to become a prerequisite for large-scale autonomy

The direction of travel is clear even if the exact pace is uncertain. McKinsey's 2025 survey already shows broad experimentation with agents. Gartner and IDC both expect substantial growth in AI expenditure. Model providers are expanding context windows, reasoning capabilities, search, tool use, computer interaction, and multimodal functionality. Each of these capabilities expands the range of tasks that can be delegated.

The constraint is increasingly shifting from “can the model do the task?” toward “can the organization trust the economics and consequences of letting the model do the task repeatedly without direct supervision?”

That is a higher bar.

A highly capable agent with unbounded execution cost is unsuitable for some production environments. A slightly less capable agent with deterministic budgets, reliable escalation, and predictable resolution economics may be much easier to deploy.

By the end of the decade, enterprise AI architectures are therefore likely to incorporate economic policy engines alongside security and authorization controls. Agents will know which models they are permitted to use, how much they can spend, which tools carry additional cost, when they must stop, and when additional expected value justifies requesting a larger budget.

Pricing and governance will increasingly converge.

Conclusion

The first generation of generative-AI economics was organized around a simple technical question: how much does it cost to run the model?

That question made sense when the model was primarily an API and the user controlled each interaction.

It becomes less sufficient as AI moves into autonomous workflows.

The customer does not ultimately want tokens, model calls, reasoning steps, or tool invocations. The customer wants a resolved case, a working software change, a reliable analysis, an approved asset, an automated business process, or another economically useful result.

The supplier, however, must recover the cost of the computation required to attempt that result.

That tension explains why execution pricing is both rational and incomplete.

The strongest evidence does not support the claim that major AI providers deliberately design failure to increase revenue. Competitive incentives, infrastructure costs, and customer retention all create significant pressure to improve reliability and efficiency. But intentional exploitation is not required for economic misalignment to exist.

The structural issue is simpler.

When the amount of billable execution required to produce a fixed customer outcome is uncertain, part of the performance risk moves downstream.

A failed first attempt can create a second paid attempt.

An uncertain answer can create a paid verification step.

A larger reasoning process can create additional output tokens.

An agent can autonomously decide that it needs another search, another model, another tool, or another branch.

A regulated deployment can then require downstream human verification and governance.

The customer's total economic exposure is consequently wider than the model's unit price.

This is where the concept of risk transfer becomes commercially useful. It directs management attention away from whether tokens are “fairly priced” and toward the more important question of whether control, information, and downside are allocated coherently between provider and customer.

In extreme cases—where autonomous systems can create substantial additional consumption while customers bear validation, remediation, and external liability and lack effective limits—the stronger language of risk laundering may be analytically appropriate. But that classification requires evidence at the product and contract level, not inference from usage pricing alone.

For most enterprises, the immediate management priority is therefore more practical.

Measure the economics of the completed business task rather than only the cost of model execution.

Separate productive consumption from failure-remediation consumption.

Track the distribution of task-level costs rather than only the mean.

Price human review and evaluation into total cost of ownership.

Require deterministic budget limits for autonomous execution.

Escalate expensive or anomalous branches.

And evaluate AI vendors on quality-adjusted outcome economics, not only token prices and capability benchmarks.

The longer-term opportunity is larger.

If providers and customers can align execution economics with delivered value, AI autonomy can scale with substantially less financial friction. Vendors can earn more because customers entrust more work to their systems. Enterprises can automate more because maximum downside becomes governable. Procurement can move from model shopping toward outcome economics. Finance can treat AI as a controllable production input rather than an opaque technical experiment.

The central economic challenge of the agentic era is therefore not that artificial intelligence consumes resources.

Every productive system does.

The challenge is that the system increasingly participates in deciding how many resources it will consume before the buyer knows whether the desired outcome has been achieved.

The companies that solve that problem—through better models, stronger observability, bounded execution, clearer unit economics, and more credible sharing of performance risk—will do more than improve AI pricing.

They will create the commercial infrastructure required for autonomous intelligence to become a trusted enterprise operating model.

References

  1. McKinsey & Company. The State of AI in Early 2024: Gen AI Adoption Spikes and Starts to Generate Value. May 2024. Survey of 1,363 participants; 65 percent reported regular organizational use of generative AI. (McKinsey & Company)

  2. McKinsey & Company. The State of AI in 2025: Agents, Innovation, and Transformation. November 2025. Reports nearly nine in ten respondents using AI, 62 percent experimenting with or scaling agents, 23 percent scaling agentic systems, and 39 percent reporting enterprise EBIT impact. (McKinsey & Company)

  3. Gartner. Worldwide AI Spending Forecast to Grow 47% in 2026. May 2026. Forecasts $2.59 trillion in worldwide AI spending for 2026. (Gartner)

  4. Gartner. Worldwide GenAI Spending to Reach $644 Billion in 2025. March 2025. Forecast covering generative-AI hardware, software, and services. (Gartner)

  5. Gartner. Worldwide End-User Spending on GenAI Models to Total $14.2 Billion in 2025. July 2025. (Gartner)

  6. OpenAI. GPT-5.4 Model and Pricing Documentation. Current pricing documentation accessed in 2026. (OpenAI Developers)

  7. Anthropic. Claude API Pricing. Model, caching, batch, and long-context pricing documentation. (Claude Platform Docs)

  8. Google. Gemini Developer API Pricing. Includes token, caching, grounding, service-tier, and agentic-loop pricing. (Google AI for Developers)

  9. Amazon Web Services. Amazon Bedrock Pricing. Includes model, token, caching, batch, provisioning, and associated service pricing. (Amazon Web Services, Inc.)

  10. Amazon Web Services. Understanding Amazon Bedrock Cost and Usage Report Data. Documents separate accounting for input, output, cache-read, and cache-write tokens and limitations of aggregated billing data. (AWS Documentation)

  11. Amazon Web Services. Monitor Model Invocation Using CloudWatch Logs and Amazon S3. Documents per-invocation model IDs, token counts, principals, and request metadata. (AWS Documentation)

  12. Amazon Web Services. Amazon Bedrock Model Evaluation Pricing. Documents inference charges and additional human-evaluation charges for evaluation workloads. (Amazon Web Services, Inc.)

  13. FinOps Foundation. State of FinOps Report 2025. Reports broad multi-infrastructure AI investment and strong early focus on understanding AI spend and quantifying business value. (data.finops.org)

  14. FinOps Foundation. FinOps for AI. Framework guidance on spend unpredictability, allocation, forecasting, token consumption, inference costs, and AI unit economics. (finops.org)

  15. FinOps Foundation. How to Build a Generative AI Cost and Usage Tracker. Guidance on token-based billing, attribution, and production-scale cost tracking. (finops.org)

  16. FinOps Foundation. GenAI FinOps: How Token Pricing Really Works. 2026. Discusses context-window growth, input/output pricing differences, provider variation, and use-case economics. (finops.org)

  17. FinOps Foundation. Token Economics: The Atomic Unit of AI Value. 2026. Discusses nonlinearity, agentic workloads, SaaS opacity, and AI consumption economics. (finops.org)

  18. FinOps Foundation. FinOps X 2026: Token Economics and the Evolving Role of FinOps. June 2026. Reports the growing institutional focus on token economics and AI spend standards. (finops.org)

  19. National Institute of Standards and Technology. Tabassi, E. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1, 2023. (NIST)

  20. Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P., and Roberts, K. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1, 2024; updated 2026. (NIST)