Skip to main content

AI Agents

Token and AI API budgets: the FinOps guide for SMEs in 2026

· 35 min read · Paul-Antoine Tual

AI Act FinOps AI ISO 42001 LLM Gateway SME tokens

Introduction: the new economic paradigm of artificial intelligence in the enterprise

The integration of artificial intelligence into business processes has crossed a tipping point. As of May 2026, the technology landscape of small and medium-sized enterprises (SMEs) is marked by a deep structural transition: the shift from a software economy based on fixed per-user licences (SaaS) to a utility-consumption economy, dictated by an omnipresent billing unit: the token. This pricing shift has introduced new volatility into financial and technology planning.

Current statistics reveal an instructive duality. On the one hand, adoption rates have risen sharply: 88% of organisations use artificial intelligence in at least one business function (Stanford AI Index 2026), a figure the McKinsey barometer had already revised from 78% to 88% in its 2025 iterations [1]. On the other hand, this ubiquity comes with a financial reality that must be faced head-on: 43% of large AI initiatives are judged doomed to fail, for want of execution (HCLTech, May 2026) [3]. An MIT study (Project NANDA, August 2025) estimated that 95% of generative AI pilots had no measurable impact on the P&L [2]; the 2026 measurements paint a less binary picture: 11% of “AI leaders” (KPMG Global AI Pulse Q1 2026) and fewer than 10% of organisations having fully scaled AI in a function (Stanford AI Index 2026). The main cause lies not in the limitations of language models (LLMs), but in the absence of an architectural and organisational framework for managing inference costs at scale. This is, fundamentally, good news: a problem of method is fixed by method.

Companies are now observing an internal phenomenon sometimes referred to as “tokenmaxxing”: the consumption of computing power by development and operations teams is sometimes wrongly interpreted as an indicator of technological velocity. The financial consequences are concrete. Some SMEs find that spending on AI tokens has become one of their fastest-growing budget lines, sometimes overtaking the cost of the automation tasks it replaces. It is not uncommon to see a cloud infrastructure bill rise sharply. One documented case [6] mentions an autonomous agent that reached its injection cap of 150,000 characters and accumulated several hundred dollars in monthly overspend on unsupervised flows. This is what is called a “shadow budget”: AI spending that escapes financial control.

In the absence of control frameworks, consumption grows asymmetrically relative to the value generated. With the proliferation of complex requests and the emergence of autonomous multi-agent systems, inference spending in engineering departments becomes a budget line in its own right. Several field reports put it at close to 10% of staff costs on user teams, although no reference institute (IDC, Gartner) has to date validated this ratio as a consolidated average. Optimising AI costs is therefore no longer a mere matter of financial hygiene relegated to FinOps teams at the end of the quarter; it constitutes an architectural discipline in its own right.

This document sets the gold standard for optimising, distributing and governing token and API budgets for teams operating within SMEs in 2026. It details the structuring of quotas, routing gateway architectures, semantic caching strategies, the secure management of autonomous agents and frameworks for evaluating return on investment (ROI). The objective is simple: to provide a technical and financial foundation that turns a structurally inflationary technology into a lever for predictable, measurable profitability.

The regulatory framework and governance: foundations of profitability

Optimising technology budgets in 2026 is intrinsically tied to a company’s ability to impose clear governance. The absolute freedom to experiment of previous years has given way to a regulated environment, where compliance guides the architecture of information systems. SMEs can no longer let each department deploy artificial intelligence models on an ad hoc basis, without centralised oversight.

The ISO/IEC 42001 standard: structuring accountability

The international standard ISO/IEC 42001:2023, dedicated to artificial intelligence management systems (AIMS), has established itself as the reference framework for structuring a responsible and financially viable use of AI [7]. Obtaining this certification (or, at a minimum, aligning rigorously with its requirements) is not a communications exercise: it is a prerequisite for budgetary control.

One of the standard’s major contributions is the obligation to maintain a complete and up-to-date inventory of all artificial intelligence systems, deployed models and third-party providers used by the organisation. Without this visibility, it is impossible to attribute token consumption costs to the various profit centres. The standard requires that risk and impact assessment be carried out at the level of each specific application, and not generically at company level. This leads SMEs to link each API request flow to a designated owner, creating a direct line between the technology spend (the cost of tokens) and managerial responsibility (accountability).

Adopting ISO 42001 also helps to fill a notable decision-making gap. On one side, the Piper Sandler CIO Survey reports that 87% of CIOs expect an increase in their AI budget [4]. On the other, the Drexel LeBow / RGP work shows that only 14% of leaders say their organisation is prepared in terms of skills, and that 14% of CFOs measure a clear impact on the P&L [5]. These two studies do not overlap exactly, but their convergence points to the same reality: AI budgets are rising faster than governance maturity. Deploying an AIMS framework in line with ISO 42001 leads management committees to take ownership of consumption metrics, and turns technology spend into an auditable strategic asset.

European regulation (AI Act) and initiatives for frugal AI

On the regulatory front, the AI Act timetable has just changed. The “Digital Omnibus” political agreement reached at the European trilogue on 7 May 2026 has postponed the entry into force of binding obligations for high-risk AI systems: 2 December 2027 for standalone systems (Annex III: recruitment, credit scoring, biometrics) and 2 August 2028 for systems embedded in already-regulated products (Annex I: medical devices, industrial machinery) [8]. The transparency obligation remains set for 2 August 2026, with one exception: the machine-readable marking of generative content (Article 50(2)) benefits from a reprieve until 2 December 2026 (Digital Omnibus).

For French SMEs, this reprieve is not an invitation to stall: it is a useful window in which to structure governance (system inventory, per-application risk assessment, documented human oversight) before these requirements become binding. The penalties remain heavy on the horizon; and the real cost of being unprepared is paid first in emergency reorganisation, not in fines.

Alongside legal compliance, the concept of “frugal AI” has taken concrete form in France through practical reference frameworks. AFNOR Spec 2314 (12 July 2024), “General reference framework for frugal AI: measuring and reducing the environmental impact of AI”, sets out methodological guidelines [9]. Technological frugality aligns with budget optimisation: by minimising energy consumption (through the use of models with fewer parameters or through the reduction of superfluous API calls), SMEs mechanically lower their token bill.

Sector initiatives, such as the work carried out by Numeum on “Ethical AI” [10], reinforce this momentum. A manifesto built around three pillars (DO, COMMUNICATE, PROGRESS) and an application guide containing 117 recommendations in its 2024 edition, these tools help companies design architectures in which the accuracy of data prevails over quantity, which limits the overloading of models’ context windows. AI governance, whether driven by ecology, ethics or the law, invariably leads to a rationalisation of data flows and, consequently, to the protection of the company’s financial capital.

The AI gateway (LLM gateway): control infrastructure

Optimising budgets and managing quotas require an architecture capable of intercepting, analysing and directing every request sent by the SME’s applications to model providers (OpenAI, Anthropic, Google, etc.). The traditional model (developers embedding API keys directly in the applications’ source code) is no longer fit for purpose: it is difficult to audit and exposes the company to uncontrolled costs. The 2026 standard rests on the use of an AI gateway (LLM gateway) acting as a centralised control plane.

An AI gateway differs from a classic API gateway (REST or GraphQL) in its ability to understand the asynchronous, probabilistic and token-priced nature of generative workloads. Without this intermediary layer, companies face unexplained outages during provider incidents, uncontrolled proliferation of high-end models, and the inability to attribute costs to the different teams.

Every request passing through an enterprise gateway must be wrapped in four logical envelopes:

  • Identity: associating the request with a user, a team, a project or a cost centre, to enable internal chargeback.

  • Policy: applying rate limits, budgets, allowlists of authorised models and dynamic routing logic.

  • Security: inspecting in real time to filter out personally identifiable information (PII) and block prompt-injection attempts.

  • Observability: recording in detail the latency, the exact number of tokens consumed (input and output) and the cost of the transaction.

Comparative analysis of AI gateways in 2026

The market offers a range of solutions addressing different constraints of latency, deployment complexity and granularity of financial controls. The table below summarises the characteristics of the dominant platforms for SMEs [15].

Gateway solutionArchitecture & deploymentLatency (overhead)Cost and quota controlUse case and SME recommendation
Bifrost (Maxim AI)Open source (Go) / fully managed~11 µs at 5,000 RPSHierarchical budgets across 4 levels (organisation, team, key, user). Strict rejection (hard block) of out-of-budget requests. Millisecond-level cost analytics.Gold standard for SMEs requiring very low latency on customer-facing applications, with enterprise-grade governance.
LiteLLMOpen source (Python)Average under light load; P99 = 90.72 s at 500 RPS, memory crash at 1,000 RPSNormalisation of requests across more than 100 providers. Spend tracking and strict enforcement of limits per virtual key and per project.SMEs with platform teams able to manage the infrastructure, favouring open-source flexibility and portability; not to be exposed to high-volume real-time traffic.
PortkeySaaS / private deployment+65% latency compared with Kong AI GatewayAdvanced observability capturing more than 40 data points per request. Strict cost segmentation by workspace, team and user.SME applications requiring complex firewalls, deep CI/CD integration and application-level rather than infrastructure-level management.
Braintrust GatewaySaaS (free beta)AverageCost attribution via customisable tags (environment, feature). Detailed tree-structured traces (span-level).Teams strongly focused on evaluating model quality (evals) and debugging reasoning chains.
Kong AI GatewayEnterprise API gateway (Lua/Go)Industry benchmarkRobust quota management and rate limiting via the existing plugin ecosystem. Enterprise security (mTLS, key rotation).SMEs already using Kong for their traditional APIs and wishing to consolidate all traffic under a single governance.
Cloudflare AI GatewayEdge infrastructureDepends on the networkReal-time dashboards for token usage. Limited hierarchical budgeting capabilities, strong DDoS protection.SMEs seeking immediate deployment and already using Cloudflare’s content delivery network (CDN).

Beyond the functional comparison, the benchmarks published by gateway vendors in 2026 (to be cross-checked) [15] make one key point clear for SMEs: under real traffic, the differences in behaviour between gateways quickly become decisive. The choice of infrastructure therefore determines the company’s financial resilience. Adopting a tool such as Bifrost or LiteLLM ensures that the financial safeguards run at the edge, and stop any excess request before the provider can even bill it.

Quota management by team: allocation and pragmatic enforcement

Treating tokens as an infinite resource is an architectural mistake. Token budgeting (Token Budgeting Architecture) consists of treating these units as a scarce and exhaustible resource, in the same way as random-access memory (RAM) in an operating system or processor time in a scheduler.

Structuring departmental quotas

The starting point is to set an overall budget not from the models’ theoretical limits (which can accept up to 2 million tokens), but from economic projections. The architectural golden rule: an application should plan to use only 85% of its theoretical maximum envelope, with the remaining 15% serving as a safety margin to absorb estimation errors or the inevitable expansion of system messages.

The breakdown of this overall budget must be carried out precisely across the SME’s teams, drawing on realistic consumption forecast models for 2026. Analysis of the workloads yields the following profiles.

Department / SME use caseEstimated task volumeMonthly consumption (tokens)Financial impact and optimisation priority
Customer service (chatbots / support)5,000 to 50,000 conversations / month15 to 250 millionVery high. Near-systematic reliance on entry-level (budget-tier) models to avoid a cost explosion. The pricing gap reaches several orders of magnitude compared with flagship models.
Finance & accounting (invoices)500 to 5,000 documents / month1.25 to 75 millionModerate. Structured extraction tasks. Using regular expressions or traditional OCR in pre-processing is recommended to limit the volume submitted to the LLM.
Software engineering (developers)Intensive daily use (copilots, agents)Hard to capCritical. The forecast budget per developer ranges, according to our field observations, between $1,000 and $3,000 per year in 2026 (MATIA estimate). Coding agents can consume 50,000 to 200,000 tokens per complex task.
Marketing (content generation)Continuous stream of copy and trend analysisVariable (high ratio of output tokens)High. Content generation involves a high proportion of output tokens (output), billed 3 to 8 times more than input tokens [14]. Strict verbosity limits are imperative.

Enforcement mechanisms: from soft limits to hard cut-offs

Quota governance does not rest on merely watching post-billing financial dashboards. It requires pre-emptive controls implemented directly in the AI gateway, orchestrated along a rigorous, graduated scale.

Warnings and soft limits. Configured to trigger when the team reaches 70% or 80% of its daily or monthly allocation. This threshold does not disrupt end users’ workflow; it triggers automated webhooks (Slack notifications, emails) that alert project managers and FinOps teams to a potentially abnormal acceleration in spending.

Conservative mode and slowdown (rate limiting). As the critical zone approaches (85% to 95% of budget), the gateway activates a throttling strategy. Requests are deliberately slowed to discourage non-essential usage. Above all, routing is altered: requests explicitly asking for access to expensive premium models are intercepted and automatically downgraded to standard models, unless the request is identified as coming from a critical process (whitelist).

Emergency mode and hard limits (hard limits & feature gating). When 100% of the quota is consumed, the gateway refuses to incur any new charges. The application undergoes a hard cut-off (hard reject) for standard requests, returning an HTTP 429 Too Many Requests code. To maintain the continuity of service perceived by users, the feature gating technique is used: advanced features are disabled in the interface, and residual basic traffic is routed exclusively to “nano” models whose inference cost is close to zero.

This hierarchical system protects the SME’s gross margins from uncontrolled consumption, while preserving a controlled operational flexibility.

Dynamic model routing: maximising yield per token

One of the most common inefficiencies in enterprise AI deployment is the routine use of the most powerful (and most expensive) models to solve trivial problems. In 2026, the cost disparity between entry-level models and top-tier models is considerable. Using a flagship model to format a text or classify a customer intent is an economic aberration: the market now offers highly capable models for fractions of a cent.

A comparison of the pricing in force in May 2026 illustrates the extent of this gap [11][12][13].

Provider and modelCost / 1M tokens (input)Cost / 1M tokens (output)Recommended use case for SMEs
OpenAI GPT-5 Nano$0.05$0.40The champion for small budgets. Ideal for classification, simple data extraction and formatting.
DeepSeek V4-Flash$0.14$0.28An ultra-economical open-weights alternative, for batch processing (batch) or high-volume pipelines.
Anthropic Claude Haiku 4.5$1.00$5.00Routing of high-volume customer support flows requiring speed and consistency.
OpenAI GPT-5$1.25$10.00General-purpose use cases, a balance between contextual nuance and moderate cost.
Anthropic Claude Opus 4.7$5.00$25.00Flagship model. New tokenizer that can consume ~35% more tokens for the same text (higher real cost); a figure from a single secondary source, to be cross-checked. To be reserved for complex analysis and deep reasoning [12].

The gap between the cheapest model (GPT-5 Nano) and the most expensive (Claude Opus 4.7) represents a cost multiplier that exceeds 60 on output and 100 on input. Given that around 70% of a company’s typical requests are basic extraction or simple question-and-answer, the absence of dynamic routing amounts to spending most of the IT budget on unused computing power.

The decision-making architecture (router logic)

Dynamic routing (Dynamic Routing) consists of inserting an algorithmic evaluation layer that intercepts the user’s request, analyses it in a few milliseconds and directs it to the model offering the best cost/performance ratio for that specific task. The execution flow of a modern intelligent router follows a logical sequence:

  • Classification of intent and complexity. A very fast “nano” model, or a set of heuristic rules, evaluates the request: a simple rewording? reading a long context? a complex mathematical problem?

  • Tier selection (tiering). The request is assigned to a capability tier. The modern SME deploys its models as a portfolio: the vast majority of traffic is directed to the core layer (the low-cost models).

  • Quality check and fallback. If the small model’s response has too low a confidence score, the gateway organises a transparent escalation to a higher model. This safety net ensures that the quality perceived by the user does not deteriorate, while achieving substantial savings on the bulk of requests handled on the first attempt.

Implementing this strategy translates into a portfolio approach: a large volume of requests routed to the cheapest models, a medium fraction to standard models, and a narrow reserve to elite models. Field reports and vendor comparisons cite API bill reductions ranging from 40% to 85% with such an architecture, without perceived quality degradation, provided the escalation confidence thresholds are calibrated correctly.

The reasoning-token caveat (thinking tokens)

2026 has seen the widespread adoption of so-called “reasoning” models (Reasoning Models), which simulate an internal chain of thought before formulating their answer. They are remarkably effective at solving software problems or complex mathematical logic.

This architecture, however, introduces an important point of caution for budget management. The “thinking tokens” generated during the internal cognitive process, although often hidden from the end user, are billed at the output-token rate (output tokens) [14]: depending on the provider, a price 3 to 8 times higher than that of input tokens.

As a result, a seemingly trivial request that triggers a prolonged reasoning loop can consume between 500 and 5,000 invisible tokens. To model correctly the budget of an SME using these advanced models, finance departments must apply a safety multiplier of 3 to 5 times the cost usually estimated for standard responses. This is why dynamic routing must formally isolate access to these reasoning models, prohibiting it for routine requests and front-line conversational agents.

Context engineering and compression: maximising the signal-to-noise ratio

Cost optimisation also involves reducing the volume of data ingested by the models. In a Transformer architecture, processing cost and latency scale quadratically with the size of the context window: doubling the amount of text provided multiplies the required computing power by approximately four. Filling this window with irrelevant documents or verbose instructions is not only costly: it also degrades the accuracy of the responses (the lost-in-the-middle phenomenon).

Traditional prompt engineering has given way to context engineering. The key skill in 2026 is no longer to craft a fine sentence, but to design the informational ecosystem in which the model operates, filtering out the noise. SMEs would do well to establish strict formatting rules.

Verbosity constraints and structured format. The most immediate technique for curbing output costs is to systematically require concise or formatted responses. Replacing long textual descriptions with instructions such as “Provide the answer as a Markdown table” or “Limit the answer to 50 words” directly reduces the most expensive part of the API bill. Likewise, using clear XML tags (, ) allows the model to isolate variables quickly without the need for long explanatory sentences.

Algorithmic prompt compression (LLMLingua). Retrieval-augmented generation (RAG) systems inject large volumes of document fragments into the context window. To avoid token inflation, programmatic tools such as LLMLingua [16] are deployed. These algorithms, which rely on small language models (SLMs), compute the perplexity of each word and remove non-essential terms (stop words, syntactic flourishes) while preserving the semantic integrity of the information. Microsoft Research benchmarks report compression rates of up to 20x with limited performance loss, and 4x savings at a compression rate of 5x, reducing in typical cases an 800-token prompt to around 160, with minimal quality degradation.

Dynamic management via reinforcement learning (ContextBudget). At the frontier of optimisation in 2026, new frameworks such as “ContextBudget” and its BACM-RL method [17] treat memory management as a sequential decision problem subject to explicit budget constraints. Instead of relying on arbitrary chunking heuristics, the system dynamically learns to compress the conversation history as it progresses, thereby avoiding capacity overflows (overflow) while maximising the retention of critical information.

The discipline imposed by context compression is fundamental. By treating the context window as a virtual bank account where every word deposited costs a few cents, software architects learn to prioritise essential data and eliminate waste at the source.

Semantic caching: the most effective saving lever

While compression reduces the unit cost of a request, caching eliminates the need to query the model altogether. In enterprise environments, a massive share of traffic is inherently redundant: users continually ask the same technical support questions, request the same summaries of HR policies, or generate reports based on identical data.

Traditional caching (Exact Match) relies on the exact comparison of character strings or their hash (SHA-256). Its limitation is well known: a tiny variation in punctuation or wording (“What is the delivery time?” vs “When will I receive my parcel?”) invalidates the cache and triggers a fresh, full call to the API. On the natural language of real users, the hit rate remains modest.

Semantic caching resolves this inefficiency by understanding the intent behind the request. It is the optimisation with the most immediate return on investment for an SME.

The three-layer architecture

The robust implementation of a semantic cache (often hosted at the AI gateway level or via in-memory databases such as Redis) is orchestrated along a defensive three-layer architecture:

  • Exact match. Fast and free. The incoming prompt is normalised (whitespace removal, lower-casing), hashed, then compared. In the event of a perfect match, the response is served in less than a millisecond.

  • Semantic similarity (Semantic Cache). If the first layer fails, the system calls on a lightweight, inexpensive embedding model to convert the sentence into a multidimensional mathematical vector. This vector is compared with the requests previously stored in a vector database. By computing the distance between vectors (cosine similarity), the system assesses the closeness of meaning; if the score exceeds a rigorous confidence threshold (for example 0.95), the stored response is reused.

  • Recourse to the LLM (LLM Fallback). Only when the first two barriers have been crossed is a paid call triggered to the large model’s API. The new response is then vectorised and stored to enrich the future cache.

Financial impact and lifecycle management

The metrics observed in production justify the integration effort. The principle is validated by the documentation of the main gateways: by intercepting redundant requests, API load and perceived latency are markedly reduced. The orders of magnitude often cited (API cost reductions of the order of 45% to 86% and latency improvements of around 88%) do not yet have a consolidated academic study as a reference; they serve as an indicative range to be validated on one’s own scope. On the measured-cost side: vector computation adds a marginal overhead of around 20 milliseconds, negligible compared with the 850+ milliseconds of an avoided LLM call.

CharacteristicTraditional cache (Exact Match)Semantic cache (Vector Similarity)
Matching methodStrict string comparison (hashing)Vector distance (cosine similarity) reflecting meaning
Handling of rewordingsSystematic failure (cache miss)Success if the similarity threshold is reached
Required infrastructureSimple key-value store (e.g. Memcached)Vector database + embedding model
Hit rate (hit rate)Low on natural language (sensitive to wording variations)High, but varies greatly with the recurrence of traffic (to be measured for one’s own case)
Latency reductionInstantaneous (< 1 ms)Strong (minimal computation overhead ~20 ms, largely offset by the gain)

Semantic caching carries one caveat: the staleness of information (staleness). Serving a cached response about a financial procedure that was changed the day before poses a real reliability problem. The gold standard therefore requires meticulous management of the time to live (TTL, Time To Live) of cache entries. Highly volatile data (prices, stock levels) must have a short TTL (a few minutes); structural information (FAQs, product documentation) can persist for several days. Event-based invalidation mechanisms (event-based invalidation) must purge the cache as soon as the source database is updated.

Finally, API providers now offer server-side prompt caching solutions (Provider-Side Prompt Caching). This feature is particularly valuable for long system messages or static RAG contexts of more than 1,000 tokens: Anthropic advertises discounts of up to 90% for repeated access to the same prefix; on the DeepSeek side, a cache hit is billed at around $0.014 for an initial cost of $0.14, that is, the same discount of around 90% [18]. Combining the local semantic cache with the provider-side prompt cache forms the most robust financial shield against cost inflation.

Mastering agentic systems: circuit breakers (kill switches)

2026 is the year of “agentic” AI. Models no longer merely generate text in response to an isolated prompt: they are embedded in autonomous workflows where they plan, use tools (web browsing, code execution) and delegate tasks among themselves, in multi-agent systems built with frameworks such as LangGraph, CrewAI or AutoGen [19]. While this evolution sharply increases productivity, it also introduces new financial and security risks that must be kept in check.

The risk of infinite loops (infinite retry loops)

Agentic autonomy alters the cost dynamics: billing is no longer linear, it becomes quadratic. With each iteration of an agent trying to correct an error, the complete history of its previous actions must be reinjected into the context window to maintain the coherence of its reasoning. An agent stuck on a task, and persisting in solving it, therefore consumes more and more tokens with each attempt.

Silent failures exist and are documented [6]. An agent programmed to analyse a code base or validate invoices, which encounters a transient API error, can enter an infinite retry loop (infinite retry loop). If it runs at night, unsupervised, it can generate thousands of useless API calls and accumulate several hundred dollars in monthly overspend on the environment concerned. The remedy is not fear: it is architecture.

Containment architecture: three levels of circuit breaker

Preventing these incidents does not rest on improving prompts, but on a containment architecture operating below the application layer. Implementing “circuit breakers” (kill switches) and firewalls is a necessity. A resilient architecture is built around three blocking layers.

The budget and threshold circuit breaker (Quota Guard Pattern). Built into the heart of the AI gateway, this circuit breaker monitors the telemetry stream in real time. It imposes an absolute, non-negotiable ceiling on the number of iterations allowed per session (for example, forced stop after 3 unsuccessful attempts) or on the amount spent (for example, a cut-off at $5 for the current task). Beyond these thresholds, the gateway blocks communication with the LLM API, freezes the agent’s state and requires the intervention of a human supervisor (Human-in-the-loop, HITL).

Cryptographic identity isolation (Identity Gate Revocation). In mature production environments, each autonomous agent is given a unique cryptographic identity (for example, SPIFFE certificates). When aberrant behaviour is detected (data leakage, excessive loops, unauthorised access attempts), the security system does not merely refuse requests: it revokes the agent’s certificate. This cryptographic cut-off is absolute: the agent loses its mutual authentication capability (mTLS), its requests to the models are rejected, its access to internal databases lapses, and the other agents refuse to communicate with it.

Time-based tool confinement (Sandbox & Data Plane Gates). The principle of least privilege must govern access to external tools (reading emails, writing to a database). The architecture prohibits perpetual access: if an agent must audit a customer file, the system issues it a strictly time-limited authorisation token (timeboxed consent), for example for 60 minutes, and confined to a specific resource. Once the deadline expires, the data plane gate closes. Thus, even in the event of a hallucination or a malicious prompt injection, the potential damage is contained in space (restricted access) and in time (rapid expiry).

Thanks to these architectural barriers, an SME ensures that error, inevitable in any probabilistic system, remains contained, with no major financial or security consequence. The infrastructure protects the application from its own failures.

Evaluating return on investment (ROI) and FinOps practices

Governance, gateways, dynamic routing and semantic caching are the tools of profitability. But to sustain the funding of these initiatives, SME finance departments (CFOs) expect quantified evidence of their impact. The debate no longer concerns the theoretical capabilities of the technology, but the return on the capital deployed.

Although 88% of organisations use AI in at least one function (Stanford AI Index 2026) [1], financial validation remains demanding: the CIO Playbook 2026 (IDC-Lenovo) measures that only 46% of POCs reach production (even though 94% of organisations anticipate a positive ROI) and fewer than 20% of initiatives manage to scale to enterprise level. This gap is explained by a poor grasp of the total cost of ownership (TCO) and by the difficulty of monetising productivity gains.

Calculating the total cost of ownership (TCO)

Financial modelling of generative AI systems differs from that of traditional software. Classic SaaS licences had fixed, predictable costs; AI generates variable costs tied to the computing intensity of each interaction. Finance departments must analyse AI through the lens of cost of goods sold (CoGS, Cost of Goods Sold) or as a variable operating expense (OpEx).

The classic return-on-investment formula applies, provided the variables are defined rigorously:

ROI (%) = [ Net benefit (Gains − TCO) / Total cost of ownership (TCO) ] × 100 The most common mistake SMEs make is to equate the TCO with only the price billed by the provider’s API (input and output tokens). The true cost (Fully Loaded Cost) is structurally broader and must include:

  • Data preparation and processing: data engineering, cleaning, structuring and vectorisation (embeddings). The Snowflake/ANZ surveys [20] identify this item as the leading operational blocker (lack of data diversity: 56%; lack of preparation: 59%) and estimate that it regularly accounts for 10% to 20% of the total budget, or even the majority of unexpected costs (vector databases, pipelines).

  • Infrastructure and orchestration: hosting the AI gateway, log storage, observability tools, server costs.

  • Technical integration and quality assurance: connector development, time spent by subject-matter experts (SME, Subject Matter Experts) annotating and evaluating the quality of responses (evals), and continuous prompt tuning.

  • Change management: training employees to ensure adoption of the tools, which often absorbs 10% to 30% of the project’s overall budget, with training costs per employee ranging from $3,000 to $20,000 according to Snowflake feedback [20].

  • Governance and compliance: monitoring AI Act-related risks and maintaining IT security.

Quantifying the benefits: from intangible to financial

On the numerator side, measuring the benefits must go beyond superficial metrics of “employee satisfaction”. To justify the investment, time savings must be converted into financial value.

The most rigorous method is to monetise the hours saved: the time gained on a task is multiplied by the employee’s loaded hourly cost (base salary plus 25% to 40% for social charges and benefits). For example, if automating a customer-returns classification process allows a team of 10 people to save 1,300 hours per year at a loaded cost of $87/hour, the gross productivity gain amounts to $113,100. Adding the reduction of manual errors and the decrease in rework, the financial value generated can easily multiply from the very first year of operation.

SMEs must also incorporate the notion of cost avoidance. If deploying a customer-support agent routed to economical models absorbs a 20% increase in the volume of incoming requests without additional hiring, the ROI includes the total cost of the salaries that did not need to be paid to support the growth.

The following table lists the key performance indicators (KPIs) useful for assessing budgetary impact by operational area.

Impact (department)Financial gains (net benefits)Operational indicators (KPIs)Target time to value (TTV)
Finance (FP&A) & operationsReduction in operating costs, increased operating leverage.Hours saved / week (e.g. 2 to 4 h/employee), reduction in forecasting cycle time (−30%).Fast (< 6 months) thanks to structured flows.
Customer service & supportAvoided recruitment costs, reduced customer attrition (churn).Autonomous resolution rate (containment rate), reduction in average response time, increase in CSAT / NPS.3 to 6 months. High-volume economical models excel here.
Security & complianceLower compliance costs, fewer high-risk errors.Number of false positives, fraud detection speed, rate of unresolved incidents.3 to 6 months.
Sales & marketingRevenue growth, higher customer lifetime value (LTV).Conversion rate, average order value (AOV), return on ad spend (MER) targeted at 5.0x.Short to medium term. Requires monitoring of generation costs.

The adoption strategy to secure ROI

Securing ROI in an SME requires methodological prudence. The “shiny object syndrome”, which drives the adoption of AI for every problem, must be set aside. The gold standard recommends initially targeting only a single use case (Single Use Case), characterised by strong potential impact, low risk and internal data that is already structured and of good quality.

It is also prudent to apply a safety discount to projections. While industry reports and solution vendors cite productivity gains of 30% to 50%, a conservative finance department will reduce these estimates by 30% to 50% in its business case, to account for adoption frictions and the performance gaps specific to the real contexts of SMEs. By framing expectations and limiting pilot projects to a horizon of 2 to 4 weeks, companies ensure that their investments translate into measurable cash rather than costly laboratory experiments.

Conclusion: engineering profitability in the age of AI

2026 marks a break in the corporate technology ecosystem. Access to the most capable artificial intelligence models has become commonplace, which erases the competitive advantage tied to merely owning the technology. The real differentiator between SMEs now lies in the ability to master the underlying economic architecture of these systems. Intelligence has become an abundant commodity; it is profitable intelligence that is scarce.

The gold standard for optimising budgets and tokens is not improvised: it is architected. It begins with a solid governance framework, aligned with demanding reference standards such as ISO/IEC 42001, which turns experimentation into accountable, auditable processes. It takes shape through the deployment of AI gateways (LLM gateways), the true backbone of financial control: these proxies ensure that every fraction of a cent spent is identified, budgeted and subject to clear consumption limits.

Cost control then rests on precise execution strategies. Dynamic request routing demonstrates that a vast majority of tasks can be accomplished by low-cost models without sacrificing quality. Context engineering and semantic caching eliminate waste at the source, converting linguistic redundancies into economies of scale. Faced with the growing autonomy of agentic systems, operational resilience is ensured by the integration of circuit breakers (kill switches) and cryptographic firewalls, which protect the organisation against technical and financial runaways.

The reprieve offered by the “Digital Omnibus” on the AI Act, until 2027-2028, is not a respite: it is a window of opportunity for the SMEs that want to move out of experimentation and into architecture. Those that adopt these principles, rigorously measure their total cost of ownership and demand a tangible, documented return on investment equip themselves with a resilient infrastructure. They turn artificial intelligence, by nature unpredictable and costly, into a lever for lasting operational efficiency, and thereby cement their competitiveness in tomorrow’s digital economy. This is precisely the logic of the MATIA Method™: structuring the method before piling up tools, and making the control of AI costs a discipline of architecture, not an emergency reaction.

Going further

The AI Express Audit & Roadmap (60 minutes by video call, no commitment) includes a review of your AI cost architecture: gateway, routing, caching and per-team quotas, with a quantified recommendation.

The white paper “AI Maturity of French SMEs 2025-2026” is available at croissance-transitions.fr.

Sources: verifiable, May 2026

[1] McKinsey & Company, The state of AI: How organizations are rewiring to capture value (and State of AI 2025/2026), late 2024 / 2025. URL: https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-how-organizations-are-rewiring-to-capture-value; “78% of respondents say their organizations use AI in at least one business function” (revised to 88% in 2025; a figure taken up by the Stanford AI Index 2026).

[2] MIT (Project NANDA), The GenAI Divide: State of AI in Business 2025, August 2025. URL: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf; “95% of enterprise AI pilots deliver zero measurable return on the P&L” (study of 300 deployments). The 2026 measurements paint a less binary picture: 11% of “AI leaders” (KPMG Global AI Pulse Q1 2026) and fewer than 10% of organisations having fully scaled AI in a function (Stanford AI Index 2026).

[3] HCLTech, survey of 467 decision-makers, May 2026. URL: https://www.hcltech.com; 43% of AI initiatives judged doomed to fail (execution gap).

[4] Piper Sandler, CIO Survey 2025/2026. URL: https://www.pipersandler.com/sites/default/files/document/cio_survey_sample.pdf; “87% [of CIOs are] expecting budget increases” for AI.

[5] Drexel LeBow / RGP, State of Data Integrity & Foundational Divide, 2025-2026., “14% of leaders responded that their organization is not prepared with the skills”; “only 14% of CFOs report clear, measurable impact”. Note: the “87% / 14%” combination is a conflation of two separate studies (see [4]).

[6] Niko Feith (Medium), The token tax: who pays when AI agents run in loops, 2026. URL: https://medium.com/@niko.feith/the-token-tax-who-pays-when-ai-agents-run-in-loops-59adef9eee1b; “total injection cap is 150,000 characters… hundreds of dollars per month in API costs… burns tokens on failed retry loops” (OpenClaw agent).

[7] ISO / ISMS.online, ISO/IEC 42001:2023, Artificial Intelligence Management System, December 2023. URL: https://www.isms.online/iso-42001/; AIMS requirements: “Conduct comprehensive AI risk assessments, AI impact assessments, Implement Ethical AI Practices”.

[8] European Commission / Modulos, Digital Omnibus Deal / AI Act FAQ, May 2026. URL: https://www.modulos.ai/blog/eu-ai-act-omnibus-deal/; high-risk obligations postponed to 2 December 2027 (standalone Annex III) and 2 August 2028 (Annex I products). The political agreement was reached on 7 May 2026.

[9] AFNOR, AFNOR Spec 2314, Référentiel général pour l’IA frugale : mesurer et réduire l’impact environnemental de l’IA, 12 July 2024. URL: https://www.afnor.org/en/news/artificial-intelligence/reference-framework-reduce-environmental-impact-ai/

[10] Numeum, Ethical AI Manifesto + Guides, 2021-2024. URL: https://ai-ethical.com/home/ ; https://ai-ethical.com/en/manifesto/; three pillars DO / COMMUNICATE / PROGRESS; 2024 edition of the guide = 117 recommendations.

[11] OpenAI / DevTk, OpenAI API Pricing Guide 2026, May 2026. URL: https://devtk.ai/en/blog/openai-api-pricing-guide-2026; GPT-5 Nano: $0.05 / $0.40 per million tokens; GPT-5: $1.25 / $10.00.

[12] Anthropic / Metacto, Anthropic API Pricing: A Full Breakdown, May 2026. URL: https://www.metacto.com/blogs/anthropic-api-pricing-a-full-breakdown-of-costs-and-integration; Claude Opus 4.7: $5.00 / $25.00 with a new tokenizer that can consume ~35% more tokens for the same text (a figure from a single secondary source, to be cross-checked); Claude Haiku 4.5: $1.00 / $5.00.

[13] DeepSeek / TLDL, DeepSeek API Pricing 2026, May 2026. URL: https://www.tldl.io/resources/deepseek-api-pricing; DeepSeek V4-Flash: $0.14 / $0.28.

[14] Anthropic Docs / Metacto, Extended thinking tokens billing, May 2026. URL: https://www.metacto.com/blogs/anthropic-api-pricing-a-full-breakdown-of-costs-and-integration; “Extended thinking tokens are billed as output tokens… charged at the standard output rate”; Input/Output ratio of 3 to 8x depending on the model (Opus 5x, R1 ~4x).

[15] Varshith V. Hegde (Dev.to), Top 5 LLM Gateways in 2026: A Deep Dive Comparison for Production Teams, 2026. URL: https://dev.to/varshithvhegde/top-5-llm-gateways-in-2026-a-deep-dive-comparison-for-production-teams-34d2; Bifrost: < 11 µs overhead at 5,000 RPS; LiteLLM: P99 = 90.72 s at 500 RPS, memory crash at 1,000 RPS; Portkey: +65% latency vs Kong.

[16] Microsoft Research, LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, 2023-2025. URL: https://llmlingua.com/llmlingua.html ; arXiv: https://arxiv.org/html/2310.05736v2; “up to 20x compression with little performance loss” and “4x savings at a prompt compression rate of 5x”.

[17] Independent researchers (arXiv), ContextBudget: Budget-Aware Context Management, BACM-RL, April 2026. URL: https://arxiv.org/abs/2604.01664; “BACM-RL, an end-to-end curriculum-based reinforcement learning approach that learns compression strategies under varying context budgets”.

[18] Anthropic / DeepSeek (Finout synthesis), Provider-side prompt caching, May 2026. URL: https://www.finout.io/blog/claude-opus-4.7-pricing-the-real-cost-story-behind-the-unchanged-price-tag; Anthropic: “up to 90% savings with prompt caching”; DeepSeek cache hit ~$0.014 for an initial cost of $0.14 (≈ −90%).

[19] Radixia AI, Designing proactive AI agents. URL: https://blog.radixia.ai/designing-proactive-ai-agents/; agent frameworks (AutoGen, etc.) and design patterns.

[20] Snowflake / Scoop, Snowflake research ANZ: More organisations investing heavily in Gen AI than the global average, 2024-2025. URL: https://www.scoop.co.nz/stories/BU2504/S00311/snowflake-research-reveals-more-anz-organisations-investing-heavily-in-gen-ai-than-the-global-average.htm; lack of data diversity: 56%; lack of preparation: 59%; training costs per employee: $3,000 to $20,000; unexpected cost drift: 30% to 50% of budget.

Article written by Paul-Antoine TUAL, AI Transformation Leader, creator of the MATIA Method™.

Paul-Antoine Tual

Paul-Antoine Tual

AI Transformation Leader · MATIA Method™ · Transition manager specialising in AI for French SMEs and mid-caps. Engineer from the École des Mines de Nantes, lawyer, developer since 1993.