MATIA Method™
Where is AI's value heading? Three signals of commoditisation that business leaders must read before the markets do
· 16 min read · Paul-Antoine Tual
By Paul-Antoine TUAL · AI Transformation Leader, Croissance et Transitions · June 2026.
Stance. What follows was not true a week ago: the launch of Gemma 4 12B (the compact variant of the Gemma 4 family released on 2 April) on 3 June 2026 shifted the frontier of what runs on a personal machine. There is much debate about whether AI is a bubble. That is the wrong question. The right question, for a business leader and an investor alike, is: into which layer of the value chain is value migrating? Three signals from 2026 sketch out a coherent answer: open models catching up, compact local models being good enough, and the software layer being cloned overnight. It does not condemn AI; it condemns certain valuations. And it points precisely to where an SME should invest.
1. First signal: an open model overtakes a proprietary flagship
On 7 April 2026, Z.ai released GLM-5.1, an open-weight model under the MIT licence (a 754 billion-parameter MoE, 40 billion active per token, 202,752-token context, autonomous execution for up to 8 hours) [1]. On release, it took the top of the SWE-Bench Pro ranking with 58.4%, ahead of GPT-5.4 (57.7%), Claude Opus 4.6 (54.2-57.3% depending on the harness) and Gemini 3.1 Pro (54.2%) [1][2]. On Terminal-Bench 2.0 it also beats Gemini 3.1 Pro (63.5% vs 56.9%), and it is the first open-weight model in history to reach the top 3 of Code Arena [2]. On coding and agentic work, precisely the two use cases driving enterprise AI transformation, a model that can be downloaded for free thus pulled level, if only for a few weeks, with the flagship models billed by usage.
The nuance that avoids sensationalism: the frontier struck back within weeks. Claude Opus 4.8 (Anthropic, 28 May 2026) retook a clear lead on SWE-bench Pro at 69.2%, Claude Opus 4.6 keeps the lead on Terminal-Bench 2.0 (68.5%), and proprietary models still dominate pure abstract reasoning (GPQA Diamond: 94.3% vs 86.2%) [2]. Epoch AI puts the average lag of the best open-weight models at 3-4 months behind frontier models [3]. The frontier has not disappeared; it has shrunk to a few points, on a few benchmarks, for a few months’ lead. A gap imperceptible for 95% of enterprise use cases.
But the economics, for their part, have already tipped. The price of a given level of performance is collapsing: Epoch AI measures a 40-fold drop in price per year to reach GPT-4 level on doctorate-level scientific questions [3]; the Gemini 3.1 Flash API costs $0.10/M input tokens where GPT-4 cost $30 in 2023, a ~99.7% drop in three years [4]; Gartner forecasts that inference on a 1 trillion-parameter model will cost 90% less in 2030 than in 2025 [5].
Interim conclusion: when the underlying asset (the model) catches up with proprietary quality and trends towards the marginal cost of compute, the rent from the model alone evaporates. What still sells at a premium is the few months’ lead, an asset that depreciates at the pace of releases.
2. Second signal: the compact local model is enough for most agentic tasks
This signal is very recent: it dates from the launch of Gemma 4 12B on 3 June 2026 (the compact variant of the family released on 2 April), four days before these lines were written. Thanks to its unified, encoder-free architecture, Gemma 4 12B runs entirely locally on a machine with 16 GB of RAM, a standard professional laptop, with a 256,000-token context and native audio/vision [6][7]. Its agentic performance is serious: 69.0% on tau2-bench (agent simulation in a real enterprise environment), against 76.9% for its bigger 31B sibling. And the 12B beats the older Gemma 3 27B, twice as heavy, on document vision and reasoning [7]. Google explicitly documents “local agentic workflows” on a laptop [8].
The strategic reading matters as much as the technical feat. By releasing for free a local model that covers most agentic use cases, Google applies the old principle “commoditise your complement” formulated by its own chief economist, Hal Varian: it deliberately cannibalises the model layer, the one from which labs with no other revenue source draw their rent, because its own value lies elsewhere: distribution, cloud, hardware, advertising. When the best-resourced player in the sector decides that the model is a loss-leader, commoditisation is no longer a market drift; it is a deliberate strategy.
How many enterprise tasks does this cover? NVIDIA Research’s position paper “Small Language Models are the Future of Agentic AI” estimates that 80 to 90% of agentic invocations fall into the “a small model is enough” category (tool calls, structured reasoning, orchestrated steps) for an inference cost 10 to 30 times lower [9].
The 2026 hardware scale (orders of magnitude observed):
| Need | Machine | Budget | What runs |
|---|---|---|---|
| Everyday local agentic work | Laptop / 16 GB mini-PC | < €1,000 | Gemma 4 12B and quantised equivalents |
| Advanced agentic + multimodal | Mac Studio 64 GB or AMD/Qualcomm mini-PC with 128 GB of unified memory | ~€2,000-3,500 | Gemma 4 31B and beyond |
| Frontier-class model run locally | Mac Studio 512 GB, option pulled from Apple’s catalogue in March 2026 (DRAM shortage); speculative secondary market or a cluster of 256 GB machines | > €12,000 (estimate) | GLM-5.1 quantised to 8-bit |
Two readings of this table. First, the new generation of unified-memory mini-PCs (AMD Ryzen AI Max+ 128 GB between ~$2,400 and $3,200, against $4,699 for the equivalent NVIDIA DGX Spark station) drives down the cost of the local token: at market prices, the ratio of tokens generated per dollar invested tilts clearly in favour of these commoditised machines [26]. Second, the disappearance of the Mac Studio 512 GB is an absurd illustration of this article’s thesis: chasing frontier infrastructure locally has become a speculative game, whereas the €2,500 machine that covers 80-90% of use cases remains on the shelf.
And the curve keeps falling thanks to quantisation innovations: TurboQuant (Google Research, presented at ICLR 2026) combines random vector rotation, aggressive quantisation and 1-bit residual correction to divide the memory footprint of the weights by 3.2. In a 4-bit + 8-bit residual configuration, the perplexity loss is strictly zero compared with the 16-bit model [10][11]. Models that required a data-centre GPU cluster now run on consumer or semi-professional hardware [11]. Quantisation does not nibble away at the commoditisation of inference: it accelerates it structurally, shifting the data-centre/local frontier by one notch every six months.
The countercurrent not to be hidden: the cost of memory. It would be dishonest to present the shift to local as plain sailing. The crisis is not a factory accident, it is a strategic choice: to serve data centres’ demand for AI chips, Samsung, SK Hynix and Micron have converted the majority of their lines to HBM (high-margin), sacrificing the classic DRAM in our PCs, servers and Macs. The result measured by TrendForce: DRAM contracts +90-95% in Q1 2026, then +58-63% in Q2 (NAND: +70-75%), with hyperscalers locking up supply through long-term contracts [23]. Gartner expects a cumulative increase of up to 130% over the year [24], and no easing is expected before the new plants ramp up in volume: late 2027 at the earliest, with a return to 2024 price standards more likely around 2028 [23][25]. It is this shortage that pushed Apple to pull the Mac Studio 512 GB from its catalogue rather than display absurd prices [26].
The practical consequence for a business leader fits in two lines. Standard needs (16-64 GB): buy despite the increase. An extra cost of around €100-150 remains marginal against the productivity gain from compact models, and this increase hits new hardware, not cloud API prices, which keep collapsing. Massive needs (128 GB and above): wait or rent. The segment is in speculative overheating; it is better to consume tokens via API over the next 12-18 months than to overpay for physical infrastructure bound to depreciate sharply when supply eases around 2028. The local/cloud trade-off is therefore made use case by use case. And quantisation partly offsets the increase by dividing the memory requirement at near-equal quality. One last point of caution: self-hosting assumes engineering skills that many SMEs do not yet have in-house. This is precisely where the value of support lies (section 5).
What still justifies heavy infrastructure: long-running frontier reasoning, real-time at large scale (voice, video, trading), very long contexts, training. This is real, but these are cutting-edge use cases, not the daily reality of 99% of SMEs and mid-caps. For most companies, the relevant AI infrastructure fits within a workstation budget, not a data-centre contract.
3. Third signal: the software layer was cloned overnight
On 31 March 2026, a .map file published by mistake in the Claude Code npm package exposed the complete internal architecture of Anthropic’s coding agent (512,000 lines, 1,906 files) [12]. Within hours, the open-source community produced a clean reimplementation of it (Claw Code, Python/Rust, without copying a single line of proprietary code), which reached 100,000 GitHub stars in 24 hours, an all-time platform record [12][13].
The lesson goes beyond the anecdote: the software layer surrounding the model (the “harness”: agentic loop, tools, permissions) is architecturally transparent. Protecting it through secrecy or intellectual property is very difficult; replicating it costs a motivated community one night’s work. What remains defensible in this layer: installed distribution, brand, trust, release cadence, deep integrations. In other words, commercial assets, not code.
4. Consequence: are the labs becoming mere compute landlords?
If the model becomes commoditised (signal 1), if inference moves down to the workstation (signal 2) and if the software is cloned (signal 3), what is left for the labs? The temptation is to answer: renting out compute power, a telecoms business, with returns compressed by the price war.
The 2026 figures already sketch out this telecoms economy. Anthropic passed $30 billion in annualised revenue in April 2026 and overtook OpenAI (~$25 billion). But the consumer subscription is plateauing, with a non-GAAP operating margin of -122% at OpenAI in Q1 2026 [14]: growth comes from the enterprise API, that is, from selling tokens by volume, literally renting out cognitive compute. And the reality is more uncomfortable still: most labs do not even own the compute they would be renting out. They themselves rent it from the hyperscalers and chipmakers, often through circular arrangements (the Nvidia-OpenAI deal reportedly represents 13% of Nvidia’s projected 2026 revenue [15]) that recall the vendor financing of late-1990s telecoms. The hyperscalers’ 2026 capex ($725 billion, including $180-190 billion for Alphabet alone, in the region of 2.5% of US GDP) [15], the applications-revenue “gap” (more than $500 billion needed to justify ~$1 trillion of infrastructure, according to Sequoia and Goldman) [15], and the NBER study of February 2026 (across ~6,000 executives: 90% measure no impact of AI on their company’s productivity over three years, for an average use of 1.5 h/week) [16] say the same thing: the current valuations of the model and application layers assume a rent that the three signals above are dissolving.
The labs’ counter-arguments exist and deserve to be taken seriously: vertical integration into products (the lead is sold as a product, not as an API), usage data, consumer brand, enterprise relationships. But each of these assets belongs to distribution and service, not to the model.
And there is an even more explicit admission: the labs and their allies are themselves launching consulting practices. The IBM-Google Cloud alliance of 4 June 2026 mobilises thousands of consultants around Gemini Enterprise [18]; the model vendors are developing their own client-facing integration and deployment arms. A model vendor that hires consultants, a service-margin business where the API promised a software margin, acknowledges three things: that the model alone no longer sells for enough, that the value is in the last mile it did not control, and that the migration thesis is correct. When the vendor moves downstream along the chain, it is because the rent from its original layer is eroding.
5. Where value takes refuge
The grid is now readable, layer by layer:
| Layer | Value trend | Why |
|---|---|---|
| Semiconductors, energy, data-centre real estate | Sustained but cyclical | Real physical scarcity; risk of overcapacity |
| Models (labs) | Compressing | Open-weight + collapse in inference prices |
| Software layer / agents | Compressing fast | Cloneable in hours; no lasting protection |
| Living proprietary data | Rising | Continuous operational cost, not replicable by code |
| Approvals, compliance, distribution | Rising | Administrative and contractual barriers |
| Workflow integration + human support | Rising strongly | The service IS the product; Infosys puts the AI services opportunity at $300-400 billion by 2030 [17]; IBM and Google Cloud form an alliance on 4 June 2026 precisely around “human expertise at scale” [18]; in France, AI already accounts for more than 11% of Capgemini’s bookings in Q1 2026 [19] |
Human support is therefore not “what is left” by default: it is the layer towards which all rational players are converging, labs included. For an SME, the consequence is direct: sustainable AI spending is neither the model licence nor the tool of the moment. It is the method, governed data and internal skills. Exactly the thesis of the MATIA Method™: structuring the organisation to absorb models and tools that have become interchangeable, rather than marrying a single vendor.
Three questions to ask first thing on Monday. 1. What share of our AI use really requires a frontier model billed by usage, and what share would run on a €2,500 machine? 2. If our main AI tool were to disappear or increase tenfold in price tomorrow, what would we lose: code (replaceable) or data and skills (our own)? 3. Does our AI budget fund the rents of vendors on the way to commoditisation, or internal assets that appreciate?
6. What this says about valuations, without doom-mongering
Should we conclude that a “bursting” is imminent? No: recent multi-method work [20] and layer-by-layer analysis [21] converge on a more useful diagnosis. There is not one bubble; there are overvalued layers and sustainable layers, with different correction horizons. The concentration of the indices and a Shiller P/E of 42.3, very close to the record of 44.2 in 2000, make the correction of the fragile layers painful for everyone: on 3 June 2026, Broadcom lost 15% and $280 billion in market cap in a single session over a forecast barely below expectations [22]. This does not change the direction of the migration. The business leader does not have to predict the date; they have to position themselves on the right side of the migration. It is a schedule of decisions, not an alarm.
In practice
Croissance et Transitions supports the leaders of SMEs and mid-caps in positioning themselves on the right side of this migration: maturity diagnostic (MATIA Method™), local/cloud trade-offs use case by use case, governance of proprietary data and upskilling of teams. These are the assets that appreciate as models and tools become commoditised.
→ AI Express Audit & Roadmap: 60 minutes over video call. croissance-transitions.fr
Paul-Antoine TUAL · AI Transformation Leader · Croissance et Transitions (SAS) · MATIA Method™
FAQ
Is an open-source model really on a par with proprietary models in 2026? On coding and agentic work, yes for some benchmarks: GLM-5.1 (MIT licence) beats Gemini 3.1 Pro on SWE-Bench Pro (58.4% vs 54.2%) and on Terminal-Bench 2.0 (63.5% vs 56.9%). Proprietary models keep a lead on other composites, a lead measured in points and months, no longer in orders of magnitude.
Does an SME need heavy cloud infrastructure for agentic AI? In 80 to 90% of agentic invocations (NVIDIA Research estimate), a compact model is enough, and runs locally from 16 GB of RAM (Gemma 4 12B), or on a Mac Studio 64 GB (~€2,500) for the 31B version.
What does quantisation change for an SME? Techniques such as TurboQuant (Google Research, ICLR 2026) divide the memory footprint of models by ~3.2 with zero to marginal quality loss: models once reserved for clusters run on semi-professional hardware. The data-centre/local frontier keeps coming down.
Can the software layer of AI agents be protected? Weakly: the architecture of Claude Code, exposed by mistake in March 2026, was reimplemented in open-source overnight (Claw Code, 100,000 GitHub stars in 24 h). Durable defence is commercial (distribution, trust, integrations), not technical.
Where should an AI budget be invested in 2026? In the layers that appreciate: governed proprietary data, compliance, workflow integration and internal skills, not in the rents of models or tools on the way to commoditisation.
Sources
[1] Z.ai, GLM-5.1: Towards Long-Horizon Tasks (7 April 2026). https://z.ai/blog/glm-5.1 · Documentation: https://docs.z.ai/guides/llm/glm-5.1
[2] DeepInfra, GLM-5.1 Model Overview (SWE-Bench Pro / Terminal-Bench 2.0 / Code Arena). https://deepinfra.com/blog/glm-5-1-model-overview · Artificial Analysis, GLM-5.1 vs Gemini 3.1 Pro. https://artificialanalysis.ai
[3] Epoch AI, Open-weight models lag state-of-the-art by around 3 months on average. https://epoch.ai/data-insights/open-closed-eci-gap · LLM inference price trends. https://epoch.ai/data-insights/llm-inference-price-trends
[4] AI Magicx, The LLM Pricing Collapse of 2026. https://www.aimagicx.com/blog/llm-pricing-collapse-developer-guide-building-cheap-ai-2026
[5] Gartner, By 2030, Performing Inference on an LLM With 1 Trillion Parameters Will Cost Over 90% Less Than in 2025 (25 March 2026). https://www.gartner.com/en/newsroom/press-releases/2026-03-25-gartner-predicts-that-by-2030-performing-inference-on-an-llm-with-1-trillion-parameters-will-cost-genai-providers-over-90-percent-less-than-in-2025
[6] Google, Introducing Gemma 4 12B: a unified, encoder-free multimodal model (3 June 2026). https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/
[7] Sierra Research, τ²-bench. https://taubench.com/ · VentureBeat, Gemma 4 12B runs entirely locally on a typical 16GB enterprise laptop. https://venturebeat.com
[8] Google Developers Blog, Bringing Gemma 4 12B to your Laptop: Unlocking Local, Agentic Workflows. https://developers.googleblog.com/bringing-gemma-4-12b-to-your-laptop-unlocking-local-agentic-workflows-with-google-ai-edge/
[9] NVIDIA Research, Small Language Models are the Future of Agentic AI (arXiv:2506.02153). https://arxiv.org/abs/2506.02153
[10] Google Research, TurboQuant: Redefining AI efficiency with extreme compression (ICLR 2026). https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/ · https://arxiv.org/abs/2504.19874
[11] AI Indigo, TurboQuant Explained: How This New Compression Method Changes Local LLM Inference. https://aiindigo.com/blog/turboquant-explained-how-this-new-compression-method-changes-local-llm-inference
[12] Zscaler ThreatLabz, Anthropic Claude Code Leak. https://www.zscaler.com/blogs/security-research/anthropic-claude-code-leak
[13] GitHub, claw-code (ultraworkers). https://github.com/ultraworkers/claw-code · Cybernews, Leaked Claude Code source spawns GitHub’s fastest repo. https://cybernews.com
[14] SaaStr, Anthropic Just Passed OpenAI in Revenue. https://www.saastr.com · Medium, Anthropic just passed OpenAI in revenue. Here is why it matters. https://medium.com/@david.j.sea/anthropic-just-passed-openai-in-revenue-here-is-why-it-matters-e3dd9bb04069
[15] Allianz Research, AI capex cycle: war-proof for now (25 March 2026). https://www.allianz.com · Goldman Sachs, AI: In a Bubble? https://www.goldmansachs.com · IDC, Circular financing has muddied the AI story. https://www.idc.com/resource-center/blog/circular-financing-has-muddied-the-ai-story-watch-the-application-layer-instead/
[16] NBER, Firm Data on AI (Working Paper w34836, February 2026). https://www.nber.org/papers/w34836 · NBER w34851, Does Generative AI Narrow Education-Based Productivity Gaps? https://www.nber.org
[17] Infosys, AI First Value Framework: AI Services Opportunity of Over $300 Billion. https://www.infosys.com/newsroom/press-releases/2026/unveils-ai-first-value-framework.html
[18] IBM, IBM and Google Cloud Announce Strategic Partnership to Scale AI with Human Expertise (4 June 2026). https://newsroom.ibm.com/2026-06-04-ibm-and-google-cloud-announce-strategic-partnership-to-scale-ai-with-human-expertise-and-ai-powered-delivery
[19] Capgemini Investors, Q1 2026 revenues. https://investors.capgemini.com
[20] arXiv, Boom, Bubble, or Buildout? A Multi-Method Evaluation (2606.01575). https://arxiv.org/html/2606.01575
[21] VentureBeat, Stop calling it ‘The AI bubble’: It’s actually multiple bubbles. https://venturebeat.com/infrastructure/stop-calling-it-the-ai-bubble-its-actually-multiple-bubbles-each-with-a
[22] TradingKey, S&P 500 valuation, Shiller P/E 42.32, Broadcom -15%. https://www.tradingkey.com/analysis/stocks/us-stocks/261950917-sp500-valuation-bubble-ai-concentration-shiller-pe-buffett-indicator-fed-hawkish-yield-market-nifty-fifty-strategy-tradingkey
[23] TrendForce, AI Server Demand to Drive Memory Contract Price Increases in 2Q26 (31 March 2026). https://www.trendforce.com/presscenter/news/20260331-12995.html · Tom’s Hardware, DRAM prices predicted to jump 63% in Q2, NAND up to 75%. https://www.tomshardware.com/pc-components/dram/dram-and-nand-contract-prices-to-climb-again-in-q2
[24] TechTimes, RAM Prices 2026: Gartner Forecasts 130% Memory Cost Surge (5 June 2026). https://www.techtimes.com/articles/317872/20260605/ram-prices-2026-buy-now-wait-gartner-forecasts-130-memory-cost-surge.htm
[25] SaaS Sentinel, RAM Shortage Could Last Until 2028 as AI Demand Reshapes Memory Markets (April 2026). https://saassentinel.com/2026/04/19/ram-shortage-could-last-until-2028-as-ai-demand-reshapes-memory-markets/
[26] Tom’s Hardware, Apple pulls 512GB Mac Studio upgrade option. https://www.tomshardware.com/tech-industry/apple-pulls-512-mac-studio-upgrade-option · MacRumors (5 March 2026). https://www.macrumors.com/2026/03/05/mac-studio-no-512gb-ram-upgrade/ · TechSpot, AMD Ryzen AI Halo mini PC, 128GB, vs DGX Spark. https://www.techspot.com/news/112287-amd-ryzen-ai-halo-mini-pc-coming-june.html · Liliputing, Ryzen AI Max+ mini PCs with 128GB. https://liliputing.com/more-ryzen-ai-max-395-mini-pcs-with-128gb-are-now-available-if-you-can-afford-one/
Paul-Antoine Tual
AI Transformation Leader · MATIA Method™ · Transition manager specialising in AI for French SMEs and mid-caps. Engineer from the École des Mines de Nantes, lawyer, developer since 1993.