Junyr Method™
Where is AI value created? Three signals for deciding without betting on one supplier
· Updated on · 8 min read · Paul-Antoine Tual
By Paul-Antoine TUAL, AI Transformation Leader, Croissance et Transitions. June 2026, revised September 2026.
The announcements of 2026 do not prove that every part of AI is becoming a commodity, but they make any strategy based on the lasting scarcity of one model, machine or software product increasingly fragile.
- Open-weight models regularly approach closed systems on particular evaluations whose meaning depends on the protocol.
- Compact models can now run multimodal and agentic workflows on compatible personal hardware.
- Orchestration patterns are becoming easier to observe, reimplement and replace.
- The practical decision is to locate each dependency, then invest in assets that survive a supplier change.
1. Open models widen the choice without eliminating performance gaps
The release of GLM-5.1 and Epoch AI’s estimate of an average four-month lag for the best open models illustrate real convergence without establishing general equivalence across licences, tasks or delivery models [1][2].
- Z.ai presents GLM-5.1 as an open-weight model designed for long-running tasks and tool use; these supplier claims still require local validation [1].
- Epoch AI’s index aggregates several evaluations and states limitations, including incomplete access to private models and benchmarks [2].
- A snapshot score says little about API stability, confidentiality, latency, cost or severe errors in your process.
- The relevant gap should be measured on representative cases with the same tools, instructions and acceptance criteria.
Older SWE-bench figures cannot support a durable hierarchy because OpenAI retired SWE-bench Verified from frontier evaluation in February 2026 and withdrew its recommendation of SWE-bench Pro in July after identifying substantial task defects [3][4].
- A ranking remains a dated observation from one protocol.
- A difference of a few points can reflect the harness, compute budget or statistical uncertainty.
- A business should retain its own regression history instead of rebuilding its architecture after every release.
The falling price of a given capability increases model substitutability, but it does not guarantee a smaller bill until usage, retries and human review are counted.
- Price per token compares one billed unit under stated conditions.
- Total workflow cost adds volume, tool calls, infrastructure, observability and human labour.
- Cost per accepted result divides that total by cases that pass review and is the useful management measure.
2. Local execution is becoming an option, not a purchasing rule
Gemma 4 12B shows that a multimodal model designed for agents can run on a computer with 16 GB of VRAM or unified memory, opening local scenarios without proving that one laptop can serve every workload [5][6].
- Google documents native vision and audio, local execution and examples of data processing and tool use on a laptop [5][6].
- The stated memory requirement establishes that the model can run; it does not guarantee throughput, battery life or application quality.
- Reasoning, context, concurrency and availability requirements can still justify an API or larger infrastructure.
- Sensitive data does not become safe merely because computation is local: access, logs, backups and updates still require governance.
A position paper by NVIDIA researchers proposes that small models could handle 80–90% of calls in an agentic system, but this is an architectural hypothesis to test rather than a representative measurement of every business [7].
- Simple tool calls, routing and structured transformations are plausible candidates for a compact model.
- Ambiguous, rare or high-impact cases can be escalated to a more capable model or a person.
- The share that can be routed depends on the corpus, agent, acceptance thresholds and cost of error.
Quantisation can reduce the memory footprint of particular components, but its gains must remain tied to the tested model, bit width, hardware and benchmark [8].
- Google Research reports TurboQuant results for KV-cache compression and vector search across several models and long-context tests [8].
- Those results do not prove that every model formerly run on a cluster preserves its quality on any workstation.
- A local-versus-cloud comparison must measure power, administration, availability, depreciation and opportunity cost alongside token prices.
3. Orchestration code travels quickly, while operations remain difficult
The Claude Code source-map episode reported by Zscaler, followed by public reimplementations, shows how quickly architectural choices can spread without proving that a product’s entire value can be cloned overnight [9][10].
- Planning loops, tool calls and permission models can be studied and reconstructed.
- Brand, distribution and user experience do not arrive with a source repository.
- Security, evaluations, connectors, support and maintenance cadence require recurring work.
- The defensible barrier therefore lies in the operated system and its adoption, rather than the secrecy of an isolated harness.
4. Model laboratories sell several forms of value
Reducing laboratories to compute landlords is premature because their position also depends on research, distribution, safety, integrated products, contracts and the ability to operate models at scale.
- API price competition can compress one margin component without determining the profitability of an entire company.
- An advanced model can command a premium on a demanding task even when a compact option suffices elsewhere.
- Usage data, developer ecosystems, brands and contractual guarantees create real switching costs.
- Private revenue, margin and valuation figures cannot support a strong conclusion without comparable audited accounts.
The partnership announced by IBM and Google Cloud in June 2026 provides a narrower signal: putting agents into production requires industry knowledge, modernisation, governance and integration in addition to a model [11].
- IBM announced a global practice involving thousands of Google Cloud-certified consultants [11].
- The release covers agents, data, cybersecurity, hybrid cloud and operational resilience.
- It is a commercial statement from the partners that documents their offer, but does not prove that every supplier will converge on consulting.
5. Map the assets that withstand a model change
The useful question is to identify each layer’s competitive pressure, dependency and required investment evidence, rather than predict a certain migration of value.
- Examine the six layers separately so that a model’s price does not conceal the cost of the complete system.
- Convert each dependency into an internal asset, contractual criterion or fallback.
| Layer | Pressure to examine | Durable asset to build |
|---|---|---|
| Compute and hosting | Price, capacity, availability, energy | Right-sized architecture and fallbacks |
| Model | Dated quality, unit price, version policy | Internal evaluation and replaceable interface |
| Agent software | Feature replicability, API dependence | Maintained connectors, controls and observability |
| Data | Quality, rights, freshness, access | Governed and traceable corpus |
| Process | Error cost, acceptance, deadlines | Rules, thresholds and human recovery |
| Skills | Supplier dependence, operating ability | Trained team and explicit responsibilities |
An SME becomes more resilient when it buys a reversible capability and retains ownership of the expected outcome, authorised data and evidence of quality.
- Version test cases, instructions and accepted outputs.
- Isolate the model call behind an interface you control.
- Measure cost per accepted case and human review time.
- Negotiate data return, migration periods and a fallback route.
Three questions for management: what share of the workflow truly needs the most capable model, what would the business lose if its supplier changed tomorrow, and which internal asset improves with every case processed?
6. What these signals establish — and what they cannot
These signals justify disciplined purchasing and close attention to dependency, but they cannot date a financial correction or determine which layers will mechanically gain or lose value.
- Performance, prices and market shares move at different rates.
- A listed company often combines several value-chain layers and sources of revenue.
- A sound operating decision can survive any market scenario: test, limit lock-in and build internal assets.
- Any financial investment decision requires regulated information and advice suited to the investor’s circumstances.
Turn the signal into a first decision
The free Junyr AI maturity audit is a 30-minute video call with no commitment that positions the business on the scale, identifies its main blocker and first sensible project, and provides a one-page follow-up.
- Bring one candidate process, its volume and two example cases.
- Use the conversation to decide what to measure before selecting a model or infrastructure.
Paul-Antoine TUAL · AI Transformation Leader · Croissance et Transitions (SAS) · Junyr Method™
Sources
[1] Z.ai, GLM-5.1 release notes, 7 April 2026.
[2] Epoch AI, Open models lag state-of-the-art closed models by 4 months, 29 May 2026.
[3] OpenAI, Why SWE-bench Verified no longer measures frontier coding capabilities, 23 February 2026.
[4] OpenAI, Separating signal from noise in coding evaluations, 8 July 2026.
[5] Google, Introducing Gemma 4 12B, 3 June 2026.
[6] Google Developers Blog, Bringing Gemma 4 12B to your Laptop, 3 June 2026.
[7] NVIDIA Research, Small Language Models are the Future of Agentic AI, position paper, version accessed 6 September 2026.
[8] Google Research, TurboQuant: Redefining AI efficiency with extreme compression, 24 March 2026.
[9] Zscaler ThreatLabz, Anthropic Claude Code Leak, 2026.
[10] GitHub, ultraworkers/claw-code repository, accessed 6 September 2026.
[11] IBM, IBM and Google Cloud Announce Strategic Partnership to Scale AI with Human Expertise and AI-Powered Delivery, 4 June 2026.
Frequently asked questions
- Is an open-weight model as good as a proprietary one?
-
The licence does not predict quality: the comparison must cover a specific version, configuration and set of real business tasks.
- Public leaderboards can shortlist candidates, but cannot identify a durable winner.
- Private hosting provides control while transferring operations, security and updates to the business.
- The right choice minimises the cost of an accepted result within quality, time and risk constraints.
- Should an SME run its models locally?
-
Local execution is credible for some workflows, but its value depends on the model, volume, latency, data and available skills.
- Google says Gemma 4 12B can run with 16 GB of VRAM or unified memory.
- A managed API often remains simpler for bursts of demand, advanced capabilities and teams without dedicated operations.
- A comparative trial should measure quality, throughput, availability and full cost before any hardware purchase.
- Does a lower price per token necessarily reduce the bill?
-
A cheaper token lowers one unit price, while the total bill also depends on volume, rework and human review.
- Measure input, output and reasoning tokens separately, along with tool calls.
- Add retries, infrastructure, observability and validation time.
- Divide the total by an accepted case, resolved ticket or approved document.
- Is an AI agent's software a durable barrier?
-
Orchestration code can be reproduced or replaced, whereas operational reliability and workflow access require continuing work.
- Agent loops, tools and permission patterns are often observable.
- Maintained connectors, evaluations, security and support remain costly to operate.
- The strongest defence combines distribution, trust, data and integration into the work.
- Where should an AI budget be concentrated?
-
The most resilient budget funds assets that remain useful when the model or tool changes.
- A governed corpus, explicit rules and access rights.
- Connectors, evaluations and process observability.
- Internal skills, adoption work and supplier exit clauses.
Paul-Antoine Tual
AI Transformation Leader · Junyr Method™ · Transition manager specialising in AI for French SMEs and mid-caps. Engineer from the École des Mines de Nantes, lawyer, developer since 1993.