Markets & AI
The end of the model race: frontier AI is becoming a commodity
· 25 min read · Paul-Antoine Tual
By Paul-Antoine TUAL, AI Transformation Leader, Croissance et Transitions. August 2026.
Where this sits. This fifth instalment of our “Markets & AI” series closes a loop opened in January: in Where is AI’s value going?, we watched for three signals of model commoditisation. Eight months later, all three have materialised, and then some. We have already followed AI’s cash and its closed loop, the bubble signals and the Microsoft-Mistral pivot, the Saaspocalypse and the real AI maturity of the CAC 40 and Nasdaq Top 10. Here we look at the race itself: who still wins what, at what rhythm models ship, and what remains worth buying when performance becomes a commodity. This article is not investment advice; figures are dated, and estimates and self-reported measurements are flagged as such.
1. Benchmarks are saturating, and nine Elo points separate the top five
For three years, the model race was measured against three simple exams: MMLU for general knowledge, GSM8K for arithmetic, HumanEval for code. Those three no longer measure anything. Frontier models now cluster in a band of roughly 88 to 94% on MMLU, exceed 95% on GSM8K and 90% on HumanEval, to the point that vendors have simply stopped publishing these figures in their announcements [1]. The saturation has a documented cause: contamination. Scale AI’s GSM1k study, which re-examined models on problems identical in spirit but never published, measured drops of up to 13 points relative to GSM8K, correlated with memorisation of the original set [2]. An exam whose papers have circulated for years no longer ranks its candidates.
The next generation of tests exists, and it is far harder. GPQA Diamond: 198 doctoral-level science questions designed to resist web search, on which human PhDs themselves plateau between 65 and 74% [3]. SWE-bench Verified: 500 real GitHub repository tickets to resolve end to end [4]. Humanity’s Last Exam: 2,500 questions written by around a thousand experts from 50 countries, published in Nature in January 2026, on which no system exceeds 65% to date [5]. These tests still discriminate. What they show, however, looks less and less like a race with a winner.
Consider the most watched leaderboard in the field, LMArena, which has millions of humans vote blind (7.8 million votes at the latest reading). As of 12 August 2026, nine Elo points separate first place from fifth, and a single point separates first from second, a full statistical tie [6]. For scale: nine Elo points translate to roughly 51% preference against 49%. The top spot itself has become musical chairs; since November 2025, every major release has held it for a few weeks before giving it up.
One last, quieter symptom: the benchmarks themselves have become a communications battleground. The same model posts three scores depending on who measures. On Humanity’s Last Exam, the latest Claude generation scores 46.4% on Scale AI’s official leaderboard (no tools), 55.5% on Artificial Analysis’s independent harness, and 64.7% in vendor figures measured with tools [5][7]. These three numbers do not contradict each other: they measure different protocols (tool access, scaffolding, compute effort), and the gap between them, up to 20 points on SWE-bench Pro between Scale’s neutral harness and self-reported figures, has grown larger than the gap between competitors [4]. When the protocol gap exceeds the performance gap, a score table reads like a press release.
2. The release tempo: a competitive clock, a financial stage
If there is no lasting lead left to take, why are releases accelerating? Because they are: on OpenAI’s main line, the interval between two versions fell from about five months in early 2025 (GPT-4.5 to GPT-5) to six or seven weeks on average in 2026 (GPT-5.1, 5.2, 5.3, 5.4, 5.5 and 5.6 followed each other at 29 to 77 days) [8]. Anthropic chained Opus 4.6, Sonnet 4.6, Opus 4.7, Opus 4.8, Fable 5, Sonnet 5 and Opus 5 between February and July 2026 [8]. The week of 17 to 23 July 2026 alone saw seven major model releases across vendors [9].
The commentators’ short answer, “they ship models for the stock market”, deserves verification rather than assertion. Here is what the facts establish, and what they do not.
What is established is the deliberate coupling of financial announcements and releases. On 28 May 2026, Anthropic launched Claude Opus 4.8 on the very day it announced its record raise of 65 billion dollars at a 965 billion valuation; the trade press headlined the simultaneity [10]. On 6 January 2026, Grok 5’s existence “in training” was made official inside xAI’s fundraising press release, 20 billion dollars at a 230 billion valuation: a financial document served as the first product announcement [11]. Three days after the Kimi K3 launch in mid-July, Bloomberg revealed that Moonshot AI was preparing a Hong Kong listing “within six months”, its valuation having quintupled in press ranges [12]. And on 24 July, Yahoo Finance titled soberly: “Anthropic debuts Opus 5 model as company preps for IPO later this year” [13].
Spring 2026 even saw the two leading laboratories run the same sequence, raise, IPO filing, model launch, within a single quarter (see Figure 2). The constraints of capital, meanwhile, are written into the contracts: the second tranche of the SoftBank round was conditional on OpenAI’s for-profit restructuring completing before 31 December 2025 (executed on 28 October), and 35 of the 50 billion dollars pledged by Amazon in the March 2026 round are conditional on an IPO or on reaching AGI [14]. When the disbursement of funds depends on deadlines, deadlines govern communication.
Now for what the facts do not establish: that release dates are set by the stock-market calendar. No investigation by Reuters, Bloomberg, the FT or the WSJ demonstrates it, and we looked. The accelerations documented from the inside are competitive: Sam Altman’s “code red” memo after the Gemini 3 launch, revealed in early December 2025, pulled GPT-5.2 forward to 11 December [15]. The same-day duels follow the same logic (Opus 4.6 and GPT-5.3-Codex both shipped on 5 February). And the counter-examples are massive: Gemini 3.5 Pro, announced at Google I/O in May, has missed three launch windows and remains unavailable in mid-August [16]; Grok 5 has still not shipped despite SpaceX’s stock-market debut, the largest in history, on 12 June [17]; OpenAI presented its “Astra” system on 1 August, ten open mathematical problems solved, without releasing or dating it, months before an IPO that would presumably have welcomed the launch [18]. When the technology is not ready, the financial deadline does not suffice.
The precise formulation fits in one sentence. The financial calendar has not become the clock of releases: it has become their stage, and the releases an argument for valuation. The clock remains competition, ticking through ever more frequent increments; meanwhile the promised generational leaps (Gemini 3.5 Pro, Grok 5, DeepSeek R2, Llama 4 Behemoth) slip, keep everyone waiting or vanish from roadmaps [8][16]. Many versions, few breakthroughs. A market that ships versions to the rhythm of its communications rather than the rhythm of its discoveries has a name in industrial economics: a maturing market.
3. Open weights have joined the leading pack, and China distributes them
The second engine of commoditisation is the convergence between closed and open-weight models. In November 2024, Epoch AI measured roughly one year of lag between the best open models and the closed frontier. In October 2025: three months. At the latest published reading, May 2026: four months, or eight points on their capability index, the equivalent of the gap between GPT-5 and GPT-5.5 [19]. OpenRouter, which sees the traffic of thousands of applications, describes “a constant gap of three to six months sustained for over eighteen months” [20]. And on the Artificial Analysis index read on 17 August 2026, three points separate the best closed model (Claude Opus 5, 63) from the best open one (Kimi K3, 60) [7]. A year’s lead has become a quarter.
This convergence has a chronology and a trigger. DeepSeek R1, published as open weights on 20 January 2025 for a derisory training cost (the reinforcement-learning phase alone cost 294,000 dollars according to the Nature-reviewed paper), erased 589 billion dollars of Nvidia’s market capitalisation in a single session, the largest one-day loss in Wall Street history [21]. Kimi K2 Thinking, in November 2025, became the first open model to beat proprietary systems on tool-assisted agentic tests [22]. And 2026 brought the heavyweights in quick succession: GLM-5 then 5.2 (topping the mid-July Design Arena leaderboard, ahead of Claude Fable 5, according to an aggregator reading), Qwen 3.5 then Qwen3.8-Max, DeepSeek V4, and Kimi K3, 2.8 trillion parameters published on Hugging Face in late July [7][23]. All five families are Chinese. Meta, meanwhile, walked the other way: Llama 4 Behemoth never shipped, and the first model from Meta Superintelligence Labs, Muse Spark, launched in April 2026, is closed [24]. On distribution, the switch has happened: Qwen derivatives on Hugging Face amount to 2.6 times Meta’s entire footprint, and Chinese models went from roughly 1% to roughly 61% of tokens served through OpenRouter in eighteen months, according to convergent platform readings [25].
Honesty requires keeping two caveats, because they decide what a business leader can actually do with all this. The first concerns measurement: the parity is benchmark parity. At equal scores, closed models remain more robust in real use, above all on long agentic tasks; the LMArena text leaderboard counts no open model in its top 30 [6][26]. Even Moonshot’s own published tables show Kimi K3 behind Claude Fable 5 and GPT-5.6 Sol on frontier software-engineering tests [23]. The second is commercial: at the precise moment open weights caught up technically, their share of American enterprise API usage fell from 19% to 11% between 2024 and 2025 (Menlo Ventures, December 2025 reading), and more than a third of large accounts now say they prefer closed models [27]. Performance has converged; trust, contractual liability and operational simplicity have not, yet. Add the prosaic: serving Kimi K3 at full precision requires on the order of 1.4 terabytes of GPU memory [23]. An open weight is not a free service.
4. Listed prices collapse, bills climb
The third engine is price. At constant performance, the fall is vertiginous and documented over several years: Stanford’s AI Index puts the drop in inference cost at GPT-3.5 level at more than 280-fold between late 2022 and late 2024, a16z named the phenomenon “LLMflation” (a tenfold division every year), and Epoch AI measures, depending on the chosen capability level, annual declines of 9-fold to 900-fold, with a median around 50 [28]. The year 2026 added an explicit price war: Anthropic cancelled the planned price rise for Claude Sonnet 5 (the launch price, 2 dollars per million input tokens and 10 for output, simply became the price), OpenAI cut its entry-level GPT-5.6 Luna by 80% in late July, and Google sells its recent Gemini Flash models at half price “until 31 December 2026”, at which point the rate will double, which is another way of saying price has become the weapon [29]. DeepSeek, for its part, bills its V4-Flash at 0.44 dollars per million input tokens, half that in off-peak hours [30].
A business leader who stopped there would commit the mirror-image error of the one who only reads benchmarks. The listed price per million tokens falls; token consumption explodes. Reasoning models bill their invisible deliberation as output tokens, three to five times the visible volume in ordinary use, ten to thirty times in extreme cases [31]. Long context costs extra: Google and xAI double the input rate beyond 200,000 tokens, and OpenAI introduced a comparable premium on its 2026 ranges at the very moment Anthropic removed its own (one million tokens at the standard rate) [29]. The structural discounts exist, around 90% on cache reads and 50% for batch processing across the big four, but they presuppose an architecture designed for them, the one we detail in our token budget guide and our article on caching and context optimisation. The net result has been measured: according to the FinOps Foundation’s 2026 survey, 73% of organisations exceeded their AI cost projections, citing precisely underestimated token volumes and the hidden tokens of reasoning models, while 98% of practitioners now say they manage these costs, against 31% two years earlier [32]. Price has become a competence.
Buying intelligence as a commodity: three reflexes
Let us recap what the three sections establish. The top of the leaderboard fits within nine Elo points and first place rotates with every release. The release tempo follows communications, competitive and financial, more than technical breakthroughs. The best open model sits three points from the best closed one, and price at constant performance is divided by ten every year. In January we asked where the value would go if models became commonplace; the answer arrived faster than expected. For the leader of an SME or mid-cap, it translates into three buyer’s reflexes.
Buy reversible. In a market where the leader changes every two months, the costly mistake has changed in nature: a poor model choice is corrected in days, while a contractual lock-in with yesterday’s best model is paid for over years. A gateway that routes each family of tasks to the appropriate model (frontier for critical reasoning, mid-range or open weights for volume) turns the price war into an advantage, rather than watching it from the stands.
Measure at home. Section 1 showed it: between the neutral harness and the vendor figure, the same benchmark can differ by 20 points. The only leaderboard that engages your cash is the one built from your own use cases, on your documents, with your acceptance criteria. Twenty real tasks evaluated blind are worth more than every press release of the season.
Move the moat. When raw intelligence becomes a commodity, it stops being a competitive advantage for anyone, including you. What remains defensible is what the model does not supply: your proprietary data and its quality, your tooled processes, your ability to verify outputs before they commit the company, the governance that makes real maturity far more than a subscription to the latest model.
And this is where open-weight convergence acquires its full meaning for an SME. Open models at the level of the leading pack, served at commodity prices, put within reach what used to belong to laboratories: specialising AI on your own know-how. The gradation is well known. Inject your knowledge base and decision rules into the context of any good model, whatever the brand; then, on a stable, high-volume domain, fine-tune an open model on your business corpus, or distil a small dedicated model you can host on your own premises. The movement is already massive: 83% of Hugging Face downloads are models under one billion parameters, built to be specialised, not to shine on leaderboards [25]. Techniques will go out of fashion, and large accounts have already cooled on systematic fine-tuning [27]; the asset remains: a clean corpus, explicit rules, an AI that answers with your house’s tone, reflexes and red lines. Giving AI the personality of the company is the investment that commoditisation finally makes affordable, and the only one no supplier will ever rent to you. That is precisely the ground of the Junyr Method™: the commoditisation of the engine makes the chassis decisive.
The model race is ending the way most technology races end: no photo finish, just product banalisation and margin migration. The losers will be those who keep buying magic. The winners will buy kilowatt-hours of intelligence on a meter, with a reversibility clause, and invest the price gap in the one thing that cannot be rented: an AI with their company’s personality.
Disclaimer. This article is educational and informative. It constitutes neither investment advice nor a recommendation to buy or sell. The rankings, prices and valuations cited are dated (readings of 17 August 2026 unless stated otherwise) and move fast; press estimates, aggregator readings and vendor self-reported measurements are flagged as such.
Sources
[1] Convergent 2026 readings on the saturation of MMLU, GSM8K and HumanEval (score bands of 88-94%, above 95%, above 90%) and vendors ceasing to publish them: Stanford HAI, AI Index Report 2025, benchmarks section (hai.stanford.edu); sector analyses of August 2026 (datavlab.ai, 3 August 2026; benchmarkingagents.com).
[2] Scale AI, A Careful Examination of Large Language Model Performance in Grade School Arithmetic (GSM1k), arXiv:2405.00332: drops of up to 13 points relative to GSM8K, correlated with memorisation (r² = 0.32).
[3] Rein et al., GPQA: A Graduate-Level Google-Proof Q&A Benchmark (2023): full GPQA 448 questions, Diamond subset 198 questions; PhD experts at 65% (74% excluding flawed questions), tooled non-experts at 34%.
[4] OpenAI and the SWE-bench authors, Introducing SWE-bench Verified (August 2024), 500 validated tasks; Scale AI, SWE-bench Pro, arXiv:2509.16941 (September 2025), 1,865 problems, SEAL leaderboard: best neutral-harness score 59.1% (GPT-5.4 xHigh) against 80.3% vendor self-reported, readings of 17 August 2026.
[5] Phan et al., Humanity’s Last Exam, Nature 649, 1139-1146 (28 January 2026); official site agi.safe.ai (CAIS and Scale AI, 2,500 questions, around 1,000 experts, 50 countries); official Scale leaderboard (labs.scale.com): Gemini 3.1 Pro 46.44% without tools, reading of 17 August 2026.
[6] LMArena (arena.ai, formerly lmarena.ai), text leaderboard updated 12 August 2026, code leaderboard of 15 August 2026, read on 17 August 2026: 7,779,985 votes, 391 models; text top 5 between 1506 and 1497 Elo points, first open-weight model at rank 33.
[7] Artificial Analysis, Intelligence Index and HLE evaluations (artificialanalysis.ai), readings of 17 August 2026: Claude Opus 5 (max) 63, Kimi K3 (max) 60; Claude Fable 5 at 55.5% on HLE (independent harness).
[8] Release chronology reconstructed from primary sources (official changelogs and blogs of OpenAI, Anthropic, DeepSeek, Mistral, x.ai, Alibaba Cloud) and dated press (TechCrunch, CNBC, Axios, Caixin), January 2025 to August 2026; intervals computed on verified dates.
[9] Digital Applied, Frontier Model Release Velocity Index and Seven Days, Seven Releases (July 2026), specialist blog analyses, cited as indicative counts.
[10] SiliconANGLE, “As Anthropic launches Claude Opus 4.8, it raises $65B in new funding” (28 May 2026); Anthropic, Series H announcement ($65bn, post-money valuation $965bn, 28 May 2026); Sherwood News (28 May 2026).
[11] xAI, Series E announcement (x.ai/news/series-e, 6 January 2026): $20bn raised at a $230bn valuation, first official confirmation of Grok 5 “in training”; CNBC (6 January 2026).
[12] Bloomberg via Nikkei Asia, “China’s Moonshot AI plans Hong Kong IPO as Kimi K3 shocks Silicon Valley” (20 July 2026); valuation ranges ($20bn to $35bn) reported by the press, unconfirmed.
[13] Yahoo Finance (Daniel Howley), “Anthropic debuts Opus 5 model as company preps for IPO later this year” (24 July 2026); CNBC (15 July 2026), investor meetings scheduled by Goldman Sachs, Morgan Stanley and JPMorgan, Nasdaq listing targeted for the autumn.
[14] CNBC (31 March 2025), SoftBank tranche conditional on restructuring before 31 December 2025; CNBC and Fortune (28 October 2025), restructuring executed; Bloomberg and CNBC (31 March 2026), OpenAI round of $122bn at $852bn post-money, 35 of Amazon’s $50bn conditional on an IPO or AGI; OpenAI, confidential S-1 filing made public (8 June 2026).
[15] Forbes (2 December 2025), Sam Altman’s internal “code red” memo after Gemini 3; Pure AI (16 December 2025), GPT-5.2 release pulled forward to 11 December.
[16] TechCrunch, “Google releases three new Gemini models, but no 3.5 Pro” (21 July 2026); Forbes, “Gemini 3.5 Pro Delay Continues” (13 August 2026).
[17] CNBC (12 June 2026), SpaceX stock-market debut (xAI included): $75bn raised, the largest in history, closing up 19%; felloai (August 2026), Grok 5 still awaited.
[18] Coverage of OpenAI’s “Astra” presentation (1 August 2026): ten open mathematical problems solved, with no release or date announced.
[19] Epoch AI, Open vs. closed AI: How behind are open models? (4 November 2024, lag of about one year); data insights of 30 October 2025 (about 3 months) and 29 May 2026 (4 months, 8 index points, real gap probably larger), epoch.ai.
[20] OpenRouter, The open-weight models that matter (official blog, 27 June 2026): constant gap of 3 to 6 months sustained for over 18 months.
[21] DeepSeek, official changelog (R1, 20 January 2025); CNBC and Bloomberg (27 January 2025), Nvidia down 17%, about $589bn of market capitalisation erased in one session; Nature and coverage (September 2025), cost of R1’s reinforcement-learning phase: $294,000.
[22] deeplearning.ai, The Batch (November 2025), Kimi K2 Thinking first open model to beat proprietary systems on tooled agentic benchmarks (44.9% on HLE with tools).
[23] Simon Willison (16 July 2026) and Constellation Research, Kimi K3 release (2.8 trillion parameters, weights published in late July, about 1.4 TB of GPU memory in MXFP4); result tables published by Moonshot (self-reported); mid-July Design Arena reading via aggregator (benchlm.ai), GLM-5.2 ahead of Claude Fable 5.
[24] TechCrunch and Bloomberg (8 April 2026), launch of Muse Spark, the first proprietary model from Meta Superintelligence Labs; December 2025 press on the abandonment of the open line; Llama 4 Behemoth announced in April 2025, never released.
[25] Hugging Face, State of Open Models: Summer 2026 (official blog, 14 August 2026): 151,448 Qwen derivatives, 2.6 times Meta’s entire footprint; convergent OpenRouter counts (officechai, datagravity, May-June 2026): about 61% of tokens on Chinese models, US share down from about 70% to about 30% in a year; platform figures, secondary readings.
[26] Nathan Lambert, Interconnects, My bets on open models (15 April 2026): superior robustness of closed models at equal scores, coding agents as the first domain of clear dominance.
[27] Menlo Ventures, 2025: The State of Generative AI in the Enterprise (9 December 2025, about 500 decision-makers): open source at 11% of enterprise API usage in 2025 against 19% in 2024, Chinese models about 1%; a16z, survey of 100 Global 2000 executives (30 January 2026), growing preference for closed models and cooling interest in systematic fine-tuning.
[28] Stanford HAI, AI Index Report 2025 and 2026 (13 April 2026): inference cost at GPT-3.5 level divided by more than 280 between November 2022 and October 2024; a16z, LLMflation (November 2024), tenfold division per year at constant performance; Epoch AI, LLM inference price trends: annual declines of 9-fold to 900-fold depending on the capability threshold, median around 50.
[29] Official pricing pages read on 17 August 2026: Anthropic (platform.claude.com, cancellation of the Claude Sonnet 5 price rise, one-million-token context at the standard rate); OpenAI (developers.openai.com, long-context premiums on the 5.4 to 5.6 ranges); Google (ai.google.dev, page of 13 August 2026, premium beyond 200,000 tokens on Pro models, promotional Flash 3.6 and 3.7 rates doubling on 1 January 2027); xAI (docs.x.ai); VentureBeat (30 July 2026), 80% cut on GPT-5.6 Luna.
[30] DeepSeek, official pricing page (api-docs.deepseek.com, read on 17 August 2026): V4-Flash at $0.44 per million input tokens ($0.22 off-peak), V4-Pro at $1.32 ($0.66).
[31] Convergent measurements of the reasoning-token overhead (2026): factor of 3 to 5 in ordinary use, 10 to 30 in extreme cases (analyses by silicondata.com, benchlm.ai, trilogyai); Artificial Analysis publishes output-token volumes per task in its methodology.
[32] FinOps Foundation, State of FinOps 2026 (data.finops.org, 1,192 practitioners): 98% of practitioners now manage AI costs (31% two years earlier); 73% of organisations exceeded their AI cost projections, citing underestimated token volumes, the hidden tokens of reasoning models, and agentic workflows.
Frequently asked questions
- Is the model race really over?
- The race continues, but it no longer produces a lasting winner, and that is the whole nuance. On the LMArena text leaderboard of 12 August 2026, nine Elo points separate first place from fifth, and first and second are in statistical tie; the top spot has changed hands with every major release since November 2025. Meanwhile the cadence of incremental versions has accelerated sharply (roughly six to seven weeks between OpenAI versions in 2026, against five months in early 2025), while the genuine generational leaps keep slipping: Gemini 3.5 Pro announced in May and still absent in mid-August, Grok 5 postponed, DeepSeek R2 never released, Llama 4 Behemoth abandoned. Many versions, few breakthroughs: that is the textbook definition of a commoditising market.
- Are AI models released to the rhythm of the stock market?
- Partly, and precision matters about what is actually demonstrated. Deliberate couplings between financial announcements and model launches are documented: Claude Opus 4.8 shipped on the very day of Anthropic's record 65 billion dollar raise (28 May 2026), Grok 5 was officially confirmed inside xAI's fundraising press release (6 January 2026), Moonshot's Hong Kong IPO plan followed the Kimi K3 launch by three days (July 2026), and both Anthropic and OpenAI chained a raise, an IPO filing and a model launch within the same quarter. However, no investigation by a reference outlet establishes that release dates are set by the financial calendar, and counter-examples abound (Gemini 3.5 Pro and Grok 5 missing despite maximal financial pressure). The financial calendar serves as the stage and the amplifier of releases, and models serve as valuation arguments, with no proof to date that the market sets the dates.
- Have open-weight models caught up with the paid ones?
- On scores, very nearly. As of 17 August 2026, three points separate the best closed model (Claude Opus 5, 63) from the best open one (Kimi K3, 60) on the Artificial Analysis index, and OpenRouter measures a stable lag of three to six months sustained for over eighteen months. Open models even take first places on public leaderboards (GLM-5.2 topping Design Arena in mid-July). Two caveats: closed models remain more robust at equal scores, especially on long agentic tasks; and adoption has not followed performance, with the open-source share of American enterprise API usage falling from 19% to 11% between 2024 and 2025 according to Menlo Ventures, while open distribution became massively Chinese.
- Why is my API bill rising while listed prices fall?
- Because the price per million tokens is no longer the cost. Reasoning models bill their thinking tokens as output: three to five times the visible volume in ordinary use, ten to thirty times in extreme cases. Some providers charge a long-context premium (double beyond 200,000 tokens at Google and xAI, and OpenAI introduced a comparable premium in 2026, just as Anthropic removed its own). And volume explodes with agents. The net result was measured by the FinOps Foundation in early 2026: 73% of organisations exceeded their AI cost projections, even though 98% of practitioners now say they manage those costs. The discounts exist (around 90% on cache reads, 50% for batch processing) but they require an architecture designed to capture them: we have published a dedicated guide to token budgets and routing.
- Should I wait for the next model before launching a project?
- No, and that is precisely what commoditisation changes. Waiting for 'the right model' made sense when each generation created a lasting gap; there is no lasting gap any more, the top spot rotates every two months and the mid-range rises faster than the frontier. The useful decision now concerns reversibility rather than picking a champion: an architecture that routes requests to several models depending on the task, evaluations run on your own use cases rather than public benchmarks, and a second supplier already tested. The defensible value moves to what the model does not supply: your data, your processes, your ability to verify outputs. The step after that, made affordable by open weights, is specialising a model on your own know-how to give the AI your company's personality.
- Should you fine-tune an open model on your company's data?
- In sequence, not straight away. The first step is injecting your know-how into the context of any good model: a proper knowledge base, a charter, decision rules, reference examples. It covers most uses and stays reversible. Specialising an open model (fine-tuning, or distilling a small dedicated model) makes sense next, on a stable, narrow, high-volume domain, with a sovereignty benefit attached: the model and the corpus stay with you. Two safeguards: large accounts have cooled on systematic fine-tuning (a16z, January 2026), a sign that the technique should be chosen case by case; and a specialised model has to be served, monitored and updated, which carries its own operating cost. Whichever route, the durable asset is the same: a clean business corpus and explicit rules. Techniques come and go; encoded know-how stays.
- Is this article investment advice?
- No. It is educational and informative: it documents, with sources, the state of competition between AI models as of mid-August 2026 (benchmark saturation, release calendar, open-weight convergence, prices). It constitutes neither investment advice nor a recommendation to buy or sell any security. Several figures cited are press estimates, aggregator readings or vendor self-reported measurements, flagged as such in the text. For any investment decision, consult an authorised financial adviser.
Paul-Antoine Tual
AI Transformation Leader · Junyr Method™ · Transition manager specialising in AI for French SMEs and mid-caps. Engineer from the École des Mines de Nantes, lawyer, developer since 1993.