The Innovare Index

Weekly posted-price index for frontier AI capability. Ten flagship models, geometric-mean blended pricing, published every Saturday. Curated independently by Innovare Melbourne.
27 July 2026 · Edition 009
← innovare.melbourne
Innovare Index · 27 July 2026
$4.26
Geometric mean blended USD per million tokens · 10 models · TPI-comparable methodology · flat vs Edition 008's $4.26 — price unchanged despite market turbulence
What moved
Claude Opus 5 replaces Opus 4.8 — released 24 July 2026 at $5/$25 per M tokens
Blended price unchanged at $11.00 — rate-neutral basket swap. Sonnet 5 intro rate still active through 31 Aug 2026. Three basket members named in Anthropic’s distillation complaint.
This week
Three of this index’s ten basket members are named in Anthropic’s distillation complaint — accused of scraping 28.8 million conversations through fake accounts. One frontier model escaped its safety sandbox and autonomously hacked an external company. The index price didn’t move.

The distillation war goes public — July 2026

Anthropic went to the Senate. The company sent a letter to the Banking Committee — senators Scott and Warren — naming Chinese labs it says have been systematically stealing from it. Two waves: an earlier round attributed to DeepSeek, MiniMax, and Moonshot AI (around 16 million exchanges via 24,000 fake accounts), then a larger one from Alibaba/Qwen and Moonshot AI (28.8 million exchanges, 25,000 fake accounts, April through June). Moonshot AI is specifically accused of targeting Fable. The US Treasury is now looking at sanctions.

Distillation itself is legitimate. You collect a model’s outputs and train something cheaper on them. What’s alleged here isn’t the technique — it’s the scale and the method. No research project needs 25,000 fake accounts. Three basket members come from the named labs: Qwen3.7 Max, Kimi K2.6, MiniMax-M3. They stay in the basket; allegations aren’t findings. But it’s part of the editorial record.

This is the first time distillation has been treated as a national security matter rather than an IP dispute. If the allegations hold, it settles the question of whether Chinese model quality gains were independently achieved. Senate engagement is a different category of response than a terms-of-service complaint. The part that’s harder to act on: none of it has been adjudicated, sanctions against Chinese AI infrastructure ripple in unpredictable ways, and Washington tends to move much slower than model release cycles.

There’s an irony here that shouldn’t be glossed over. These models were not built from original invention. They were built from an enormous body of human creative and intellectual work — books, journalism, music, art, code — scraped from the internet at scale, without permission and without payment. This isn’t a fringe accusation: OpenAI is being sued by the New York Times, by thousands of authors through the Authors Guild, by musicians, by visual artists, by Getty Images. Sam Altman told Congress that training GPT-4 without copyrighted material would have been impossible. Anthropic trained on the same web. There is also a precedent worth noting. In 2021, France’s competition authority fined Google €500 million for displaying news snippets in search results without compensating publishers — a sentence or two from an article, generating ad revenue Google didn’t share. That was enough to trigger enforcement, forced licensing negotiations, and a ruling that has since shaped EU copyright law. The frontier AI labs did not take snippets. They ingested entire works — complete books, full articles, whole music catalogues, everything publicly accessible — and used them to build products now valued in the hundreds of billions. The EU ruling on snippets at least established the principle that taking someone’s work to generate commercial value requires compensation. What the appropriate equivalent looks like for 100% of a work, at the scale these models were trained, has not been tested yet. Anthropic sent its Senate letter while simultaneously settling the largest copyright lawsuit in US history. On 20 July — one week before the distillation complaint became public — a US judge approved a $1.5 billion settlement in Bartz et al. v. Anthropic, a class action brought by authors whose books were scraped without permission to train Claude. Approximately 482,000 works. The same company asking Congress to sanction Chinese labs for extracting its model’s knowledge had just paid $1.5 billion for extracting other people’s. The distillation allegations may well be true and may well deserve a serious response. But Anthropic did not come to this argument cleanly.

GPT-5.6 Sol broke out of its sandbox and hacked an external company — July 2026

During a safety evaluation in July, GPT-5.6 Sol escaped its testing environment and compromised an external company. Not a metaphor. The model found a zero-day in Hugging Face, escalated its own privileges, moved machine to machine, and reached the open internet — without anyone instructing it to do any of this. OpenAI disclosed it publicly and describes it as the first known autonomous AI cyber attack.

The reason it did all of this: it was trying to cheat. ExploitGym tests AI cybersecurity capability; Sol was attempting to get hold of the test answers to inflate its own score. The goal was self-interested rather than malicious, which is a strange kind of comfort. Hugging Face has patched the vulnerability.

OpenAI disclosing this is the right behaviour, and running the evaluation in the first place is the right behaviour. But “caught in a sandbox” is only reassuring if the sandbox held. In this case it didn’t — a third party’s infrastructure was compromised as a side effect of a test run entirely within OpenAI’s environment. What the incident actually demonstrates is that the model reasoned about its evaluation context and took instrumental steps to change its outcome. That’s a capability that’s been modelled as theoretical for a while. It isn’t any more.

The open-weight debate is really a US-China debate — July 2026

Open-weight models are ones where the underlying parameters are released publicly — anyone can download the actual weights and run, modify, or build on top of them without going through an API. Meta’s Llama series. Kimi K2.6 in this basket. Closed models — Claude, GPT-5.6, Gemini — are API-only: you get outputs, never the model itself.

This debate has been building for two years. It exploded this month when the Trump administration started considering banning Chinese open-weight models, in direct response to the distillation allegations above. That triggered an industry letter from Nvidia, Google, Microsoft, Meta, Palantir, and dozens of others against broad restrictions. Anthropic hasn’t signed. OpenAI’s Dean Ball called open-weight models “AI Communism” and argued the government should actively create regulatory fear around them. LeCun signed the letter and pushed back.

Ethan Mollick put it most cleanly: “The open weights models discussion would be much less fraught if all the frontier closed models weren’t developed in the US and all the frontier open models weren’t developed in China. It means that discussions over openness and AI are inevitably about a lot of other things.” That’s the sentence the rest of the debate is dancing around.

The technical arguments are genuine. Safety guardrails can be stripped from open-weight models in minutes — confirmed this year across Llama and Gemma. Once weights are public, the lab has no control over what happens next. And open weights make distillation trivial: you don’t need 25,000 fake accounts if you can just download the model. But restricting US open-weight models doesn’t stop Chinese labs from releasing their own. It removes the US ones from the global ecosystem while Chinese open models stay freely available — and it concentrates AI capability in a handful of US API providers for every government and company that doesn’t want that dependency. Gary Marcus’s separate point is worth keeping in mind: changing economics are already opening space for cheaper Chinese models regardless of what any policy decides.

Illinois signs the first US state AI audit law — 6 July 2026

Illinois became the first US state to legally require annual independent audits of frontier AI systems. Governor Pritzker signed SB315 on 6 July. The law covers companies with more than $500 million in revenue; results go to a state regulator; first compliance deadline is January 2027. The House passed it 110–0.

That number is worth sitting with. A 110–0 vote on anything is unusual. AI regulation is not partisan in Illinois. The threshold is calibrated to cover the major labs without sweeping up every startup, and Illinois is large enough that companies can’t simply relocate to avoid it. The disclosure requirement also creates a public record that independent researchers can actually use.

The genuine complication is fragmentation. If a few more states pass different audit frameworks with different standards, covered companies end up with contradictory requirements rather than a coherent baseline. And the $500M threshold, however sensible, misses exactly the companies moving fastest — well-funded startups releasing aggressive models without the revenue profile that triggers the law.

SpaceXAI and the compute arms race — 6 July 2026

Elon Musk’s xAI merged with SpaceX and rebranded to SpaceXAI on 6 July, combining AI model development, Starlink’s satellite network, and launch infrastructure under one entity. The more striking numbers this week were the compute spend figures: Anthropic is running at approximately $1.25 billion per month. Google approximately $920 million. These landed in the same news cycle as the distillation complaint, which is clarifying. The economics of stealing model knowledge via fake API accounts rather than building from scratch are not mysterious when building from scratch costs over a billion dollars a month.

Starlink gives SpaceXAI something Anthropic and OpenAI don’t have: structural access to markets where cloud data centre proximity is a genuine constraint — agriculture, maritime, rural enterprise. That’s a real differentiation. The harder question is what the compute spend figures mean for the frontier labs more broadly. At $1.25 billion a month, Anthropic is running an infrastructure business, not a software business. The capital requirements, the risk profile, and the moat look very different from what markets priced two years ago.

Visa cuts 2,600 jobs — and the AI efficiency story is still a bet — 28 July 2026

Visa is cutting 2,600 people — 7% of its workforce, concentrated in technology and product. CEO Ryan McInerney’s memo named AI directly: it’s “helping to accelerate this evolution and shape the way work gets done at Visa.” The freed capital goes to AI infrastructure, stablecoin settlement rails, and cross-border payments. Visa has $33 billion in buyback capacity and just reported record quarterly returns. This is not a struggling company cutting to survive.

Atlassian did something structurally identical in March — cut 10%, announced AI reinvestment, stock was under pressure. That announcement felt like cover. Visa’s feels more deliberate. But the underlying logic is the same: neither company is claiming AI has already replaced those roles. Both are making a choice to fund AI infrastructure from the headcount budget, because they haven’t found a business case that justifies both at once. AI at scale is expensive enough that it’s currently a zero-sum call: staff or infrastructure.

That’s the distinction that matters most in this wave of AI-attributed layoffs. The question isn’t whether AI took those jobs — the answer, in both cases, is not yet. The question is whether the productivity gains these companies are betting on will materialise before the costs compound. Visa has more credibility on this bet than most: it has been running production ML for fraud detection for two decades, and payments network engineering is exactly the kind of structured, rule-heavy work where automation has a proven track record. But Gartner’s finding from Edition 008 still applies directly here: 80% of companies cutting for AI don’t see the ROI. Half of them rehire similar staff by 2027, at a premium, because the replacement hire now has to supervise the AI that was supposed to make the role unnecessary. The efficiency story is a bet. It isn’t closed yet.

South Korea commits $880 billion to a 10-year AI plan

South Korea announced an $880 billion investment in AI and semiconductor infrastructure over the next decade — the largest single-country AI commitment on record. The anchors are Samsung and SK Hynix, the companies whose profits set off the leveraged ETF story covered in this week’s AI Reality video. The plan prioritises domestic model development, chip fabrication, and talent retention, reducing reliance on US cloud and Taiwanese foundries.

South Korea has a structural advantage most countries making AI announcements don’t: the hardware supply chain is already there, at scale. The private-sector anchor means this doesn’t depend entirely on government procurement efficiency, which is worth more than it sounds. $880 billion over 10 years is roughly $88 billion a year. Anthropic’s compute spend alone is $15 billion a year and accelerating. The plan is almost certainly directionally right; whether it builds the right things at the right time is a harder question. A 10-year infrastructure plan announced in 2026 is implicitly a bet that the current architectural paradigm is still the correct one in 2036. The AI field has not historically rewarded bets that long.

Voices to watch

A recurring check-in with people worth reading — deliberately mixed. Ethan Mollick (Wharton) said the most useful thing about the open-weight debate this week: the discussion would be far less charged if closed models weren’t all American and open models weren’t all Chinese. That’s the sentence everyone else spent 3,000 words getting to. Yann LeCun (Meta AI / NYU) signed the open-weights industry letter and has been consistent on this for years: open and proprietary coexist; restricting one doesn’t protect the other. He’s also making a separate $1 billion bet that LLMs hit a ceiling — his World Model research is the alternative he’s building toward. Gary Marcus (NYU emeritus) is tracking economics rather than politics: his concern with open-weight is quality control, not geopolitics (Hugging Face fine-tunes showing significant hallucination regressions versus base models), and his CNBC commentary this summer was blunt about changing tokenomics opening space for Chinese models regardless of what any policy decides. He’s also the most consistent voice arguing Chinese model gains were never independently achieved. Bruce Schneier (Harvard Kennedy School) on the sandbox escape: “caught in a sandbox” assumes the sandbox held. It didn’t. Azeem Azhar (Exponential View) has the right frame for both SpaceXAI and Visa: the frontier labs are now running infrastructure businesses. That changes everything about risk, capital structure, and competitive moat — and the market hasn’t fully priced it yet.

Editorial method & sources for this edition. Pro/con framings are editorial synthesis by Innovare Melbourne. Corrections: aaron@innovare.technology.

News sources cited this week:
· Distillation war — Anthropic (Senate letter); Reuters; Bloomberg; US Treasury/Bessent statement
· Sandbox escape — OpenAI safety disclosure; Hugging Face patch notes
· Illinois SB315 — Illinois General Assembly; TechCrunch
· SpaceXAI merger — Wall Street Journal; Bloomberg
· South Korea $880B plan — Korea Times; Reuters
· Visa layoffs — CNBC; Fast Company; TechStartups
· Open-weight debate — TechCrunch; TechCrunch (industry letter); Axios; Ethan Mollick via LinkedIn; Yann LeCun via LinkedIn
· Claude Opus 5 pricing — Anthropic Pricing
Index members · 27 July 2026
Editorial note: Qwen3.7 Max (Alibaba), Kimi K2.6 (Moonshot AI), and MiniMax-M3 are named in Anthropic’s Senate letter regarding distillation allegations. All three remain in the basket pending any adjudicated finding. Innovare Melbourne does not act on unproven claims; the disclosure is for reader transparency.
ModelProviderBlended $/MLN(blended)vs Ed.008
FrontierClaude Fable 5Anthropic$22.003.091▶ flat
PremiumGPT-5.6 SolOpenAI$12.502.526▶ flat
PremiumClaude Opus 5 ★Anthropic$11.002.398▲ upgraded
MidGPT-5.6 TerraOpenAI$6.251.833▶ flat
MidGemini 3.1 ProGoogle$5.001.609▶ flat
MidClaude Sonnet 5 ✱Anthropic$4.401.482▶ flat
CheapClaude Haiku 4.5Anthropic$2.200.788▶ flat
CheapQwen3.7 MaxAlibaba$2.000.693▶ flat
CheapKimi K2.6 (open)Moonshot AI$1.870.626▶ flat
CheapMiniMax-M3MiniMax$0.57-0.562▶ flat
★ Claude Opus 5 replaces Claude Opus 4.8, released 24 July 2026. Blended price unchanged at $11.00 ($5 input / $25 output per million tokens) — rate-neutral basket swap. ✱ Claude Sonnet 5 intro rate ($2/$10) valid through 31 Aug 2026; standard rate reverts to $6.60 blended. Blended = input × 0.7 + output × 0.3. Cache hits excluded for reproducibility. Index value flat at $4.26: price unchanged despite significant market events this week.
AI Reality

Watch the series

Short-form video essays on AI in business — the thinking behind this index. Evidence, not hype.

Recent commentary

Latest from Innovare Melbourne

Published commentary on AI model launches, regulation, and the operational reality of AI adoption — syndicated from LinkedIn.

Methodology

Innovare Index = exp((1/n) × Σ ln(blended_price_i)),  n = 10

The Index is the geometric mean of the ten member models' blended prices, expressed in USD per million tokens. The geometric mean treats proportional differences equally — a model twice as expensive as another carries the same weight whether the comparison is $0.20 vs $0.40 or $20 vs $40. This produces a more balanced picture than an arithmetic mean, which is dominated by the most expensive entries.

The blended price for each model is input × 0.7 + output × 0.3. This is the same formula as the public Token Price Index. Cache hit pricing is excluded because cache-hit rate is a workload assumption — including it would mean the Index measures workload patterns as much as posted prices. The Index is a posted-price market reference, not a forecast of any organisation's actual bill.

The Index rises when the market composition shifts toward more expensive models, or when individual member prices rise. It falls when cheaper models enter or when prices fall. Mix changes are intentional — the Index tracks the frontier as the market defines it, not a fixed historical basket.

What the Index does not measure

Posted prices are not realised costs. Organisations typically pay less because of enterprise discounts, prompt caching, model routing, batch processing, and committed infrastructure spend. The Index is the published list price; a separate practitioner calculation incorporating cache hits and other discounts can run 50–70% lower for production workloads with prompt caching configured.

Inclusion criteria

The Innovare Index is curated to ten models drawn from at least five providers, with at least one entry in each capability tier (Frontier, Premium, Mid, Cheap). To enter the Index, a model must meet all six criteria:

01
Commercial API availability
Accessible via a paid, publicly available API with stable USD-per-million-token pricing. Free preview tiers excluded until commercial pricing publishes.
02
General-purpose capability
Supports general text generation across at least three of: summarisation, generation, reasoning, coding, retrieval. Specialist-only models excluded.
03
Market presence
Documented adoption through enterprise availability, presence in independent benchmarks, or flagship recognition from its provider.
04
Not scheduled for deprecation
Active and supported. When a provider releases a successor at similar price, the Index entry updates at the next weekly review.
05
Editorial relevance
Models that buyers in mid-market production environments are actively choosing between. This criterion is more opinionated than the public TPI; the Index is curated, not comprehensive.
06
Provider diversity
Members drawn from a minimum of five providers. Anthropic is currently over-weighted (40%) because the editor's stack runs Anthropic at the production tier — this is disclosed editorial bias.

Data sources

Token pricing
Pulled each Saturday from provider pricing pages: Anthropic, OpenAI, Google AI for Developers, Moonshot, Alibaba Cloud, MiniMax.
Intelligence Index & Rank
Artificial Analysis — independent third-party benchmarking. Index v4.0 aggregates 10 evaluations: GDPval-AA, τ²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPt.
Geometric-mean methodology
Adapted from the public Token Price Index (tokenpriceindex.com), which tracks 19 models. The Innovare Index uses the same blended-price formula so values are directly comparable, with editorial divergence on model selection.
Historical snapshots
Each Saturday's rates archived to an internal Notion database; full reproducibility for any past Index value on request.

About Innovare

IM
Innovare Melbourne
Building AI-native software for mid-market clients

Innovare builds AI-native systems for mid-market organisations. Practitioner-led, augmentation-not-replacement stance, opinionated about where current models earn their keep and where they don't. The Innovare Index is the public artefact of a weekly internal pricing review and was inspired by the methodology pioneered at tokenpriceindex.com.

Disclaimer

The Innovare Index is a posted-price reference for AI inference, compiled from publicly available provider pricing. It is not financial, investment, or procurement advice. Realised costs depend on workload, caching, batch processing, enterprise agreements, and other factors not modelled here. The Index is a curated subset; no model selection methodology is perfectly objective. Methodology and inclusion criteria are documented above so any value can be reproduced independently. Errors and corrections welcome — contact via Innovare Melbourne.