The Frontier Is Getting Expensive. The Rest of Us Are Catching Up.
I was half-watching one of those AI-model-leaderboard breakdown videos this week, the kind where an enthusiastic narrator scrolls through the Artificial Analysis Intelligence Index like it's a football table, when a detail stopped me mid-coffee. Grok 4.7 had just landed in fifth place, a few points behind the top three frontier models. Sitting right above it, at fourth place, was something called Muse Spark 1.3 Max. Open source. Open weights. Nearly indistinguishable, on the benchmark that mattered, from a model backed by a company burning through capital at a rate that would make a small nation's central bank nervous.
Right behind Grok 4.7 sat GLM 5.3 Max and Kimi K3, both open weight, both described in the same breath as "phenomenal." Four models, clustered within a handful of points on the leaderboard, and two of them cost nothing to download, fine-tune, and run on hardware you already own.
This is not a fluke. It's the shape of the next five years, and if you're making infrastructure decisions on the assumption that frontier labs will always be a generation ahead, you're planning against a trend that's already reversing.
The Bit Everyone Skips Past
Buried in the same leaderboard discussion was a smaller, less flattering detail about Grok 4.7: its gains apparently come with substantially higher token usage. It's doing more work, generating more tokens, to get to roughly the same place as models that use fewer. Which means the cost per completed task is higher too, an inversion of what you'd hope a new model generation would deliver. Then there's the context window: 500K tokens for Grok 4.6 and 4.7, against a million for most of its frontier rivals. So while it climbs the intelligence leaderboard, it's simultaneously falling behind on the boring, unglamorous metric that actually determines whether you can use the thing for a real job: how much of your codebase or your case file it can hold in its head at once.
None of this is a knock on Grok specifically. It's a symptom. When a frontier lab needs to burn more compute, more tokens, and accept narrower practical limits just to inch up a leaderboard by a handful of points, that's the law of diminishing returns introducing itself politely, before it kicks the door down.
Diminishing Returns Is Not a Metaphor Here
Scaling laws for language models were never a promise of infinite headroom. They were an observed relationship, more data plus more parameters plus more compute yields predictably better performance, and predictable relationships eventually run into the wall that all predictable relationships run into: the marginal unit costs more than the last one bought you.
We are watching that wall get built in real time. The publicly available high-quality text corpus is not growing exponentially; it's roughly fixed, and frontier labs have already ingested most of what's worth ingesting, which is why synthetic data generation and increasingly exotic RL environments have become load-bearing parts of every major lab's training pipeline rather than a nice-to-have. Compute keeps getting thrown at the problem, but each additional order of magnitude of FLOPs buys a smaller slice of benchmark improvement than the order of magnitude before it. This is not speculation, it's the visible pattern across every major release cycle since 2024: bigger training runs, smaller jumps, and an increasing reliance on post-training tricks (better RLHF, better tool-use scaffolding, better reasoning traces) to squeeze out gains that used to come from scale alone.
Meanwhile the cost of being at the frontier keeps compounding. Training runs that cost tens of millions two years ago cost hundreds of millions now, inference costs for reasoning-heavy models are eye-watering (that $2,300 figure to run Grok 4.6 through a single benchmark suite should tell you something about what it costs at scale), and the talent required to squeeze out the next few points of benchmark performance is some of the most expensive labour on the planet. You are paying an exponentially increasing price for a linearly, then sub-linearly, increasing return. Eventually that trade stops making sense, for anyone except a company with a strategic reason to hold the "world's smartest model" crown regardless of unit economics.
What Open Weights Get to Skip
Here is the asymmetry that matters. Frontier labs are trying to move the entire curve outward, to find capabilities that have never existed in any model. Open-weight labs, DeepSeek, Alibaba's Qwen team, Moonshot (Kimi), Zhipu (GLM), Mistral, Meta when it's feeling generous, are mostly trying to catch up to a curve that already exists and has been extensively mapped by the frontier labs' own published papers, leaked training details, and the simple fact that once a capability is known to be achievable, reproducing it is dramatically cheaper than discovering it.
This is not a new pattern. It's the oldest pattern in technology diffusion: the innovator pays full price to discover that something is possible; the fast follower pays a fraction of that price to figure out how. Pharmaceutical patents exist precisely because of how catastrophically lopsided this asymmetry is. There is no patent regime protecting "the general shape of a transformer architecture trained with RLHF," so the fast followers in AI are following faster than in almost any other industry in living memory. DeepSeek's V3 and R1 releases in early 2025 already demonstrated that a team with a training budget in the single-digit millions could produce a model competitive with releases that cost a hundred times as much. Eighteen months on, that gap in relative capability has narrowed further, not widened, and the leaderboard I mentioned at the top of this post is the receipt.
Add to that the fact that open-weight teams don't need to discover the next paradigm to stay relevant. They need to be good enough, cheap enough, and controllable enough, and for the overwhelming majority of real commercial use cases (customer support, coding assistance, document processing, the unglamorous 90% of what businesses actually pay for), "good enough at a tenth of the cost, run on your own hardware, with no data leaving your premises" beats "marginally smarter, ten times the price, running on somebody else's GPUs" every single time procurement gets involved.
Why the Frontier Can't Simply Outrun This
The tempting counter-argument is that frontier labs will just keep running faster, and the gap, even if fast followers are closing it, stays roughly constant because both sides are improving. That would be true if the cost of staying at the frontier were constant. It isn't. It's convex. Each additional point of benchmark performance costs meaningfully more than the last one, while each point of catch-up performance costs meaningfully less than the frontier lab paid for the equivalent capability, because the fast follower isn't paying the R&D tax on the blind alleys, the failed architectures, and the compute spent on ideas that didn't pan out.
Put those two curves on the same chart and they have to cross. Not might. Have to, given current dynamics, unless something changes the underlying economics (a genuine new paradigm that open-weight labs can't quickly reverse-engineer, for instance, or an regulatory moat that prevents weight release entirely). Absent that kind of structural break, the frontier keeps getting more expensive to defend while the following pack keeps getting cheaper to keep up with it, and the distance between them shrinks with every cycle. We are not at the crossing point yet. We are close enough to see it from here, which is more than could be said eighteen months ago.
A model sitting one or two leaderboard places behind the frontier, at a tenth of the inference cost and available to self-host, is not "behind" in any sense that matters to a CTO signing off on an infrastructure budget. It's simply the better trade.
What This Means If You're Actually Building Something
I've spent thirty years watching "the leading vendor's proprietary technology is unassailable" turn out to be a temporary state rather than a permanent one, and this pattern rhymes with all of them. A few practical takeaways, since I try not to leave a post entirely in the abstract:
Don't architect yourself into a single frontier API. If your product depends on being 3% smarter than everyone else on a general benchmark, you are building on sand that a Chinese lab with a much smaller budget will quietly wash out from under you within a year, possibly less. Build the abstraction layer that lets you swap the model underneath without a rewrite.
Watch the open-weight release cadence, not just the frontier one. Qwen, GLM, Kimi, DeepSeek and their peers are shipping meaningful upgrades every few months, and each release closes more of the gap than the previous one opened. That trend line is more informative for planning purposes than any single leaderboard snapshot.
Self-hosting stops being a compromise and starts being the sensible default for a growing share of workloads, the moment an open-weight model clears whatever capability bar your specific task requires. For a huge amount of enterprise work, that bar is already cleared. It was cleared eighteen months ago for anything short of frontier-grade reasoning, and the bar keeps rising to meet more use cases every quarter.
Treat frontier capability as a temporary lead indicator, not a permanent moat. The labs racing to stay ahead know this too, which is precisely why the economics around subscription pricing, enterprise lock-in, and platform bundling have all been tightening. That's not confidence. That's a business model responding to a threat it can see coming.
The narrator in that video I mentioned said he hoped Meta kept pushing open weights until they eventually hit the absolute frontier. It's a nice sentiment, and I share it, but I don't think it requires hope so much as patience. The economics were always going to bend this way. Diminishing returns don't care whose logo is on the model card, and the gap between "the best model that exists" and "the best model you can actually afford to run at scale" has been narrowing for two years straight. Somewhere in the next few cycles, on some benchmark, for some workload that matters to you, that gap is going to close entirely. Best to have the swap-in-place architecture ready before it does, rather than discovering the hard way that you built your entire product on the assumption that somebody else's expensive lead would last forever.