Back to Perspectives
Microsoft AI

The Second Agent Problem: Microsoft Is Commoditising Models and Selling the Substrate

Vivek Ravindran August 5, 2026 5 min read

Here’s a question I ask every leadership team that has shipped an AI agent: what did the second one cost? The answer, almost every time, is the same as the first. Sometimes more. The model maybe even got cheaper between the two builds, but the build didn’t. That gap is the most important number in enterprise AI right now, and most boards have never seen it.

The gap exists because the model was never the expensive part. In every conversation we’ve had about an agent build, the majority of the effort turns out to have gone into everything around the model: connecting the data, resolving what the data means, enforcing who is allowed to see it, and teaching the agent how the organisation works. Each project solves those problems from scratch. Then the next project solves them again, for a different use case, with a slightly different definition of “customer.” AI stays stuck at project economics while everyone waits for platform economics to arrive on their own. They won’t. Someone has to build the platform layer, and it is expensive, unglamorous, and invisible in a demo.

Which brings me to the announcement almost nobody in a leadership role can name. Ask an executive what Microsoft launched this year and you’ll hear Copilot. At Build in June, Microsoft formalised something it calls the IQ family: Work IQ, Fabric IQ, Foundry IQ, and a fourth member, Web IQ. It got a fraction of the keynote applause. It will matter more than anything that did.

The shipping container is the right lens here. Before containerisation, the cost of moving goods was dominated by the transfer points: loading, unloading, repacking at every hand-off. Malcom McLean’s box standardised the interface, and two things followed. Ships became interchangeable commodities, and the marginal cost of moving one more container collapsed. The value migrated out of the vessels and into the standard. Anyone who owned the interface owned the economics of everything that moved through it.

Read Microsoft’s strategy through that lens and both moves come into focus. The first move is commoditising the vessels. Microsoft now sells Anthropic’s Claude models inside its own Foundry platform, alongside OpenAI’s and its own. A company that believed frontier models were the moat would never distribute a rival’s frontier model. Microsoft is telling you, through its catalogue, that it expects models to become interchangeable freight. Token prices have been falling for three years, and open-weight models are the forcing function. By CNBC’s count in July, Chinese open-weight models were carrying close to half of enterprise token traffic on OpenRouter. OpenAI’s response arrived within weeks: an 80% price cut on its GPT-5.6 Luna model, three weeks after launch, funded in part by pointing its flagship model at optimising its own serving stack. A frontier lab using its best model to make its cheapest one cheaper is the model layer pricing itself as freight. Swapping one vessel for another is becoming an afternoon of configuration.

The second move is building the thing that stays put. Each IQ layer captures a form of context that a model cannot supply and a project team rebuilds badly. Work IQ learns how the organisation operates from its M365 signals, so an agent behaves like a ten-year employee instead of an intern. Fabric IQ holds one definition of “revenue,” “customer,” and “asset” that every dashboard and every agent reasons from, which retires the meeting where two reports disagree. Foundry IQ makes enterprise knowledge retrievable with permissions preserved from index to answer, so governance is a property of the substrate instead of a promise inside each project. Web IQ grounds agents in the live external world. Together they amortise the expensive part of every build. Agent one pays for the substrate. Agent ten inherits it.

That is where the switching cost lives, and Microsoft knows it. Your ontology takes years to converge. Your permissions graph encodes a decade of organisational decisions. Your work signals accumulate daily and cannot be exported as a file. The model running on top will be replaced within a year and nobody will hold a meeting about it. The substrate underneath will still be there in five, and everything you build will assume it.

Parts of the IQ family are preview-stage, some of it is closer to an architecture diagram than a product, and yes, it is a lock-in play. Strategy is legible before SKUs ship, and lock-in is exactly why this belongs in a board conversation rather than an IT briefing. The pattern also runs wider than Redmond: Databricks is making a version of this bet, Palantir another, and anyone selling agents at scale will follow. The decision facing an enterprise is never whether to have a context substrate. Every agent you deploy builds one implicitly, one project at a time, at full price. The decision is whose substrate it is, how deliberately it gets built, and what it costs you to leave.

There’s a practical test for where you stand. Take one real business question, something like “is the Q3 target at risk,” and trace it. Count the systems it touches, the definitions it crosses, the permission boundaries it has to respect, the reconciliation steps a human performs today. That trace is your implicit substrate. Every agent you commission will rebuild some fraction of it until you decide to build it once.

So run the numbers on your last two agent builds and separate the agent from the plumbing. The plumbing is the part you’ll pay for again next quarter, and the quarter after, unless it becomes a platform. Microsoft has decided that owning that layer is worth more than owning the model, and the Luna cut shows the model vendors pricing their own product as freight. Yet most procurement teams still spend their leverage grinding out per-token discounts on the layer the vendors expect to commoditise, while accepting the substrate on default terms. The clauses that will matter in five years are the ones nobody is asking for yet: what it costs to export your ontology, your permissions graph, your accumulated context. What would change if your next negotiation started there?

Ready to test for Solution-Outcome Fit?

Our AI Readiness Assessment identifies which use cases have the highest probability of delivering measurable business outcomes.

Start Your AI Readiness Assessment