In short
- Microsoft cut internal Claude spend by over a third. Meta cut Claude Code use roughly in half. Neither stopped the work.
- In the same year, Anthropic went to 40% of enterprise LLM API spend and 54% of coding spend. The headline and the market point in opposite directions.
- Both companies ship their own models now. Moving your engineers onto your own model is a product decision before it is a cost decision.
- Routing cheap work to cheap models is sound, and it is not where enterprise AI is failing. The gains are real but concentrated: a few people get enormously faster, and an average across everyone else shows up as nothing in a P&L.
- A $200 seat is subsidised in a way an enterprise contract is not. Enterprise deals discount the API price, not the seat, which is why engineers buy their own.
A link went round our Slack on a Monday: Meta and Microsoft are scaling back internal use of Claude. The thread that followed was better than the article, and it came down to a disagreement about what the story was even about.
The reporting itself is solid. The Information found that Microsoft cut internal Claude spending by more than a third and is cancelling most Claude Code licences across its Experiences and Devices division, pointing engineers working on Windows, Microsoft 365, Outlook, Teams and Surface at GitHub Copilot CLI instead. Meta cut Claude Code usage roughly in half. The reason given in both cases is cost. The tools got popular, and the token bill followed.
The reading most people took from it is the one I keep hearing inside other companies too: the spend is hard to justify when nobody can point at what it bought. That is a real problem and I will come back to it, because it is the part of this story that actually applies to everyone.
But cost is the reported reason. It is rarely the whole one.
Two of Claude’s biggest internal users cut back in the year enterprise spend on Claude went up
If Claude were losing on the merits you would expect the market to move with it. It moved the other way. Menlo Ventures put Anthropic at 40% of enterprise LLM API spend in 2025, up from 24% in 2024 and 12% in 2023.
Share of enterprise LLM API spend. Anthropic and OpenAI swapped places somewhere in 2024, and nothing in the Microsoft or Meta story bends that back.
Source: Menlo Ventures, 2025: The State of Generative AI in the Enterprise. Remaining share sits with Meta, Cohere, Mistral and others.
Foundation model API spend reached $12.5 billion for the year. In coding specifically, the category this story is actually about, Anthropic held 54% against OpenAI’s 21%.
So the two facts sit side by side: a handful of very large customers reduced internal usage, and the category kept consolidating toward the same vendor. Both are true. Treating the first as a verdict on the second is the mistake.
When a company builds its own model, its own engineers are the training set
Here is what I think is actually happening, and it is not mainly about the bill.
Meta is no longer only a customer. Meta Superintelligence Labs introduced Muse Spark in April 2026, shipped Muse Code as a terminal coding agent alongside 1.2, and released Muse Spark 1.3 in September, priced at $1.25 per million input tokens and $4.25 per million output tokens. A company shipping a coding model needs its own engineers on it, for two reasons that have nothing to do with the invoice. The first is data: your engineers at work are the highest-quality trace data you will ever get for a coding model, and you cannot collect it while they are using somebody else’s. The second is simpler. You cannot credibly sell a coding model that your own engineers decline to use.
There is a contractual edge to this too. Anthropic’s commercial terms say a customer may not access the services to build a competing product or service, including to train competing AI models, and every major provider has a version of that clause. None of this means anyone has broken it. It means that once you are shipping a rival model, the tidiest position is the one where the question never has to be asked, and moving your engineers off is how you get there.
So this is a product decision before it is a procurement one.
Microsoft is in a more awkward spot than Meta
Microsoft is a genuinely mixed case. Its customer-facing AI products lean heavily on OpenAI, while internally a lot of its developers reached for Anthropic and kept reaching. The reporting says Claude Code had become, in the words of one account, a little too popular. That is not a procurement failure. That is a preference showing up in a cost line.
Microsoft would like a home-grown model carrying that load, and so far nothing it has built has closed the gap. So what reads as a retreat is better described as a reallocation: the same engineering work moved to a cheaper tool, not stopped. My guess is that the enterprise-scale contracts get trimmed first and individual seats stay, because for most organisations a per-head subscription is cheaper than a negotiated deal at volume.
There is a word for what both of them are doing, and it is dog-fooding. Meta and Microsoft are moving their own engineers onto their own tools, which is the oldest way there is of finding out whether the thing you are selling is any good. Reported as a retreat from a vendor, it is mostly a company eating its own cooking.
Cutting a bill is not the same as cutting the work. Those get reported as one thing and they are not.
Most of your AI work does not need your most expensive model
The part of this that generalises is the routing, and here the economics are genuinely on your side. Stanford HAI’s 2025 AI Index found the cost of querying a model at roughly GPT-3.5-level performance on MMLU fell from $20 per million tokens in November 2022 to $0.07 by October 2024. That is about 280x in 18 months, and the capability floor has kept rising since.
What that buys you is a Pareto curve to ride. Summarising, classifying, extracting, tagging and first-draft writing clear the quality bar on cheap models now, and they are the bulk of what a non-engineering organisation actually does with AI. Engineering is the exception and will stay the exception for a while. One thing worth correcting about how these get described: GPT-6 Sol sits in the middle of OpenAI’s own lineup, which is not the same thing as being a mid-tier model. Turned up to its maximum settings it goes toe to toe with anything out there. The frontier is crowded now and the models sitting on it are all good, which is itself part of why the pricing behaves the way it does. If you are paying frontier prices for work a cheap model handles, that is worth fixing, and it is the one lesson from the Microsoft story that transfers cleanly.
It is worth being concrete about what sits where, because “use a cheaper model” is not a plan. This is the split I work to when I look at a questionnaire workflow:
| The work | Usual default | What it actually needs |
|---|---|---|
| Reading a questionnaire into structured questions | Frontier model | Cheap. It is parsing, not judgment. |
| Drafting a first answer from approved sources | Frontier model | Mid-tier. The constraint is the source, not the reasoning. |
| Deciding an answer needs a human | Nobody decides | A classifier and a threshold you set |
| Reconciling two approved answers that disagree | Frontier model | Frontier. This is the judgment call. |
| Production code | Frontier model | Frontier, still |
Two of those five need the expensive model. Getting that ratio right is most of what the Microsoft story is actually about, and it is the part you can act on this quarter.
That table is not theoretical for me. It is how we route inside Tribble, and it is why we are not tied to a single provider. The bill is the smaller half of that decision: a system that cannot change models cannot take the next 280x when it arrives, and on this evidence another one is coming.
Routing is not free, though. Somebody has to own the decision about which work goes where, and be willing to be wrong in public when a cheaper model ships a worse answer to a customer. Teams that skip that ownership step end up with a routing layer nobody trusts and a quiet drift back to the expensive default.
The spend is not what is failing. The measurement is.
Now back to the point about justifying the spend, because it is the one that matters most.
MIT’s Project NANDA reviewed more than 300 public initiatives, interviewed 52 organisations and surveyed 153 senior leaders for a report called The GenAI Divide. It found that 95% of enterprise generative AI pilots produced no measurable P&L return. The authors put the gap down to brittle workflows, weak contextual learning and poor fit with how people actually work, not to model quality and not to price.
Gartner forecasts that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.
Read those next to each other and the cost story inverts. A pilot that returns nothing is expensive at any token price. A deployment that takes four hours off a process somebody already reports on every month is cheap even at frontier rates. The question was never cost per token. It was whether anyone could tell you what the thing was worth, and in 95% of cases nobody could, because nothing was being measured before the pilot started.
There is a second thing under that 95% and it is not measurement, so I want to be careful not to oversell the measurement point. The gains are real. They are just concentrated. In a lot of companies a handful of teams or individuals have got genuinely, enormously faster, and the budget gets approved because somebody watched that happen and wants it for everybody. That is usually what is behind a spend number that looks irrational from outside. Even a 20% lift across a team is worth the money. What those companies are not getting is the lift for every employee, and an average taken across everybody is exactly the shape that shows up as nothing in a P&L.
It is also not an infinite money glitch. There is not infinite cash in the world, and everyone is buying the same capability at the same time. If every company in a market gets faster together, some of that gain is competed away before it ever reaches the bottom line, and the companies involved end up roughly where they started relative to each other. That is a real effect and it is part of the picture. It is not the whole picture either, because we are also seeing organisations that built AI into their product pull away on ARR at the fastest pace any of us have watched.
None of this happens on its own, which is the part people skip. The tool is not magic. It still takes skill, a learning curve, and a fair amount of wrangling, and the distance between the person getting extraordinary results and the colleague sitting next to them getting very little is mostly that wrangling.
That gap is the one Tribble Respond is built to close, and it is also why it points at questionnaires and not at knowledge in general. An RFP has a clock on it, a queue behind it, and a reviewer who already knows what share of answers they had to rewrite. The baseline is sitting there whether or not anybody has called it one, which means the before and after are arguable in a way that “the team feels faster” never is.
The shadow AI economy is the actual forecast
The same MIT report found something that gets quoted less and predicts more: staff at over 90% of companies use personal AI tools for work, while only 40% of companies provide official access.
That is the end state I would bet on for organisations where the product is not built by engineers. Not a large enterprise agreement. A budget line, a sensible list of approved tools, and people picking what suits the task. It is cheaper than a negotiated contract and it is already how most of these companies operate, whether or not anyone has written it down.
A $200 seat is a better deal than your enterprise contract
Here is the part the coverage keeps missing, and it is why I think the subscription end state is close to inevitable rather than merely likely.
Frontier subscriptions are heavily subsidised. A heavy user on a $200 seat gets through a volume of work that would run to many thousands of dollars a month at list API prices, and the seat does not reprice when they do it. That is not an accident, it is the point of the seat.
Now hold an enterprise agreement up next to it. Enterprise deals discount the API price. They do not discount the $200 seat. So you can negotiate your API rate down by half, be pleased with yourself, and still be nowhere near what a seat already hands one person for $200, because the seat was never priced off the same curve you just negotiated on.
An engineer who expenses their own subscription is not going around procurement. They are doing arithmetic.
The adoption path is just as predictable, and I have watched it run end to end more than once. Somebody tries the free tier. They hit the ceiling. They move to the $20 tier and burn through that in a day. The next day they are on $200. Then they show the person next to them what they just did with it, and it spreads sideways across a team rather than downwards from a decree. It expenses cleanly, so nobody has to run a procurement cycle to find out whether it was worth having.
Giving every engineer a $200 seat is almost always cheaper than an enterprise plan, and it arrives a great deal faster.
What a product can add on top of that is the part nobody wants to do themselves. We have already worked out the routing, the cost per token, which tier each step needs and what to do when an answer is not good enough. What a team gets from Tribble is the finished workflow, not a bill to understand.
It also has a cost that does not show up on the invoice. When the work happens in personal accounts, the answers people depend on live in someone’s chat history. There is no source on them, no approval, no record of who decided what, and no way to correct an answer once it is wrong in forty places. That is fine for drafting an email. It is not fine for the answer you send a regulator, a security reviewer or a customer.
That gap is what the Brain is for. Every answer Tribble gives carries where it came from, who owns it, and who was allowed to see it, so “who approved this, and when” is a question with an answer. A personal chat account cannot do that, and it is not trying to.
Which is where this stops being a story about model vendors.
The teams getting a return picked work that already had a number on it
The deployments I see produce a measurable result are not the ambitious ones. They are the ones pointed at a process the business was already counting: how long an RFP takes, how many security questionnaires are in the queue, how much of a response needed a specialist to rewrite it. The number existed before the AI did, which is the only reason anyone can tell whether it moved.
That is the work Tribble Respond is built for. A questionnaire lands, Tribble reads the questions out of whatever file they arrived in, drafts from approved sources only, shows the source on every answer, and flags the ones that need a human instead of guessing. It scores how sure it is, and anything below the line you set goes to a person. Customers started calling the governed layer underneath it the Brain before we did.
“I completed 90% of a 200-question RFP in just an hour, and it handles one-off technical questions from sales with incredible accuracy. Now I can focus more time on strategic customer engagement instead of searching for information.”
Shaundra Toy · Sales Engineer, Clari
That is the shape of number I mean. Not “faster”, but 90% of a known workload in a known time, against a process Clari was already measuring.
The number I find more interesting is the other one from the same deployment: 10-20% of security responses still needed a specialist. That is the figure that makes the first one trustworthy. A system that cannot tell you which tenth needs a human has not saved you the work, it has moved the risk somewhere you cannot see it.
Then the second-order thing happens, which is the part I did not expect when I started building this. The repeated questions turn into a signal of their own: what buyers keep asking, where an approved answer has gone stale, what the company is not yet ready to answer. That is what Tribblytics reads. Clari’s team went from treating responses as administrative work to reading them for product gaps. That only works because the answers were governed to begin with, which is the argument of this whole post one layer down.
What Tribble does not do is make the measurement problem go away. If nobody can tell you today how long your RFP process takes or what share of answers get rewritten, that is the first week of work, and it is yours. A tool cannot retrofit a baseline you never had.
The companies in the headline are not telling you Claude stopped working. They are telling you they would rather own the model. You probably would not, and that is the one part of their strategy worth copying: know exactly what the spend is for.