Skip to content
All guides
AI Litmus5 min read30 July 2026

Why most corporate AI pilots show no measurable return

By Shobhit Khandelwal·Founder, VMS Culture Labs
Why most corporate AI pilots show no measurable return
The short answer

Most corporate AI pilots show no return, and it is almost never because the tools are bad. Over 80 percent of companies already have them. The value stalls because companies measure the wrong things: licenses bought, tools deployed, courses completed. None of those tell you whether the work got better. The pilots that pay back measure fluency, whether the actual output on a role's real tasks became faster and more trustworthy, and tie that to a rupee figure. That is the entire premise behind AI Litmus.

The divide

Two numbers, from two independent sources, describe the same expensive gap.

95%of enterprise AI pilots delivered no measurable return, MIT, 2025
30%of generative AI projects expected to be abandoned after the pilot stage, Gartner

The first is from MIT's The GenAI Divide: State of AI in Business 2025, built on 150 leadership interviews, 350 employee surveys, and 300 public AI deployments. Only about 5 percent of pilots moved profit and loss at all. The second is Gartner's July 2024 forecast, which put at least 30 percent of generative AI projects on track to be dropped after proof of concept by the end of 2025, with unclear business value named as a leading reason.

If your team ran an AI pilot this year, the base rate says it is far more likely to have shown you nothing on the number that matters than to have paid for itself.

It is not a tools problem

The easy story is that the tools were not good enough. The research says the opposite. In the same MIT work, more than 80 percent of the organizations had already explored or piloted tools like ChatGPT and Copilot, and roughly 40 percent had deployed them. Access was never the bottleneck.

MIT traced the divide to organizational learning gaps and weak integration into real work, not to the quality of the AI model.

Read that again, because it is the whole point. The gap between the 5 percent that got value and the 95 percent that did not was not a better model. It was whether the people using it actually got good at using it, on the tasks their job is made of. That is fluency, and almost nobody is measuring it.

What the idle licenses cost you

A license is a cost the day you buy it and a return only if it changes how the work gets done. Those two things are drifting apart. Enterprise surveys reported by Fortune put weekly use of licensed Microsoft Copilot seats at just 20 to 30 percent. You are paying for every seat and getting the benefit from a fraction of them.

For whoever owns enablement, that is the uncomfortable line item. Leadership approved the spend and the workshops on the promise that the team would get more productive. When the review comes, activity dashboards and attendance sheets will not answer the only question being asked, which is whether the work got measurably better and by how much in rupees.

  • Seats bought tells you what you spent, not what you got back.
  • Course completions tell you who sat through training, not who can now do the job faster.
  • Tool logins tell you who opened the app, not who produces better output because of it.

Why buying more or training harder does not close it

The two reflexes when a pilot stalls are to buy a better tool or run another workshop. Both treat the symptom. If the problem is that fluency never formed, a new tool inherits the same gap on day one, and another generic workshop adds another completion certificate to a pile that already failed to move the number.

Completion is not fluency. Someone can finish every module and still open a blank prompt on Monday with no idea how to make the tool do their actual job, while a quiet colleague two desks over has quietly rebuilt half their week around it. Until you can tell those two people apart, more spend just widens the gap you cannot see.

Measure fluency, not access

The pilots that pay back share one habit: they measure the work, not the rollout. Not how many seats, not how many trained, but whether the output on a role's real tasks got faster, cleaner, and more trustworthy, per person and per team.

That is exactly what AI Litmus does. It reads fluency on the tasks a role is actually made of, shows you who is ahead and who is stuck and where the specific gaps are, and maps the gap to a rupee figure your leadership can act on. Not a completion rate. A picture of whether your AI spend is turning into better work, which is the number the pilot was supposed to prove in the first place.

See this on your own teams.

A private walkthrough, calibrated to your roles. About two weeks.

Frequently asked

Why do most corporate AI pilots fail to show a return?

The most cited evidence is MIT's The GenAI Divide: State of AI in Business 2025, which found about 95 percent of enterprise AI pilots delivered no measurable profit-and-loss impact, while only around 5 percent did. The report traced the difference to organizational learning gaps and weak integration into real workflows, not to the quality of the AI models. In other words, the pilots stalled because the people using the tools never became fluent on their actual tasks, not because the tools were inadequate.

If employees already have the AI tools, why is productivity not improving?

Because access is not fluency. In the MIT research more than 80 percent of organizations had already piloted tools like ChatGPT and Copilot and about 40 percent had deployed them, yet the value still stalled. Separately, enterprise surveys reported by Fortune put weekly use of licensed Microsoft Copilot seats at only 20 to 30 percent. Buying a seat or finishing a course does not mean a person can make the tool do their job better, and until that skill forms, the spend does not convert into output.

How do you measure AI fluency instead of just tracking licenses or course completions?

You look at the work rather than the rollout. Licenses, logins, and completion rates measure inputs; they do not tell you whether a role's real output got faster or more trustworthy. Measuring fluency means assessing whether a person can apply AI to the specific tasks their job involves, comparing that across people and teams, and expressing the gap as a value figure. That is the approach AI Litmus takes, which is why it can tell a team that is genuinely fluent apart from one that simply holds a lot of licenses.

Shobhit Khandelwal
Shobhit Khandelwal
Founder, VMS Culture Labs

Shobhit Khandelwal is the founder of VMS Culture Labs, on a mission to measure what most leaders only guess at: how fluently their teams truly work with AI, and the hidden cost of how people behave at work. He is out to replace workplace guesswork with evidence, and build the kind of workplaces the next generation deserves.

Connect on LinkedIn
Get the intelligence briefing

Occasional, high-signal notes on measuring AI fluency and culture cost. No spam.

Keep reading