Most AI projects stall. The studies say why.

S&P Global surveyed 1,006 companies and RAND interviewed 65 engineers about projects that failed. The causes they found are older than the models, and a trial can test for each one.

Tyler Gibbs. . 4 minute read.

A monorail car stopped on its elevated track just short of the station, a person waiting on the platform beside a blue signal lamp

A hospital, a bank and a law firm each ran an AI project last year. Each one had a demo that worked. A year later, the people who do the work are doing it the way they did before. Nobody cancelled the project. It stopped somewhere between the demo and the desk.

That story is common enough to have been counted. The counts disagree about the size of the problem, and they agree about its causes.

The number everyone quotes is the weakest one.

You have probably seen that 95% of AI projects fail. The figure comes from a July 2025 report by MIT’s NANDA project, which says 95% of organizations are getting zero return. The report labels its own findings preliminary. It rests on interviews at 52 organizations, 153 survey responses collected at four conferences, and a review of public announcements. It calls its figures “directionally accurate,” and it states the 95% three different ways: of organizations, of pilots, and of custom tools.

Read it as a signal that something is wrong. Do not plan a budget around it.

The better counts still show most projects stopping early.

S&P Global surveyed 1,006 IT and business leaders in North America and Europe and published the results in May 2025. The share of companies abandoning most of their AI projects before production rose from 17% to 42% in a year. On average, 46% of projects were scrapped between proof of concept and broad adoption.

Where the tools did reach people, the work did not always change. Two economists, Anders Humlum and Emilie Vestergaard, linked a survey of 25,000 Danish workers at 7,000 workplaces to government pay records. Two years after chatbots arrived, they found no effect on earnings or recorded hours in the occupations most exposed, and ruled out effects larger than 2%. It is a working paper and has not been peer reviewed.

The causes are older than the models.

RAND asked why. In 2024 its researchers interviewed 65 data scientists and engineers, each with at least five years of experience, about projects that failed. They found five causes.

The problem was misunderstood. The people building the system and the people who wanted it meant different things, and nobody wrote down what success was.

The data was not there. The organization did not have enough of the right records to train or test on.

The technology came first. The team chose a tool and went looking for a use, instead of starting from a problem a person has.

The plumbing was missing. There was no reliable way to get the data in and run the system where the work happens.

The problem was too hard. Some tasks are beyond what the technology can do, and nobody found out until late.

RAND also repeats an estimate that more than 80% of AI projects fail. That number is one it cites from elsewhere. Its own contribution is the list, and the list would have been true of a software project in 1995. None of the five is about how capable the model is.

A trial can test for each cause before you pay for the rest.

If the causes are known, a first project can be built to find them early, while stopping is still cheap. This is how we run a trial, and each step answers one line of RAND’s list.

We sit with the people who do the task and write down, with them, what a good result is. That is the first cause. We run on your real records, so a data gap shows up in the first week. We start from one task and choose the model last. We run it where it will live, on the systems you already have. And we measure it against the way your team does the task today. If it does not win, we say so, and the project ends there.

Report the result the way you would want to read it.

The studies above are useful because each one says who was asked and how. A result from your own trial should meet the same standard. For one health system, we trained a model to rank patient records by risk and compared its calls with the clinicians’ own. It agreed with them about 93% of the time. The comparison is to the people who do the work, on their records, and the work page says what was built.

A project that stalls has usually skipped one of these steps. The studies do not say AI fails. They say most organizations found out too late which of five ordinary problems they had.