If your company has an AI pilot that impressed everyone in a demo and then quietly stopped being mentioned, you're in good company. Pilots are easy to start and remarkably hard to finish. In 2026 the biggest AI companies launched units whose whole purpose is getting AI out of the demo and into daily operations. Here's what the research does and doesn't say about why pilots stall, the six patterns behind it, and what changes when the engineers building a system work inside the business that will use it.
What the numbers actually say
The figure everyone quotes comes from MIT's NANDA initiative. Its 2025 report, The GenAI Divide, found that only about 5% of the generative AI pilots it studied achieved rapid revenue acceleration, while the vast majority delivered little or no measurable impact on profit and loss (Fortune). It drew on 150 interviews with leaders, a survey of 350 employees and an analysis of 300 public AI deployments.
Treat the headline with care. It rests on interviews, a survey and public deployments rather than measured before-and-after results, and "little to no measurable impact" can simply mean nobody measured. That is the practical lesson: a pilot without a baseline can't prove anything, however well it works.
Two details from the same report matter more than the headline:
- Partnering beat going it alone. Buying from specialised vendors and working with partners succeeded about 67% of the time. Internal builds succeeded only a third as often.
- Budgets went to the wrong place. More than half of generative AI budgets went to sales and marketing tools, while the biggest returns came from back-office automation.
Agentic AI is heading the same way. Gartner predicted in June 2025 that more than 40% of agentic AI projects will be cancelled by the end of 2027, and estimated that only about 130 of the thousands of vendors selling agentic AI are the real thing. It calls the problem agent washing: existing assistants, RPA tools and chatbots rebranded as agents (Gartner). The same analysts expect agentic AI to make at least 15% of day-to-day work decisions autonomously by 2028, up from none in 2024. The technology is coming either way. The question is whether your projects are among the ones that survive.
Six reasons pilots stall
- There was no baseline. Nobody measured how long the process took, how often it went wrong or what it cost before the pilot started. When the pilot ends, the only evidence is enthusiasm, and enthusiasm doesn't survive a budget meeting.
- It was built beside the workflow, not inside it. A chatbot in a separate tab asks people to change how they work, and most won't. MIT describes this as a learning gap: generic tools that don't learn from or adapt to a workflow stall once they move beyond individual use.
- It aimed at the visible work, not the valuable work. Customer-facing demos win approval. Back-office processes, where the hours actually go, get overlooked.
- Data, security and legal arrived last. The prototype ran on a spreadsheet export. Production needs real access, a processor agreement and a privacy review, and that's where timelines die. Among Dutch companies that considered AI but held back, 49% cited privacy and 42% legal consequences (CBS).
- Nobody owned it after the demo. The pilot belonged to an innovation team or one enthusiastic person. When they moved on, so did the system. In MIT's study, the organisations that succeeded gave line managers, not just a central AI team, the job of driving adoption.
- What was bought wasn't what was sold. A rebadged chatbot or RPA script sold as an agent hits its limits the first time a real case goes off script.
What embedded engineers do differently
It's the gap OpenAI, Anthropic, AWS and Microsoft are targeting with the forward deployed engineering units they launched in 2026. Putting the builders inside the customer's organisation addresses most of the list above:
- The baseline starts on site. Watching the work means counting it: volumes, minutes, error rates. The result is measured against those numbers, not against expectations.
- The system goes where the work already happens. Into the inbox, the ERP, the phone line or the booking system people use every day, instead of a new tool they have to remember to open.
- The target is chosen on evidence. Candidates are scored on six questions, from volume and time to data access and the cost of a mistake. We explain the method in which business processes to automate with AI first.
- Privacy is designed in from the first sprint. Data flows, access rights and hosting are settled in the scope, before a line of code, not negotiated after the demo.
- It's tested like software. Every change is checked against a set of real examples, the way we describe in evaluating LLM quality, and the system runs alongside the old process before it replaces it.
- Someone owns it from day one. The people who will run it are trained during the build, with monitoring and a documented plan for when the model is wrong.
We call this production, not pilots: systems that run your business on Monday, not demos that impress on Friday.
What a baseline actually looks like
A baseline sounds like a research project. It isn't. For most processes a week of honest counting is enough, and five numbers cover it:
- Volume. How many cases arrive per day or week: emails, orders, calls, documents.
- Handling time. Minutes per case, timed on a sample of twenty or thirty real ones rather than estimated in a meeting.
- Error and rework rate. How often a case has to be corrected, chased or done again.
- Turnaround. How long a customer or colleague waits between asking and getting a result.
- Who does the work. Which roles spend the time, because an hour of a specialist's day costs more than the timesheet suggests.
Write the numbers down before anyone builds anything, and agree which one the system has to move. That single decision does more for a pilot's survival than any choice of model.
A realistic path from idea to production
- One to three days on site, if the work spans several people or teams. The engineer who will build it watches the work, sets up the baseline and scores the candidates.
- A written scope. One workflow, a fixed price, a date, and the success measure agreed in advance.
- Weekly demos on real data. You see it working, or not working, every week rather than at the end.
- A parallel run. The system works alongside your team, and people check its output until the numbers say it can take on more.
- Deployment and handover. Training, monitoring, and either ongoing support or the keys and the documentation.
Measuring the result is covered in our guide to AI automation ROI. The important part is that the measure is agreed before the build, not invented after it.
Is your pilot heading for production?
Answer these honestly. More than two noes and it probably isn't:
- Do you know what the process cost before the pilot started?
- Does the system work inside a tool people already use every day?
- Has it touched real data, with real access rights?
- Has someone reviewed privacy and the processor agreements?
- Is there a named owner outside the project team?
- Is there a plan for when the model gets something wrong?
- Is there a date for switching off the old way?
When to stop a pilot instead
Not every pilot deserves rescuing. Stopping is the right call when the process turned out to be rare or unpredictable, when the data the system needs can't be reached or shouldn't be, or when people still redo most of the output after weeks of tuning. It's also right when an off-the-shelf product has caught up and now does the job well enough. A clean stop, with the lessons written down, is a better result than a zombie pilot that nobody dares switch off.
A stalled pilot is rarely wasted. It usually tests the idea and exposes the obstacles. What it needs next is an engineer who treats production as the goal from day one, which is how Neurova AI builds custom AI applications. We'll take it to production, or tell you honestly that it shouldn't go. To start, ask for an on-site visit.
Frequently asked questions
Why do most AI pilots fail? Rarely because of the model. Pilots stall when there is no baseline to prove impact, when the tool sits beside the workflow instead of inside it, when data access and privacy are dealt with last, and when nobody owns the system after the demo. These are delivery problems, and they can be designed out.
Is it true that 95% of AI pilots fail? The figure comes from MIT's NANDA initiative, which reported in 2025 that only about 5% of the generative AI pilots it studied achieved rapid revenue gains while most showed little or no measurable impact on profit and loss. It rests on interviews, a survey and public deployments rather than measured before-and-after results, so read it as a warning, not a law.
How do you get an AI pilot into production? Measure the process before you build, choose one workflow with clear volume and rules, build inside the real systems with real data, run it alongside the old way, test it against real examples, and hand it to a named owner with training and monitoring. Plan for production from the first week, not after the demo.