Why AI projects stop halfway, and how to spot it early
A demo has to answer well on chosen examples. A system has to be right on a Monday with real data, be watched, cost a predictable amount, and connect to software built long before anyone said AI. The second is most of the work and it is usually not in the quote.
The gap is not the model
Everyone has access to the same models now. That is exactly why access is not the hard part any more, and why a demo is easy to produce.
What separates the demo from the thing your team uses in March is unglamorous: testing that catches wrong answers before a customer sees them, monitoring so you find out from a dashboard rather than a complaint, control over who can use it and what it can touch, a handle on what it costs to run per month, and integration with systems that were built long before anyone said the word AI.
None of that shows in a demo. All of it decides whether the project reaches anyone.
Agree what 'working' means before anybody writes code
This is the single highest-value thing you can insist on. Before the build, write down what a good answer looks like on twenty real examples from your business, including the awkward ones. That set becomes the test.
Without it, 'done' is an opinion, and the argument at the end is unwinnable in both directions: you think it is wrong too often, the vendor thinks it is fine, and neither of you has anything to point at.
Data first, honestly
Most stalled projects were never a model problem. The data was scattered across a system nobody exports from, or it was inconsistent in ways nobody had needed to care about, or it simply was not there.
The cheapest possible outcome is finding that out in week one. A short paid study that looks at your actual data and says 'this is not ready, here is what would make it ready' has saved you a six-month build. Treat a vendor willing to say that as a signal in their favour.
Give it less power than you think, at first
The safest shape for anything that acts on your systems is: it reads and suggests, a person approves, and it earns the right to act on its own by being right over time.
Alongside that: limits on what it can touch, a threshold above which a human must approve, every action recorded, and every action reversible. This is not caution for its own sake — it is what makes it possible to switch the thing on at all.
Questions to ask before you commit
How will we agree what 'working' means, and who writes that list? What happens when it gives a wrong answer — how do we find out, and how fast? What does this cost to run per month at our volume, and what makes that number go up? Which of our existing systems does it have to talk to, and has that integration been done before? Who is responsible after it goes live, and for how long?
You are not testing technical knowledge with these. You are testing whether the person has done this past the demo.
What people ask us about this.
A study of one to two weeks, then typically six to twelve weeks to a first version people actually use, depending on how much has to connect to your existing systems. Anyone promising a working system in two weeks is describing a demo.
Yes, with smaller models, and it is the right answer where your data is not permitted to leave. It has to be designed for from day one rather than discovered at the end, because it changes which models are realistic and what the hardware costs.
Measure it per task rather than per month, cache what repeats, use the smallest model that passes your tests rather than the largest available, and set alerts on volume. Cost surprises almost always come from a loop nobody metered, not from the price per request.
If this is your problem, start here.
Bring the version of this that is happening in your business.
A senior consultant, not a salesperson. If the answer is short we will just answer it, including when the answer is that you do not need us.
