From POC to Production
28 Apr 2026 · RS Management
TL;DR
- Most AI prototypes end at the demo.
- The reasons are rarely technical. Usually it is a missing definition of success, testing on idealised data and nobody named as the person who accepts the result.
- Prototypes that do reach production carry a scale-or-stop decision date written down at the start.
The prototype works. The demo runs without a stumble, the room asks good questions, someone says this will change how the whole department operates. Three months later nobody remembers where the link was. This pattern repeats often enough that it stops being bad luck for one team and starts being a property of how prototypes are run.
How often this actually happens
In July 2024 Gartner forecast that at least 30% of generative AI projects would be abandoned after the POC (proof of concept) stage, the prototype whose only job is to show that an idea can work at all, by the end of 2025.1 The reasons listed were poor data quality, inadequate risk controls, escalating costs and unclear business value. The forecast never received a public reckoning, so we treat it as a description of the mechanism rather than a measurement of the market.
The direction matches what shows up in delivery work: the barrier rarely sits with the model. It sits in the parts of the work a prototype deliberately skipped in order to move faster.
A definition of success written before the start
The cheapest thing you can do for a prototype’s chances takes half a page and happens before the first line of code. One metric, a numeric threshold, an agreed way of measuring it and the name of the person who signs off on the result.
Instead of “let’s see whether AI can handle document verification,” the note reads:
Across 100 cases sampled at random from last quarter, extraction of the key fields must be correct in more than 90% of cases, and average verification time must drop from 14 minutes to under 5.
Without that, the prototype ends in an argument about whether the output is good enough, and nobody wins that argument, because every participant is holding a different threshold in their head.
There is a genuine trade-off here. At the start you often have no idea what threshold is even achievable. In that case, write a provisional threshold together with the date you will revise it. A number you expect to correct works better than no number at all.
Production data, not a sample
Prototypes are usually tested on data someone has already tidied up: clean scans, complete fields, typical cases. Production looks different. A scan rotated 90 degrees, two documents in one PDF, a handwritten note in the margin, a 2019 file in a template nobody uses any more, a field filled in a way no form anticipated.
The practical rule is to sample at random from the real stream and keep the long tail of odd cases in the sample. A set hand-picked by someone who wants to show a working result tells you about that set and nothing more. If the data cannot leave the organisation, the prototype gets built inside the organisation’s environment. Anonymisation and agreeing a lawful basis for processing take weeks, so that thread is worth starting in week one.
Sometimes only synthetic data is available at the start. Treat the first result as a signal about technical feasibility and postpone the conversation about quality until a production sample exists.
A business owner with skin in the game
A sponsor who supports the initiative is not enough. What the prototype needs is an owner: someone whose numbers change if the solution works, who will allocate their people’s time for acceptance testing and who has the authority to change the process.
The test is simple. Whose figures in the quarterly report will look different. If the answer does not arrive within a few seconds, the prototype has no addressee. IT can build the solution and can run it, but it cannot accept a business outcome on behalf of someone who does not need that outcome.
This single point explains a large share of abandoned prototypes. They were built out of technical curiosity, detached from a specific problem someone genuinely wanted solved.
Security from day one
Three questions belong at the start of a prototype rather than at the handover to production. Where does the data go and under which contract. Who has access to logs, prompts and outputs. Who maintains this in 18 months, and out of whose budget.
A 30-minute conversation with security and legal in week one costs less than discovering in week eight that the chosen model provider will not clear a risk assessment. In regulated organisations, add system classification, a data protection impact assessment and documentation that has to survive an audit. Specific obligations are worth confirming with a qualified adviser.
Maintenance gets skipped even more often than security. Model versions change, prompt quality degrades as input data shifts, integrations break. It is reasonable to assume a standing annual budget line for maintenance instead of treating deployment as a one-off cost.
A scale-or-stop decision within weeks
The decision date belongs in the scope document alongside everything else. A six to eight week window works well in practice, with three possible outcomes: scale, fix one specific element and repeat the measurement once, or close it down.
Closing is a normal outcome and worth reporting as a measurement result, without attaching the label of team failure. The worst case is a prototype that drifts for six months until everyone quietly stops mentioning it. Nobody learns anything, and the organisation remembers only that AI did not work here.
When closing, a one-page note earns its keep: what was tested, the result against the threshold, what would be done differently. Three of those notes after a year are worth more than one prototype that never received a verdict.
What this adds up to
All five points cost very little: a few conversations, one page of writing, one data sample and a decision date in the calendar. Skipping them saves a week at the start and routinely wipes out the value of the following three months of work.
It is also worth starting from the problem before choosing the tool. If a successful prototype would not move a number that somebody cares about, the technical result stays a curiosity.
Footnotes
-
Gartner, 29 July 2024 prediction on generative AI projects abandoned after proof of concept: https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025. ↩
RS Management is an advisory practice run by one person. Who stands behind it and with what experience: About.
Blog content is informational and educational. It does not constitute legal or tax advice, nor individual business advisory. The scope of our services is described in the terms.
This topic is covered by the AI Automations & Agents package: scope + quote + build + acceptance + handover to the team.
See the package: AI Automations & Agents