Skip to main content

Dev Station Technology

AI Agents in Production: What Breaks After the Demo

TL;DR

  • A demo quietly removes the four things that make production hard: curated inputs, a friendly operator, no deadline, and nobody who has to answer for a wrong answer.
  • Five things break on the way to production: field conditions, review capacity, ownership of errors, drift, and a process that never actually changed.
  • More output does not help when review capacity stays flat. Review becomes the new constraint, and plausible-but-wrong is the failure mode that costs the most.
  • We shipped five AI features into inspection apps this year and removed three. The removals carry the argument.
  • An agent earns its place when a person can explain and defend what it produced, not when it produces the most.

01

The demo is not a lie, it is a different system

Every agent demo works. That is what makes them hard to argue with in a buying meeting.

The demo is real code doing real work. What it is not is the same system you will run in six months, because the conditions that made the demo clean are exactly the conditions production removes.

This is written from building AI features into field inspection software, where the output ends up inside a document somebody signs. The pattern holds in most B2B work: the model was never the hard part.

The cost of learning this late is the part worth planning around. A feature that fails in a demo costs an afternoon. The same feature failing after rollout costs the trust of the people who have to use it every day, and that is considerably harder to buy back.


02

Four things a demo quietly removes

Curated input. Demo data is the data somebody chose. Field data is a photograph taken at arm length, in rain, with a gloved thumb across one corner.

A friendly operator. The person driving the demo built it. They know which phrasing works and which button to avoid. A technician on their fourth job of the day knows none of that and should not have to.

Time pressure. Nothing in a demo has a deadline. In production the crew is leaving site in ten minutes and the answer is needed before the van moves.

Consequence. Nobody signs the demo. The moment output reaches a document with a name on it, every wrong answer has an owner, and the tolerance for confident nonsense goes to zero.


03

Five things that break between the demo and production

None of these are model problems. All five are system problems that only show up once volume and consequence arrive together.

  1. Field conditions do not resemble the training set. Low light, partial frames, a lens covered in dust, an asset photographed from the only angle the scaffold allowed. Accuracy measured on clean inputs tells you almost nothing about accuracy on the inputs your crew will actually produce.
  2. Review capacity becomes the constraint. An agent raises output immediately, and output has to be checked by somebody who understands the domain. If the reviewer does not understand offline sync, or what a certificate has to prove, they will approve something that reads correctly and costs a day of field work to unwind. Plausible and wrong is the expensive failure mode, because it passes every check a busy person applies.
  3. Nobody owns a wrong answer. In a regulated document the question is never what generated the finding. The question is who decided to let it through, and whether that decision can be reconstructed a year later.
  4. Conditions drift while confidence stays flat. New device, new site, a different camera, a season with different light. The model keeps returning the same confident scores while the accuracy behind them has been sliding for weeks, because nothing in the system was built to notice. Loud failures stop the line and get fixed. Quiet ones keep producing output that looks exactly like the output from when it worked.
  5. The process never changed, so the cost never moved. Individual people get faster. The approvals, the handoffs and the sign-off stay exactly as they were. Usage goes up, the operating cost line does not move, and nobody can explain why. The useful question is not whether people use AI. It is which step disappeared, who stopped waiting on whom, and what the team no longer does at all.

04

Five features we shipped, three we removed

The removals are more useful than the launches, because they mark where the line sat in real use.

Feature Outcome What decided it
Photo defect tagging Shipped Suggests type and severity while the inspector is still standing at the asset
Voice to structured notes Shipped Hands stay on the tool and the fields fill themselves
Drafted report summary Shipped The engineer reviews and signs instead of writing from scratch
Out-of-pattern reading alerts Shipped Flags the odd result before the crew leaves site
Plain language history search Shipped Finds last year at this site without anyone building a query
Automatic pass or fail decisions Removed Nobody will sign a certificate that a model decided
In-app chat assistant Removed A button beats a prompt in gloves and rain
Predictive maintenance scores Removed Too little history behind it, so the output was confidently wrong

Read the removed rows together and a rule falls out. Everything that survived assists a decision. Everything that was removed tried to make one.


05

The layer that decides whether any of it pays back

Agents in production need a review step designed as carefully as the model, and most projects design neither.

Four checks belong in front of anything that reaches a customer document: is the evidence present, is it attached to the right asset, is it recent enough to still describe reality, and did a named person accept the output. None of those are model work. All of them decide whether the output survives a challenge.

The reviewer matters as much as the checks. Review handed to whoever has capacity turns into a rubber stamp within a month, because the reviewer has no way to tell a good output from a convincing one. It belongs with someone who has sat through an audit and knows which detail the auditor returns to.

Three things are worth putting out of reach entirely. An agent should not issue a pass or fail on its own, should not edit a record after sign-off, and should not be the only account that touched a finding. We make the same argument about AI visual inspection: the model earns its place when a person can explain and defend what it produced.

A confidence score is a property of the model on data it has already seen. It is not a statement about the asset in front of the inspector, and it cannot carry a signature.


06

Where a generic platform stops

Model access is a commodity now. Orchestration frameworks are mature. The part that stays specific is everything the agent has to touch: your asset hierarchy, your approval chain, your evidence rules, the format your auditor expects, and the system that actually schedules the work.

That layer is where projects stall, and it is the reason so many agent pilots run fine for one team and never reach a second. The pilot worked because one engineer cleaned up the edges by hand every evening.

Dev Station builds that layer as AI agent development services and as custom inspection software, with the review step, the audit trail and the integration into an existing ERP or CMMS treated as part of the build rather than a later phase. If an agent pilot is working for one team and stalling everywhere else, send us the workflow around it and we will tell you which part is the blocker.


07

Questions buyers ask at this stage

Does a better model fix any of this? It moves the first failure and leaves the other four untouched. Review capacity, ownership, drift detection and process design are not properties of a model.

How do we test an agent before production? Build a set of the inputs your crew actually produces, including the bad ones, and measure against that rather than against clean samples. Then run it beside the current process for a period and compare what each one caught.

What should a pilot prove? Not that the output looks good. That a second team with no involvement in building it can run the same workflow and get the same result.

Where should a first agent go? Somewhere with high volume, low consequence, and a fast feedback loop. Drafting, summarising, searching and flagging all qualify. Deciding does not.

How do we know when to remove a feature? Watch whether people route around it. A feature that gets skipped, undone or corrected on most jobs has already been rejected by the people using it, and keeping it in the product only costs maintenance and trust.

Ask an AI about this

Want an AI assistant to summarize or cite this guide?

Click any link below to open the AI with a pre-filled prompt referencing this article:

Ready to Build Your Field App?

Contact Dev Station Technology to discuss your project requirements and receive a development roadmap within 48 hours.

Get a Quote →

Related articles

Let's Talk