6 Oct 2026 · 5 min read
Agentic AI in Production: 7 Lessons From Real Client Projects
Most agentic AI projects stall before production. Seven lessons from building AI products for clients on what makes agents reliable, affordable and trusted.
On almost every agentic AI product I've shipped to production, the code that calls the model ended up being a small fraction of the codebase. The rest was data lookups, queues, retries, permissions, billing and the screens people use to check the agent's work. I've built these systems for clients in healthcare, media, marketing and legal services, and that ratio surprised me at first. Now it's the first thing I plan around.
It also explains a lot about why so many agent projects stall. Gartner has warned that more than 40 percent of agentic AI projects could be canceled by the end of 2027. In my experience the model is rarely the reason. These are the seven lessons that changed how I scope and build agents, including a few that go against the usual advice.
1. Price the plumbing, not the prompt. When a founder asks me to estimate "an agent," the conversation used to center on the model and the prompt. Now I ask them to walk me through everything the agent has to read, write or trigger. On a healthcare assistant, that list ran to medical data sources, document scanning, triage rules, location search, payments, encrypted storage and multi factor login. The reasoning step was maybe a tenth of the effort. Estimates that only cover the clever part are the main reason I've seen agent budgets run out halfway.
2. Let the model explain, not decide. The most useful design change I've made is to take the facts out of the model's hands. In the healthcare product, whether two medications interact comes from a trusted database, not from the model's memory. The model's job is to explain that answer in plain language. On a legal billing platform, the model proposes billable entries, and each one carries the exact message it came from. A lawyer can check it in seconds, and that's why they were willing to use it at all.
3. Most steps don't need a model. This is where I disagree with a lot of agent demos. One product I built turns a short brief into finished video ads, and the first version asked a model to handle almost every step, including things like scene durations and output formats. Those steps were slower, cost money on every run and failed in strange ways. Moving them into ordinary code changed nothing for the user and made the whole pipeline cheaper and easier to debug. Today I keep the model for decisions that need judgment, like the script and the scene plan, and nothing else.
4. When the agent misbehaves, fix the tool before the prompt. On a video editing agent, the model would sometimes ask to cut a clip at a time that didn't exist in the video. My first instinct was to add more instructions to the prompt. It helped for a day. What fixed it was making the cutting tool check its inputs and return a clear error the agent could read and correct. Prompts are suggestions. Tool validation is a rule, and rules hold up better at two in the morning.
5. Assume the user will close the tab. Video generation takes minutes. People refresh, lose signal or walk away, and a request that only lives inside one web call disappears with them. I now treat every long agent task as a background job from day one, with its own saved state, progress updates and a way to resume. It's the same habit I picked up building real time backends for multiplayer games, where you design for disconnects first because they're guaranteed to happen.
6. Retries are where agents quietly do things twice. Everyone adds retries. Fewer people ask what a retry repeats. If a step that sends an email or creates a charge times out after it has already succeeded, a naive retry sends a second email or charges the customer again. Every step that touches the outside world in my pipelines has a unique key, so running it twice has the same effect as running it once. It's a small piece of code, and it's what keeps an agent from becoming a support ticket.
7. Write the "won't do" list before the feature list. Clients rarely ask how accurate an agent is. They ask what happens when it's wrong. For the healthcare product, we agreed early that it would never diagnose. It routes symptoms to the right level of care and says so plainly. For the legal product, nothing the model extracts reaches an invoice without a person approving it. Those limits shaped the interface, the data model and the security work, and they did more to win the client's trust than any improvement in accuracy.
None of this is glamorous, and that's the point. The agent projects I've seen succeed were the ones where the model was treated as one part of a normal software system, with the same care around data, failure and permissions. The ones that struggled gave the model all the attention and left the rest for later.
If you're about to start an agentic AI project, try this before you write a single prompt. List every source the agent reads, every action it can take and what happens when each one fails or runs twice. Then draw a line through everything on that list that doesn't need judgment and hand it to plain code. What's left is the real agent, and it's usually much smaller and much easier to trust than the one you started with.