Where AI is actually paying off in enterprise software — and where it isn't.
TL;DR — Two years into the LLM hype cycle, the patterns are clear. AI pays off when it's a faster, cheaper version of human reading or human typing — bounded by deterministic rules and reviewed at the points that matter. AI loses money when it's asked to make business decisions, set prices, or be the only voice in a one-shot pipeline. Six wins, four misfires, and the architecture pattern that separates them.
The pattern that separates wins from misfires
I've shipped AI features into production in five engagements over the last two years. The ones that paid off shared a structure: an LLM doing what humans used to do between systems — reading, summarising, drafting, classifying — with deterministic rules controlling what actually got committed.
The ones that failed all looked the same too: an LLM placed in the decision path of something that costs the company money — pricing, approvals, payments, content authorship — without a deterministic fallback or a human review gate. They produced demo magic that crumbled in production within three weeks.
Here's the rough rule I use now: if the LLM is wrong, what does it cost? If the answer is "a human spends 30 seconds correcting it," you have a candidate. If the answer is "we issue a refund," you don't.
Six places it's paying off
1. Lead qualification before human handoff
One of the cleanest wins. An LLM in front of WhatsApp or a website chat triages incoming leads — captures the buyer's intent, asks the four questions a junior salesperson would ask, hands a structured payload to the CRM, and only escalates warm leads to a human. The cost of being wrong: a salesperson sees one extra unqualified lead, rejects it in 5 seconds. The cost of being right: the team's after-hours coverage went from "none" to "always."
This is the engagement model behind the AI Sales Agent platform. The pricing decisions are not made by the LLM — they're a deterministic engine downstream. The LLM owns the conversation, not the commercial outcome.
2. Documenting legacy code
Reading old code is exactly the kind of mechanical work a human reluctantly does and an LLM does cheerfully. Pointed at a VB.NET module, a well-prompted LLM produces a serviceable explanation of the business logic, the inputs, the outputs, and the side effects in roughly an hour of agent time per module. A human reviews and corrects in 15 minutes, which is faster than producing the doc from scratch.
Done at scale on Jobscope, this turned a documentation backlog estimated at six months into something the team finished in three weeks. The full recipe is here.
3. AI-assisted code migration (with structured handoff)
Different from "AI writes the code, ships to production." This pattern is "AI produces a translated module + a parity test plan + a list of known unknowns; human reviews and merges." The output is a reviewable PR, not a commit. On the .NET 8 + React migration in Jobscope this took the team from a 12-week module-per-month cadence to a module-per-week cadence — once they accepted that the AI's job was first draft, not final.
The mistake teams make here is treating the LLM output as a commit-ready artefact. Make it a PR with a checklist. The review takes time; less time than writing the code from scratch.
4. Customer-support summarisation and triage
Long support threads collapsed into a one-paragraph summary, tagged with intent and urgency. The savings are minutes-per-ticket, multiplied by the ticket volume. The cost of being wrong is low — a human reads the original thread anyway. The reason this works is that summarisation is what LLMs are best at; classification with a closed label set is a close second.
What kills the ROI is over-architecting. You don't need a fine-tuned model. You need a stable prompt, a 7B-class model on cheap inference, and a feedback loop where the support team can flag wrong tags. Three months of that beats a four-month fine-tuning project.
5. Internal search across messy unstructured corpora
Wikis, old PDFs, slide decks, recorded meeting transcripts. The traditional answer was Elasticsearch with a custom analyser nobody maintained. The current answer is embeddings into a vector store + a small LLM that turns retrieved chunks into a coherent answer with citations. The retrieval part is what does the heavy lifting; the LLM is just the language layer over it.
What kills the ROI is using the LLM without the retrieval — i.e., asking the model to "remember" what was in the wiki. Don't. Always retrieve, always cite, never let the model invent a source.
6. Anomaly description from existing detectors
The detectors you already have — fraud rules, log anomaly thresholds, alerting systems — produce structured signals that humans then have to interpret. An LLM in the loop, given the signal and the relevant context, produces a one-sentence explanation that a human can act on faster than they could from raw output. The detector still owns the decision; the LLM is just the explanation layer.
This is the pattern that pays off in security operations and SRE in particular. The detector is deterministic; the explanation is not. The cost of a wrong explanation is a human reading slightly more context than they needed to.
Four places it consistently misfires
1. As the final word on price
Letting an LLM decide what to quote a customer is a textbook misfire. The model is trained to be helpful, not commercial. It will quote into negative margin if a buyer pushes hard enough. It will offer a discount that wasn't authorised, because somewhere in its training data, "discount when the buyer hesitates" was associated with closing the deal.
The fix is structural: the LLM owns the conversation, a deterministic counter-offer engine owns the price. On the AI Sales Agent platform, the LLM hands the buyer's hesitation to a rule engine, which checks whether a pre-authorised counter is allowed and issues it. The LLM never sees the floor price; it can't leak it because it doesn't know it.
2. Triggering payments without a human gate
Same family of failure as pricing. An end-to-end agentic flow that converts a buyer's "OK send me the link" into a generated payment link, without a verification step, is one prompt-injection away from collecting money for the wrong amount or the wrong product. The cost of being wrong is reputational and financial, both of which are larger than the human-time saved.
The pattern that works: the agent prepares the payment payload (amount, item, customer); a human (or a deterministic rule, for trusted items) confirms; only then is the payment link generated. This is an extra 30 seconds of human time per payment that protects an unbounded downside.
3. Authoring the canonical version of any legal/regulated content
Privacy policies, terms of service, medical advice, financial disclosures, regulatory filings. The LLM's failure mode here is not "wrong" but "plausibly wrong" — the output reads correct, the legal team approves it, three months later you discover a clause that doesn't mean what someone thought it meant. Use the LLM as a drafting tool for human-authored content, never as the source of truth.
I've seen one company ship an LLM-authored privacy policy that contained a sentence about "third-party advertising partners" they didn't have. They had to issue a correction. The savings from skipping the legal review were rounded down to zero.
4. As the only voice in a one-shot, irreversible pipeline
Whenever an LLM's output flows directly into an action that can't be cheaply reversed, you have a misfire waiting. Examples I've seen:
- An LLM auto-resolving customer-support tickets, including refund-eligible ones, without a human glance
- An LLM drafting and sending outbound emails, including to sensitive accounts
- An LLM producing JIRA-ticket creation events that became part of the audit trail without a human review
The fix is the cheapest design pattern in AI: insert a human or a deterministic rule between the LLM and the irreversible action. That gate doesn't have to be heavy — a simple "review queue, anyone with role X can approve" pattern works for nearly all cases.
The architecture pattern that scales
Across every successful AI engagement I've shipped, the production architecture looks roughly the same:
Inputs
↓
LLM (the conversational / interpretive layer)
↓
Structured artefact (JSON, schema-validated)
↓
Deterministic rule engine (the "is this allowed" layer)
↓
Action (commit, send, charge)
Plus one of:
• Human-in-the-loop review queue at the action gate
• Async eval comparing LLM output against ground truth
• Audit log capturing prompt + output + action for every run
The LLM is in the middle of the pipeline, not at the end. It expresses intent in a structured form. The rule engine validates that intent against the rules of your business — pricing, eligibility, compliance, scope. The action only happens if the rules pass.
This pattern is dull. It is deliberately dull. It is what pays off, because dull is what survives the part of the year when the LLM provider releases a new model with subtly different defaults and your zero-shot prompts stop working.
What I'd push back on if a board was pitching it
Three pitches I've heard recently and would refuse to lead:
- "AI will replace tier-1 support." No. AI will reduce tier-1 ticket time by ~40% and let your team focus on the harder ones. Pitching it as a replacement loses the team's trust before you ship anything.
- "AI will write our marketing copy at scale." Only if you don't care that the marketing copy will get worse, in ways your search rankings will eventually punish. Use it as a drafting aid, not a replacement.
- "Agents will run our procurement / our finance ops / our DevOps." Not yet. The autonomy bar for those domains is too high; the failure modes are too expensive; the audit requirements are too strict. Run agents as drafters, with humans approving every action that costs more than a few rupees.
Where this leaves you
If you're being asked to "add AI everywhere" — push back with a different question: where in our product is human reading or human typing the bottleneck? Find those, and you'll find your AI ROI in three months. Pursue everything else, and you'll burn the budget on demo theatre that gets quietly turned off after the launch press release.
Pressure to add AI but unsure where it pays off?