Skip to content

Operational AI vs Generative AI: A Buyer's Distinction

A credit team ran a flawless evaluation, picked a genuinely good tool, and eighteen months later nothing had changed. The one question that would have caught it.

Craig Davis

Managing Partner and Founder, Catalyze Labs | Co-Founder, CovenantFlow

5 min read

Ask one question. Does it write back?

A credit team we worked with spent most of a year evaluating AI. They ran a good process. Three vendors, a scored matrix, a bake-off against real loan documents. The winning tool read a credit agreement and pulled the covenants out of it with genuinely impressive accuracy, and everyone in the room could see the hours it would save.

Eighteen months later the covenant review still happened in a spreadsheet, once a quarter, the way it always had. Nothing about the operation had changed. The tool worked exactly as demonstrated.

What nobody asked in the bake-off was where the extracted covenant went. The answer turned out to be: into a panel on a screen, for a person to read and then type somewhere else. That's not automation. That's a reading task with extra steps, and it was never going to move a cycle time.

This is the normal outcome, not an unlucky one

MIT's 2026 study of enterprise AI, The GenAI Divide, tracked what happens to these programs. Of the organizations evaluating enterprise-grade tools, 60% got through an evaluation, 20% reached a pilot, and 5% reached production. Ninety-five percent of pilots produced no measurable P&L impact. The authors blame brittle workflows and misalignment with day-to-day operations, which is a polite way of saying the tools didn't touch the work.

The instinct is to read that as a technology failure. It isn't. Those models mostly did what they claimed. The failure happened earlier, at the point where somebody decided what to buy, and it happened because two genuinely different products are sold under one word and the difference doesn't surface in a demo.

Why the demo can't tell you

A demo shows a question going in and a good answer coming out. Generative AI does that on day one, with no integration and no permissions model and no agreed definitions, because it's answering from the document you handed it ninety seconds ago. Operational AI does the identical thing, but only after months of work that has nothing to do with the model: reading from the system of record, agreeing what the terms mean, getting entitlements right, building somewhere for the answer to land.

In a half-hour meeting they're indistinguishable. So the format quietly favors the one that can't carry the workflow.

This is why the write-back question does so much work. It's the one thing a demo can't fake, because answering it honestly requires the vendor to describe an integration that either exists or doesn't. Everything else in an evaluation is downstream of it. If nothing changes state in a system anyone else relies on, you've bought a very good reading tool, and you should price it as one.

What to ask instead

Once write-back is settled, the useful questions are all about the unhappy path. Does it read the system of record at the moment of the decision, or a copy that synced overnight? What happens when confidence is low, and does the answer name a threshold and a person or just assure you the model is accurate? Are permissions enforced before the model sees the data or filtered out of its answer afterwards, because only the first is a security model.

Then the one that separates a pilot from a system. When it's wrong at two in the morning, who finds out, and how? The generative answer is usually some version of "the user reviews the output". Fine for a drafting tool. Useless for anything running unattended, which is the only kind worth buying.

Ask about the exception rate too. Enterprise processes are mostly exceptions, and a system that handles the happy path has automated the cheap part.

Generative AI is excellent, at other things

None of this is an argument against generative AI. For open-ended language work where a fluent answer is the deliverable, it's remarkable, and the pilot costs almost nothing. Drafting, summarizing, research synthesis, code assistance. If that's the actual deliverable, buy the generative tool and don't let anyone sell you an architecture you don't need.

MIT's own data supports this, incidentally. The widely-adopted tools did lift individual productivity. What they didn't do was move an enterprise number, because individual productivity and operational throughput were never the same quantity.

The expensive part isn't the failed pilot

A pilot that goes nowhere costs less than people assume. The damage is what it does to the next attempt. A visible AI program that produced nothing makes the following proposal harder to fund, and the following proposal is frequently the one that would have worked. Budgets shrink, scope gets trimmed to something safe, and the safe version doesn't touch the constraint either.

Worse is what it does to the people. Operators who were asked to adopt something that didn't hold up are measurably harder to bring along a second time, and adoption decides whether an operational system produces anything at all. You get to spend that credibility roughly once.

Back to the credit team

They restarted, and the first fortnight went on something that had nothing to do with AI. They wrote down how a covenant test was actually performed. Within two weeks they found that two desks in the same bank were testing the same ratio differently, and had been for years, and that neither considered itself wrong.

No model would have caught that. A model would have automated both versions, faster, and produced a portfolio view that quietly disagreed with itself. Getting the definition agreed was unglamorous and took a fortnight, and it's the reason everything built afterwards worked.

That sequence, workflow first and model last, is most of the difference between the 5% and the 95%.

Sources

  1. The GenAI Divide: State of AI in BusinessMIT NANDA (2026)

Related capabilities

Related capability

If your organization is ready to move from AI experimentation to AI execution, we'd love to talk.