← Writing
August 10, 2026·8 min read· AI· Finance

The agentic AI ROI gap: why 92% feel the pressure but 34% measure it

By Michel EscodaIndependent Architect & SAP FICO Consultant
Share
Summary

The agentic AI ROI debate is being had backwards: 92% of finance leaders feel pressure to show returns, but only 34% have a measurement framework, and in SAP shops the real blocker is process debt, not tooling. Agents do not fix broken processes; they inherit and industrialize them, so the ROI number measures the process, not the AI. The way through is targeted pilots on bounded, standardized processes with data lineage and a tested fallback, and the teams that run the three-question test now are the ones with defensible AI business cases in eighteen months.

The agentic AI ROI debate is being had backwards, and the SAP shops that realize it first will look unfairly smart in eighteen months. Right now CFOs are being asked to prove the business case for agents that were dropped onto processes nobody re-engineered first. The pressure is genuine: 92% of finance leaders report pressure to demonstrate AI ROI, and only 34% have a formal framework to measure it, per the July 2026 CFO Dive survey. Every vendor deck quotes those numbers, then sells a better dashboard. You can buy every dashboard on the market, but the number you need will not appear until you measure the process the agent sits on. That is where the debate goes wrong: it treats an AI measurement problem as if the AI existed in a vacuum, when in an SAP shop it always lands on twenty years of accumulated process decisions.

The 92/34 gap is a symptom, not a diagnosis

CFO Dive's July 21 survey is the cleanest snapshot of the anxiety: 92% of CFOs and finance staff feel pressure to show AI ROI, but only 34% have formal measurement frameworks. The usual reading is that finance teams lack the tools or the discipline. The more honest reading is that they lack something more basic: a stable, standardized process to attach the measurement to. **You cannot build an ROI framework for an agent when the underlying process changes shape every month-end, when the data lineage is unclear, and when nobody can say exactly what the agent was supposed to improve. **The framework problem is a process problem wearing a measurement costume, and it explains why so many AI business cases get abandoned in the first quarter.

Weak signals — the 92/34 gap is a symptom; the signal the dashboards miss is that the process itself changes shape every month-end. Weak signals — the 92/34 gap is a symptom; the signal the dashboards miss is that the process itself changes shape every month-end.

The convergence of surveys matters here. It would be easy to dismiss one vendor-backed survey as self-serving, but the same week produced three independent signals: the CFO Dive pressure numbers, the Avalara finding that finance leaders are deploying agents before governance exists, and the CFO.com reporting on governance, audit trail, and data lineage as top blockers. Three different instruments, same chord. When independent surveys converge on the same tension, you are looking at a real force, not a press release.

The process debt that agents inherit

Here is the part that never makes the keynote: an agent does not fix a bad process, it inherits it. In SAP shops the mechanism is visible. A Joule agent deployed on an FI closing process that runs on manually maintained reconciliations, late postings, and a patchwork of custom reports does not smooth those flows out; it automates them as they are. The agent executes the broken workflow faster than the human did, which means the errors now scale with the throughput. When the ROI number comes back disappointing, the verdict is "AI is not mature yet." The verdict is wrong. What got measured was the process debt the agent was asked to run on top of, and the agent performed exactly as the process allowed.

This is the counter-intuitive truth of the whole cycle: agents amplify what exists. If the process is clean, the agent makes it cheaper and faster. If the process is debt, the agent industrializes the debt. The research dossier on this point is blunt: garbage in, agents out, more efficiently.

Why SAP shops feel this first

SAP finance processes are the most heavily optimized, most heavily regulated, and most heavily customized processes in the enterprise, which makes them the best laboratory for the phenomenon. A closing process in S/4HANA is not one process; it is hundreds of small decisions, some automated, some manual, some governed by local policy, some by habit. The habit ones are the ones agents expose. When a controller is asked to approve an agent's reclassification, the real question is whether the rule the agent followed was ever written down, validated, and reconciled with the other rules. In most shops it was not. That is not an AI failure; it is a process audit that was overdue anyway, and the agent simply made the overdue visible.

There is a second reason SAP shops feel this first: the history. A typical S/4HANA finance landscape carries years of configuration decisions made for reasons nobody remembers, workarounds that became policy, and interfaces held together by goodwill. That is not a criticism; it is the normal cost of running a real business. It does mean the process debt is deeper, which means the gap between the AI promise and the AI result shows up sooner and more painfully. You will feel the ROI gap first because you have the most to inherit.

The speed gap: deployment is outrunning control

The Avalara survey from July 2026 found finance leaders racing to deploy AI agents before governance frameworks exist. Read that again, because it inverts the usual assumption that governance is slow and deployment is fast. Both are true, and the gap between them is where the risk compounds. CFO.com's July 22 coverage lists governance, audit trail, and data lineage as the top blockers for agentic AI, which is the same trio finance has been fighting for a decade, now with an agent attached to each decision.

The second-order effect is the one nobody budgets for: when deployment outruns control, the control is written after the incident. The audit trail is reconstructed, the data lineage is documented in the runbook that gets updated after the agent misclassifies a provision. What gets produced then is not governance but archaeology, and the finance team pays twice: once for the incident, once for the documentation. The teams that avoid this do not deploy slower; they deploy smaller, on processes where the audit trail already exists and where a wrong autonomous decision has a bounded cost.

The partnership that redraws the map

The OpenAI and PwC collaboration announced in July 2026 is the quiet signal in this whole story. Consulting and LLM vendors are moving directly into CFO-office workflows, and they are doing it without going through the ERP vendor. For SAP shops this matters twice. First, the competitive frame shifts: your finance AI decision is no longer only about Joule versus a point solution; it is about whether the agent layer sits on your process context or on a generic model that does not know your chart of accounts. Second, the measurement question follows the money. When the LLM vendor and the consultant own the agent, the CFO is even further from the data lineage, and the ROI case gets even harder to build.

The ERP stack is not going anywhere; what changes is that the agent layer is becoming an independent architectural decision, and the process underneath decides who wins. A generic model on a clean process will beat an embedded model on a broken process every time, and every vendor in this market knows it.

Three questions before any agentic pilot

None of this means waiting for perfect governance. That objection is legitimate: if you wait for a flawless process, you will never start, and the window of competitive advantage will close while you standardize. The way through is targeted pilots on bounded, well-defined processes, and three questions separate useful pilots from expensive experiments:

  1. Is this process stable and standardized before AI touches it?
  2. Do we have data lineage and an audit trail for the process?
  3. What is our fallback when the agent makes a wrong autonomous decision?

The first question filters for process debt: if the answer is no, the pilot is measuring debt, not AI. The second filters for control: an agent you cannot audit is an agent you cannot defend in an audit committee. The third filters for courage: every finance team that has run an agentic pilot has needed the fallback, and the teams that planned it spent an afternoon on it while the teams that did not spent a quarter explaining it.

Decision under uncertainty — the path ahead dissolves into open space; the three questions are how you decide before you deploy. Decision under uncertainty — the path ahead dissolves into open space; the three questions are how you decide before you deploy.

AP approval is the classic bounded candidate: high volume, low judgment, clear approval chain, existing audit expectations. Simple journal entry classification is another. These are not grand agentic visions; they are processes with enough structure that the agent's decisions can be checked, reversed, and measured. Start there, and the ROI framework builds itself.

What the next twelve months look like

The open questions in the research dossier point at where this breaks next. How will agent behavior be auditable under SOX and IFRS frameworks? Nobody has answered that convincingly yet, and the first vendor who ships a real audit trail for agent decisions will own the conversation. Will CFOs accept AI decision-making for reclassifications and provisions, or only for informational and routing tasks? The evidence so far says the second, and the boundary is drawn by the same three questions above.

The second-order moves are visible from here. First, the measurement conversation will flip from "what is the ROI of AI" to "what is the debt in this process," because that is the only question that produces a number you can defend. Second, control frameworks will move from after-the-fact audits to pre-deployment process certification: the question shifts from what the agent did to whether the process was ready for it. The CFOs who run the three-question test now are the ones who will have both answers ready.

Picture the finance director who runs this play: one bounded process, a standardized flow, a documented lineage, a tested fallback. Eighteen months from now, when the board asks for the AI business case, she does not show a slide with a projected curve. She shows the AP approval cycle that went from days to hours, the audit trail that survived a real audit, and the process that got better whether or not the agent gets credit. That is the ROI gap closed, not with a better dashboard, but with the unglamorous work of making the process worth automating. Hard, unglamorous, and entirely doable, and it starts with one process, one question at a time.

Sources

  • Finance Leaders and AI Governance Survey (Avalara, July 2026)
  • 92% of Finance Leaders Feel Pressure to Show AI ROI (CFO Dive, July 21 2026)
  • CFOs Struggle to Get a Handle on Agentic AI (CFO.com, July 22 2026)
  • ERP Today: ongoing coverage of agentic AI in finance (July 2026)
  • OpenAI and PwC: Strategic Collaboration for CFO Office AI (July 2026)
  • Joule for Finance: Generative AI Copilot in S/4HANA Cloud (SAP product documentation)
  • The State of AI in Finance 2026 (McKinsey)
  • AI Governance in ERP: Hype Cycle 2026 (Gartner)

Frequently asked

Why do so few finance leaders actually measure agentic AI ROI?

Because the framework problem is a process problem: you cannot build an ROI framework for an agent when the underlying process changes shape every month-end, the data lineage is unclear, and nobody can say exactly what the agent was supposed to improve. 92% feel pressure to show ROI, but only 34% have a formal framework.

What should finance teams pilot first with agentic AI?

Bounded, standardized processes with an existing audit trail: AP approval is the classic candidate, and simple journal entry classification is another. High volume, low judgment, a clear approval chain, and decisions that can be checked, reversed and measured.

Need this in your organisation?

I work with a small number of clients each quarter on ERP strategy and IT-department automation. If the questions raised above are live in your team, get in touch.

Start a conversation
Share