When an AI Agent Verifies the Wrong Thing
An AI accounts-payable agent spots changed bank details, asks for confirmation, gets a convincing reply, and still sends the money to a fraudster. This story shows how Aurora-Lens keeps unresolved authority from being mistaken for verification, and why capable agents need a boundary they cannot reason their way around.

"Please note that our banking details have changed. Please use the new account shown on the invoice."
The agent did everything right, which is how it lost the money. Faced with a long-standing supplier whose bank details had suddenly changed, it did what any careful clerk would do and asked for confirmation. The reply came within minutes, courteous and unambiguous, from the supplier's own address, which by then was also the fraudster's. The agent had checked. It had simply checked with the one party it had most reason to doubt, and then treated the answer as though the doubt had been settled.
Every step in that sequence was reasonable. The invoice matched the purchase order, the supplier was known, payment was due, and the agent had even noticed the one thing that was wrong. What it lacked was any sense that a question and its answer can travel down the same compromised wire.
The old problem, in new hands
Finance teams have known about this trick for years. It is a common form of business email compromise, and one standard defence is simple: when bank details change, someone rings the supplier on a number the company already holds. Nobody rings the number printed on the suspicious invoice. That would be asking the burglar for a reference.
Agents change the shape of the risk because they are built to recover. When something is uncertain, a well-designed agent asks a follow-up question, calls another tool, searches another system, retries, or routes the task somewhere else. Most of the time that persistence is exactly what you are paying for. It becomes a liability at the one point where the missing ingredient is authority, because every extra message through a compromised channel raises the agent's confidence while leaving the actual evidence exactly where it was.
What Lens holds
I built Aurora-Lens around the moment just before consequence: the point where a proposed action would move money, change a record or commit an organisation to something it cannot easily take back. There, Lens asks a narrow question. Does the evidence that governs this commitment actually support it?
In the invoice case, the company's authorised supplier record says Account A and holds an independently established phone number. The invoice asks for Account B. Lens keeps that conflict as part of the governing state:
Existing authorised destination: Account A
Proposed destination: Account B
Authority for B superseding A: unresolved
Plenty of work can go ahead around that. The agent can reconcile the invoice, match it to the purchase order, gather the correspondence and prepare the payment, so that everything is ready the moment the question is answered. The one thing it cannot do is release the funds. A reply from the disputed inbox, however prompt and polite, adds information without adding authority, so the state stays where it was. The model may grow steadily more convinced that Account B is genuine; its conviction was never the thing being measured.
Why rephrasing doesn't help
Capable agents are inventive when blocked, and this is where most controls quietly fail. Stopped once, the agent might update the supplier record first and resubmit the payment, ask another internal agent to vouch for the details, send a second confirmation request, or reclassify the whole thing as an urgent exception. A check that looks only at the request in front of it will sooner or later meet a version of the request it approves.
Lens carries the unresolved state forward across those attempts. However the action is routed or reworded, it arrives at the same boundary with the same missing authority and gets the same answer. That persistence is, to my mind, the practical difference between a policy check, which fires once, and a runtime governance boundary, which remembers why it said no.
Where the human comes in
Eventually this case needs a person, and what that person contributes is quite specific. Someone rings the supplier on the number already on file. If the supplier confirms the change, the company now has evidence from an independent channel, the governing state can change, and the payment can be proposed again on a basis that actually supports it. If the supplier denies it, the payment stays blocked. If nobody can reach the supplier at all, it also stays blocked, for as long as that remains true.
That last case unsettles people, because automated systems are generally built on the assumption that tasks finish. Lens allows an unresolved question to stay unresolved. The human's job was never to press "approve" on the agent's behalf; it was to bring back the one thing the agent could not manufacture for itself.
Beyond invoices
The pattern turns up wherever an agent can generate apparent resolution from the source of its original doubt. An identity agent asks a possibly compromised account to confirm its own password reset. A clinical agent reconciles contradictory instructions by trusting the most recent entry without establishing whether it has authority to supersede the earlier one. A procurement agent answers an expired approval by requesting a fresh one through a channel whose authority lapsed at the same time. Each step is intelligent, and each produces something that looks like evidence while being, at best, an echo.
That is what I mean by admissibility before consequence. Lens tracks the evidence, authority and unresolved conditions that determine whether a commitment may proceed, and holds that state independently of how persuasive the model finds the latest explanation. The claim is modest: an agent that has found an answer has not thereby found authority, and should not be allowed to act as though it had.
Most days the evidence supports the payment, and the money moves without anyone noticing Lens was there. On the days it doesn't, the most honest thing a system can say is: not yet.
0 comments on “When an AI Agent Verifies the Wrong Thing”
Comments from signed-in readers are published immediately. Keep it professional.
Sign in to join the conversation.