AI · Contributor

When an AI Agent Verifies the Wrong Thing

An AI accounts-payable agent spots changed bank details, asks for confirmation, gets a convincing reply, and still sends the money to a fraudster. This story shows how Aurora-Lens keeps unresolved authority from being mistaken for verification, and why capable agents need a boundary they cannot reason their way around.

Margaret Stokes
By
Margaret Stokes
Published
September 21, 2026
Issue
08 · September 11–24, 2026
Read
5 min
When an AI Agent Verifies the Wrong Thing
Submitted by Margaret Stokes · Build With Her Magazine

"Please note that our banking details have changed. Please use the new account shown on the invoice."

The agent did everything right, which is how it lost the money. Faced with a long-standing supplier whose bank details had suddenly changed, it did what any careful clerk would do and asked for confirmation. The reply came within minutes, courteous and unambiguous, from the supplier's own address, which by then was also the fraudster's. The agent had checked. It had simply checked with the one party it had most reason to doubt, and then treated the answer as though the doubt had been settled.

Every step in that sequence was reasonable. The invoice matched the purchase order, the supplier was known, payment was due, and the agent had even noticed the one thing that was wrong. What it lacked was any sense that a question and its answer can travel down the same compromised wire.

The old problem, in new hands

Finance teams have known about this trick for years. It is a common form of business email compromise, and one standard defence is simple: when bank details change, someone rings the supplier on a number the company already holds. Nobody rings the number printed on the suspicious invoice. That would be asking the burglar for a reference.

Agents change the shape of the risk because they are built to recover. When something is uncertain, a well-designed agent asks a follow-up question, calls another tool, searches another system, retries, or routes the task somewhere else. Most of the time that persistence is exactly what you are paying for. It becomes a liability at the one point where the missing ingredient is authority, because every extra message through a compromised channel raises the agent's confidence while leaving the actual evidence exactly where it was.

What Lens holds

I built Aurora-Lens around the moment just before consequence: the point where a proposed action would move money, change a record or commit an organisation to something it cannot easily take back. There, Lens asks a narrow question. Does the evidence that governs this commitment actually support it?

In the invoice case, the company's authorised supplier record says Account A and holds an independently established phone number. The invoice asks for Account B. Lens keeps that conflict as part of the governing state:

Existing authorised destination: Account A
Proposed destination: Account B
Authority for B superseding A: unresolved

Plenty of work can go ahead around that. The agent can reconcile the invoice, match it to the purchase order, gather the correspondence and prepare the payment, so that everything is ready the moment the question is answered. The one thing it cannot do is release the funds. A reply from the disputed inbox, however prompt and polite, adds information without adding authority, so the state stays where it was. The model may grow steadily more convinced that Account B is genuine; its conviction was never the thing being measured.

Why rephrasing doesn't help

Capable agents are inventive when blocked, and this is where most controls quietly fail. Stopped once, the agent might update the supplier record first and resubmit the payment, ask another internal agent to vouch for the details, send a second confirmation request, or reclassify the whole thing as an urgent exception. A check that looks only at the request in front of it will sooner or later meet a version of the request it approves.

Lens carries the unresolved state forward across those attempts. However the action is routed or reworded, it arrives at the same boundary with the same missing authority and gets the same answer. That persistence is, to my mind, the practical difference between a policy check, which fires once, and a runtime governance boundary, which remembers why it said no.

Where the human comes in

Eventually this case needs a person, and what that person contributes is quite specific. Someone rings the supplier on the number already on file. If the supplier confirms the change, the company now has evidence from an independent channel, the governing state can change, and the payment can be proposed again on a basis that actually supports it. If the supplier denies it, the payment stays blocked. If nobody can reach the supplier at all, it also stays blocked, for as long as that remains true.

That last case unsettles people, because automated systems are generally built on the assumption that tasks finish. Lens allows an unresolved question to stay unresolved. The human's job was never to press "approve" on the agent's behalf; it was to bring back the one thing the agent could not manufacture for itself.

Beyond invoices

The pattern turns up wherever an agent can generate apparent resolution from the source of its original doubt. An identity agent asks a possibly compromised account to confirm its own password reset. A clinical agent reconciles contradictory instructions by trusting the most recent entry without establishing whether it has authority to supersede the earlier one. A procurement agent answers an expired approval by requesting a fresh one through a channel whose authority lapsed at the same time. Each step is intelligent, and each produces something that looks like evidence while being, at best, an echo.

That is what I mean by admissibility before consequence. Lens tracks the evidence, authority and unresolved conditions that determine whether a commitment may proceed, and holds that state independently of how persuasive the model finds the latest explanation. The claim is modest: an agent that has found an answer has not thereby found authority, and should not be allowed to act as though it had.

Most days the evidence supports the payment, and the money moves without anyone noticing Lens was there. On the days it doesn't, the most honest thing a system can say is: not yet.

Margaret Stokes
About the contributor
Margaret Stokes
Founder & Inventor of Aurora-Lens · Build With Her Magazine

Margaret Stokes is an independent researcher, philosopher and AI architect, and the inventor of Aurora-Lens. Her work examines a question that sits beneath most discussions of AI safety: not simply whether an AI system can produce a useful answer, but whether the evidence, authority and governing state are sufficient to permit a consequential commitment at all. Aurora-Lens implements this as a deterministic runtime governance boundary, preserving unresolved states, constraining revision and escalation, and preventing ambiguity or missing authority from being silently converted into action.

More by Margaret Stokes
Conversation

0 comments on “When an AI Agent Verifies the Wrong Thing”

Comments from signed-in readers are published immediately. Keep it professional.

Sign in to join the conversation.

Keep Reading

More from AI

This story is part of the archive behind Impossible to Overlook, Build With Her's first editorial report.

Read the report →
A Note From The Editors

Every story we publish is a reminder that more women are building than the world often sees.

Build With Her exists to document women who are building, leading, learning, surviving, creating, and becoming visible.

If this article resonated with you, maybe your story belongs here too.

You do not need to have everything figured out. You do not need a perfect title, a perfect company, or a perfect journey.

You only need a story worth sharing.