Skip to main content
Allan Buntoengsuk

Deciding where an AI agent can act on its own, and where it has to stop and ask.

AML compliance platform

Year
2026
Role
Lead Product Designer
Tags
AI SurfacesRegulated SystemsFintechAudit & Compliance

Overview

An AI agent reads the transaction history behind a money-laundering alert and returns a scored report ending in cleared or escalated. A compliance investigator then signs it. But a regulator may open that file two years later, and the name on it is the investigator's, not the model's.

A seed-stage compliance platform brought me in to run discovery across the product, then design the first build phase. The brief asked for a faster correction loop but I realized that speed wasn’t the constraint.

Strategic Frame

The feedback loop already exists, but investigators had abandoned it for a Google sheet.

Notes and flag corrections fed the model's evals, but nothing came back to the person who wrote them. The work moved into a shared Google Sheet that had the case notes, audit trail, and client comms at once. None of this was readable by the model.

Two things discovery moved that weren't in the brief:

  • Investigators read raw transactions first, flags last as a sanity check. The brief assumed narrative first.
  • A verified fact is not a risk finding. A registry hit is a fact; it might produce a flag. The product rendered both identically. I made the distinction a first-class primitive and it carried through both phases.
The model's evals
Two independent sources named the same missing primitive

Decision 1

Asking and changing are different actions, so I didn’t put them both in the chat box

A regular ol’ query would just output an answer. But a correction would open a chain: parse input, propose a plan, rerun on approval, write to a log.

Investigators can reach the co-pilot three ways. A persistent panel for work that starts with a question. A hover trigger on any block, pre-loading that block as context. Inline selection, so a correction attaches to the disputed sentence rather than the entire block, which can be the difference between correcting a fact and challenging an entire conclusion.

The founder's reference was Cursor: AI would plan, approve, execute. I kept that but dropped the rest, because Cursor edits files you can revert privately. The plan would need to scope out the entire correction so a user would know what is affected.

Decision 2

Correcting one fact can change five conclusions. That’s the moment when the AI should stop and ask.

A correction to an occupation field moves through whatever the model concludes downstream, but the investigator may not be able to see how far it reaches from the sentence they're editing. The copilot proposes a reprocessing plan that names every block the correction touches. And nothing would run until the investigator approves the list.

The autonomy line isn't just drawn at capability. Parsing unstructured input, asking follow-ups, calculation of blast radius would all be automated. And no writing would be done without explicit approval.

Decision 3

The audit trail is a byproduct of correcting

Every correction writes a change card with an undo action, preserves the prior version, and lands in the case history.

Undo didn’t feel right at first because audit trails you can undo felt weak. But the undo is itself a logged event, and making corrections should be easy to do which allows investigators to trust the AI.

Case history, designed

What I decided not to do

Rerunning the full investigation was the obvious choice, but it would have affected credibility.

Engineering surfaced a constraint I would have missed: the model's scores drift between runs on identical inputs. So reruns scope to the blocks named in the plan.

The layout breaks below 768px and investigators run this in split-screen. I didn't solve this—it went into the handoff as a named constraint.

Only one investigator was available to interview, who supervised agents as opposed to being in the client queue. I couldn't tell whether her reading order was the natural mental model or an adaptation to the workaround process.

Split-view layout at the breakpoint that breaks. Designed; the underlying constraint was handed off unsolved

Impact

The panel, the hover trigger and correction chain are live.

A correction will parse, ask for approval, write a note with an undo action, and queue a rerun scoped to affected blocks. The plan card doesn't render yet, so approval happens in the chat back and forth. Inline selection was designed but hasn't been built.

The handoff was annotated Figma covering both intents, the approval flow, layout at each breakpoint, the case history surface, and the co-pilot components.

The measure I set was whether corrections actually move out of the Google Sheet—that data doesn't exist yet. What's true now is the record has somewhere to live inside the case.