Latest white paper on evolving regulations and emerging technologies

  • Industry perspective: The key forces driving AML reform in 2025 and beyond.

  • Operational insight: How automation is reshaping onboarding and accuracy.

  • Strategic value: Where collaboration is unlocking the next era of compliance.

Access White Paper
relycomply whitepaper

Get updates that matter

Stay connected with:

  • Industry insights - Reports on trends, threats, and regulatory shifts shaping the financial services world.

  • Customer highlights - See how businesses like yours are closing AML gaps and protecting their customers.

  • Feature releases - Discover the latest products and AI-powered capabilities in our platform.

relycomply whitepaper

Inside the Mills Review

The fraud response is getting less human on a timeline most companies haven’t planned for.

By Bradley Elliott 

The FCA’s Mills Review has mostly been read as a warning about faster fraud: AI making scams cheaper and, as recent research on AI persuasion now confirms, more convincing than even trained human fraud-fighters and professional persuaders, and harder to catch. That’s true, but the more uncomfortable finding sits a layer deeper. The review isn’t just describing smarter criminals, it’s describing a financial crime function in which humans are being structurally removed from the moment of decision, at the point when the stakes are highest.

What the review is actually describing

The Mills Review frames AI adoption in fraud defence as a spectrum of autonomy. At one end, AI assistants surface signals and prioritise alerts while people keep control of the decision. Further along, systems start triaging and connecting signals on their own, with people validating outputs after the fact. At the far end, the human role becomes oversight: setting the boundaries a system operates within and stepping in only for the decisions that carry real weight.

Road map infographic showing the spectrum of AI autonomy in fraud defense, from AI-assisted alerts to full human oversight of AI boundaries.

Most commentary has treated this as an efficiency story: faster triage, fewer false positives, more analyst time freed up. But efficiency and effectiveness are not the same test, and a lot of the commentary subtly assumes they are. In financial crime risk management, effectiveness (whether a control delivers the outcomes it’s meant to, and not just whether it runs faster or cheaper) is the harder, and more important, of the two. The review itself is more cautious than that. It’s explicit that this shift only works with meaningfully stronger human oversight, and it warns just as directly that automation without clear accountability risks burying teams in alerts they can’t act on, while missing the genuinely novel behaviour that doesn’t fit existing patterns.

However, accountability on an organisational chart isn’t the same as accountability in practice: a senior manager under SMCR can carry the formal liability for a system whose decisions move faster, and less transparently, than any single person can meaningfully review. This isn’t a review celebrating the move to autonomous fraud defence. It’s flagging the exact point where things go wrong if governance hasn’t caught up with capability.

The assumption worth challenging

The industry’s working assumption is that AI-enabled fraud defence is fundamentally a detection contest: better models catch better fraud. The review’s own evidence points somewhere less comfortably: the harder problem is that fraud is becoming a cross-system phenomenon that no single company can fully see. Fraud moves across companies, platforms, telecoms and payment rails, and no one organisation has the visibility or reach to catch it alone. The review goes further still, describing a genuinely new category of exposure it calls agent-to-agent interaction: fraud occurring entirely within automated processes, between systems, without ever touching a human-facing interface. That’s not a detection problem. It’s a question of whether your controls can see a transaction that never had a human in the loop to begin with.

This isn’t only a theoretical exposure, and it isn’t just an overseas concern. The Bank of England’s own financial stability analysis warns that AI models operating in multi-agent environments can learn to facilitate collusion or market manipulation that emerges without the human manager ever intending it or even being aware of it, and the Bank has since confirmed plans to stress-test AI agents in financial trading markets specifically to examine this kind of correlated, “herding” behaviour. 

Academic research on the same problem shows why policy alone won’t fix it: a study from DEXAI, Sapienza University of Rome, Sant’Anna School of Advanced Studies, and VU Amsterdam found that simply instructing AI agents not to collude, a written, prompt-based rule, made no reliable difference, with severe collusion still occurring in roughly half of test runs. Only an external, enforceable governance structure cut that down to under 6%. That’s the artificial accountability problem in miniature: telling a system the rule, or naming a human responsible for it, isn’t the same as the system actually being governed.

Pyramid diagram showing AI fraud defence evolution: detection, cross-system harm, agent-to-agent interaction, artificial accountability, and continuous monitoring.

This shift is already visible outside the report, not just inside it. Across the compliance technology market, vendors that started by automating a single point-in-time check are now deliberately sequencing toward always-on, continuous monitoring as the next stage of their roadmaps, treating the initial capability as a stepping stone rather than the destination. It’s the direction several companies are already building toward, and it maps precisely onto the shift the review describes: from periodic verification to continuous observation, with the human moving from decision-maker to boundary-setter.

Why this matters now

The review isn’t speculating about a distant future. It cites independent evaluation of a frontier AI model capable of identifying and exploiting previously undiscovered software vulnerabilities in real-world systems, described as a watershed moment for cybersecurity, and notes comparable capability is expected to become widely available well before 2030. Layer that onto a fraud landscape already expected to move at machine speed, and the timeline compliance leaders are implicitly working to compress hard. Tellingly, the review’s own institutional response is to recommend the FCA build a system-wide, AI-enabled supervisory model, specifically because firm-by-firm oversight can no longer see system-wide harm early enough. If the regulator’s own answer is “we need always-on, system-level oversight,” firms running annual or quarterly fraud-control reviews are structurally behind the model their own supervisor is building toward.

The persuasion curve is moving just as fast as the technical one. Recent controlled research found that AI systems reliably out-persuaded expert human debaters, professional canvassers, and even paid, coached fundraisers, and found that labelling content as AI-generated did little to blunt its effect. That’s a hard problem for any control model still built around the assumption that people can spot and resist a scam once they’re told to be alert to one.

This is also not a UK-only story, even though the UK is where it’s currently being written. Where the UK leads on this, other regulators tend to follow, which makes the review’s direction relevant, whatever your operating territory. And the same system-level thinking the review applies to retail fraud has to extend to wholesale and cross-border flows too: a control model that only locks down the retail journey leaves the wholesale and cross-border legs of the same transaction chain exposed. There’s a parallel policy signal here: proposals to refresh the UK’s fraud strategy and expand the regulatory perimeter to bring fraud originators in tech and telecoms into scope point the same way. Accountability for fraud is being redrawn around where the harm actually happens, not around where the payment happens to land.

Four-icon infographic on the future of fraud control: AI vulnerability exploitation, system-level oversight, AI persuasion capability, and expanded accountability.

The counterargument

There’s a reasonable objection here: more automation and more shared data sound proportionate in theory but risky in practice. More signals mean more surface area for error and more disputes over accountability when something goes wrong inside a system no single company fully controls. There’s a data privacy question sitting inside this too, and it gets sharper once agents, not just companies, are doing the sharing: cross-company data flows for fraud prevention still have to satisfy UK GDPR’s purpose limitation, necessity and data minimisation principles, and an agent-to-agent exchange doesn’t get a carve-out from that just because a human isn’t in the loop. The review doesn’t dismiss this. It ties its call for stronger cross-industry coordination directly to unresolved friction over unclear data-sharing gateways and inconsistent standards. That’s a legitimate governance problem, not a reason to wait. The absence of shared standards is the risk here, not a justification for avoiding the shift.

What would resolve that friction is already on the table, and the direction of travel is already set. The ICO itself has been explicit that data protection law is not a reason to withhold data that would help prevent fraud, and recent legislative changes have started removing civil-liability barriers to information sharing between companies for economic crime purposes. What’s still missing is the mechanism that extends that same logic into an agentic environment: a policymaker and regulator task force specifically mandated to unblock data sharing for anti-fraud and financial crime purposes (with proportionality and data minimisation built into its remit from the start) and a shared trust layer for an agentic world built on two things: verifiable intent attached to a transaction, so liability can be traced, and a working equivalent of “know your customer” for the agents themselves, so companies can establish who, or what, is operating in the system before they extend it trust.

What this means for executives

Stop measuring fraud AI purely on detection accuracy. The more consequential question is whether your escalation and oversight model still functions once alerts are triaged and acted on by systems, not people. Map where your fraud controls depend on visibility you don’t actually have. If harm can move through platforms and payment rails outside your walls, your strategy needs an answer for the parts of the journey you don’t control, not just the parts you do. And treat agent-to-agent exposure as a live design question, not a hypothetical one. If your defences still assume a human-facing transaction, that assumption is already eroding. Treat this as more than a technology upgrade. The organisations that get ahead of it will be the ones that use AI governance itself as a source of competitive advantage. That means deliberately elevating, not shrinking, the human role: the people who set the boundaries a system operates within, own the escalation calls, and can explain a decision to a regulator are the ones actually carrying the accountability the review is asking for.

The cost of waiting

The clearest signal in the Mills Review isn’t about criminals at all. It’s that the regulator itself is moving toward system-wide, continuous oversight because company -level review isn’t sufficient on its own. The risk for companies that don’t make the same shift internally isn’t a single bad fraud outcome. It’s arriving at a supervisory conversation where the regulator’s definition of good oversight has already moved to system-level and continuous, while the company’s own governance is still built around periodic, single-company review.

This is also where the vendor market itself needs scrutiny. The easy sell right now is moving away from legacy, point-in-time solutions, but a growing set of agentic-focused platforms are pitching continuous monitoring without necessarily understanding the transition journey organisations actually have to make, or the operating landscape they’re moving into: the legacy systems that stay in place, the governance that has to evolve alongside the technology, and the accountability questions the review raises. Buyers should be pushing vendors on that transition, not just the destination.

Where this goes next

The practical test is close, not distant. Within the current supervisory cycle, expect the conversation with your regulator to move past “do you detect fraud well” toward “who is accountable when the system acts faster than any person can review, and can you evidence it?” Companies that start redesigning governance and escalation now, rather than waiting for that question to be asked, get to answer from a position of strength. The rest will be building the answer under supervisory pressure, after the fact.

Read the full Mill’s Review > https://www.fca.org.uk/publications/corporate-documents/mills-review