top of page

The First AI Agent Job in Payments Isn't Checkout. It's the Exception Queue

11 minutes ago
7 min read

SEO title: The First AI Agent Job in Payments Is the Exception Queue Meta description: Sibos 2026 showed where AI agents are reaching production first: payment repair and trade exceptions. Learn what merchants should automate and measure. Suggested slug:first-ai-agent-job-payments-exception-queue Primary keyword: AI agents in payments Secondary keywords: payment repair, payment exceptions, agentic payments, AI payment operations, payment reconciliation automation, human-in-the-loop payments

The most useful AI agent in payments may not be the one that completes a purchase. It may be the one that stops a failed payment from becoming a manual investigation.

That was the clearest operational signal from Sibos 2026 in Miami. The AI agents reaching production were working behind the scenes on payment repair, trade processing and document checking. They were not given unrestricted authority to move money. They reasoned over the problem, checked their proposed answer against deterministic rules, and escalated judgement calls to people.

For merchants and payment operators, the lesson is practical: your first AI agent should probably work the exception queue.

Short answer: why start with exceptions?

Exception handling is the best first AI use case because it combines three characteristics:

  1. The input is relatively structured.

  2. The correct repair can often be verified against a rulebook.

  3. The payment can remain paused while the system proposes and checks a fix.

That is a safer starting point than allowing an agent to initiate or approve a payment autonomously at checkout.

The exception queue is where payment operations already expect investigation, evidence, approval and escalation. AI can reduce the time spent gathering information and diagnosing the problem, while deterministic controls and human oversight remain responsible for the final outcome.

Neon payment data streams passing through a deterministic repair gate

What Sibos 2026 revealed about production AI

The strongest examples came from banks that have moved beyond demonstrations.

BNY described a digital employee on its Eliza platform that repairs payment instructions that fail straight-through processing because the beneficiary, reference data, address or payment purpose is incorrect. The AI reasons over the fields and the payment context. ISO 20022 syntax and scheme rules vet the proposed repair. People retain the judgement call.

According to Microsoft’s BNY customer story, one digital employee now handles more than 10% of BNY’s payment repair issues globally. That is a vendor-published figure, but it is significant because it describes production work rather than an innovation lab experiment.

Other Sibos examples followed the same pattern:

  1. BNP Paribas reduced one securities-services trade processing flow from ten steps to six. Within two weeks, an agent handled 80–85% of the work. Every model interaction passed through an observability layer with guardrails set by compliance, legal and HR.

  2. Deutsche Bank’s Ada framework helped a German lending client onboard and issue a loan in one day rather than one month.

  3. HSBC’s Smart Checking is live in Hong Kong, the UAE and the UK. It uses AI, ICC rules and the tacit knowledge of trade specialists to check documents. HSBC has not published volume or time-saved figures.

  4. Citi said supervisors are walked through each use case, including whether the human sits in, on or out of the loop.

The Fintech Times’ Sibos day-two report captured the common architecture: AI does the reasoning, deterministic controls do the vetting and people make decisions where context, risk or customer impact cannot be reduced to a rule.

Why payment repair is the right first job

Payment repair is not easy. But it has a defined boundary.

A failed instruction may contain a missing address, an invalid account format, a scheme-incompatible field or a reference that does not match the expected transaction. The agent can compare the message with the relevant data model and scheme requirements, identify the likely cause and propose a correction.

The important control is that the proposed correction does not go straight to the payment rail. It first passes through a deterministic verifier.

That distinction matters as the payments calendar becomes less forgiving. Swift’s MT and ISO 20022 coexistence period ended in November 2025. From November 2026, additional validations and charges apply to institutions still sending MT payment instructions. Swift also plans to stop accepting unstructured postal addresses in CBPR+ messages from the same month. Address quality is one of the failure causes BNY identified.

The Swift ISO 20022 guidance should be treated as a live implementation reference because standards and timelines can evolve. The direction, however, is clear: structured payment data is becoming an operational requirement.

Bottomline’s 2026 Payments Intelligence Gap report reinforces the problem. Among more than 300 banking and payments professionals:

  1. 36% reported no AI integration in payment systems.

  2. 48% called AI integration a critical or high priority for the next 12 months.

  3. 65% identified structured data quality as the biggest barrier to ISO 20022 enablement.

The opportunity is not simply to add a language model to a payment hub. It is to make payment data structured enough for an agent to reason over, and rules-based enough for the business to verify the answer.

The merchant equivalent of a bank repair queue

Merchants have the same pattern, even when the terminology is different.

A merchant exception queue may contain:

  1. Failed card authorisations that need routing or retry analysis.

  2. Refunds that do not match the original transaction.

  3. Settlement references that cannot be matched to orders or invoices.

  4. Reconciliation breaks between the payment provider, ledger and bank account.

  5. Duplicate captures, partial settlements or disputed transaction records.

  6. Subscription payments where the retry path depends on customer, product or account context.

These are not merely technical defects. They are operational decisions with financial consequences.

An AI agent can gather the transaction history, identify the failed field, compare the event with routing and accounting rules, prepare a proposed action and package the evidence for approval. It should not quietly invent a new payment route or change a customer’s financial position without a controlled release step.

This is where Quantum Payments’ unified commerce platform matters. When checkout, orchestration, reconciliation, accounting and business intelligence share one environment, the agent can access both the payment event and the business context around it. The proposed fix can be tested against the relevant rules and logged from checkout through reconciliation.

That is a more useful foundation for agentic payments than starting with a fully autonomous checkout experience.

The human gate can still fail

Keeping a person in the loop is not enough. A tired operator approving hundreds of similar recommendations can develop automation bias, assuming the agent is correct because it is usually correct.

The counterargument, raised at Sibos by EY’s Stephany Kirkpatrick, is that humans also make mistakes every day. The answer is not to pretend that human review is perfect. It is to reserve human attention for the questions that rules cannot decide.

For example:

  1. Is this beneficiary plausible for this customer?

  2. Does the payment purpose match the customer’s established pattern?

  3. Is the proposed refund consistent with the order and service history?

  4. Does the settlement break indicate a data issue or a deeper control failure?

The deterministic layer should handle syntax, scheme rules, field requirements and duplicate checks. People should handle meaning, context and accountability.

Auditability is equally important. Forrester reports that BNY has more than 130 digital employees in production, with every action and rationale logged. Bottomline’s Natasha Lapierre makes the broader point that AI outputs must be traceable, auditable and understandable to the people responsible for acting on them.

Human approval node illuminating one auditable payment repair path

Build or buy? Own the verifier

The build-versus-buy decision is less about which model is most impressive and more about who owns the verifier.

BNY built its capability on Eliza. HSBC built Smart Checking in-house. JPMorgan and Lloyds use Cleareye. Deutsche Bank and SMBC use Traydstream. BNP Paribas and ING use Conpend for trade documents.

The key question is where scheme-rule validation lives. If it already lives in your payments hub, a vendor repair module connected to that hub may inherit the right controls. If you build on a general-purpose agent platform, you own the ongoing work of keeping the rulebook current through every Swift standards release, scheme change and internal policy update.

Operator checklist: your first exception agent

  1. Choose one narrow queue, such as failed authorisations, refund mismatches or reconciliation breaks.

  2. Define what the agent may read, what it may recommend and what it must never execute.

  3. Connect the agent to structured transaction, order, ledger and settlement data.

  4. Put deterministic validation before any human approval step.

  5. Create a mandatory escalation path for low confidence, conflicting evidence or unusual beneficiaries.

  6. Log the input, the recommendation, the rules applied, the human decision and the final outcome.

  7. Measure time to resolution, first-time repair acceptance, operator overturns, queue ageing and unresolved value.

  8. Review whether human approvals are becoming repetitive enough to create review fatigue.

  9. Expand only when the first queue has a stable error rate, clear ownership and complete audit evidence.

FAQs

Should merchants start with checkout automation?

Usually not. Checkout is customer-facing, high-volume and commercially sensitive. Exception handling is contained, measurable and often governed by an existing rulebook. It provides a safer route to production value.

Does an exception agent make the final decision?

It should not by default. The strongest production pattern is AI reasoning, deterministic verification and human approval for decisions that require context or carry financial impact.

What should merchants measure first?

Start with resolution time, first-time repair acceptance, manual touches, queue ageing, escalation rate and the value of payments recovered. Add audit completeness and human overturn rate as governance measures.

How does this relate to fraud?

Repair is not a substitute for fraud controls. A plausible-looking correction can still be suspicious. The agent should pass relevant cases to fraud and risk workflows, consistent with the concerns outlined in Your Fraud Model Was Built for Humans. AI Agents Break It at Machine Speed.

What about disputes and liability?

Keep approval boundaries explicit. The Agent Liability Gap: Banks Have Principles, Merchants Have Disputes explains why merchants need a clear record of what the agent recommended, what the operator approved and what the system ultimately did.

The broader objective is not to remove every human from payments. It is to make human attention more valuable by removing the repetitive work that rules and context can safely resolve.

Sources and further reading

 
 
bottom of page