How replacing a broken dispute resolution flow with a conversational AI experience resolved 80% of claims in under 5 minutes.
Disputes were one of the top case drivers for Credit Karma Money members. 210K+ members had filed a dispute, and 1 in 3 cases were being filed through member support — accounting for 70% of the total dollar amount disputed. Each support call cost $50, lasted more than 30 minutes, and required members to wait in queues before even speaking with an agent. Disputes were the second leading cause of member churn, behind only card declines.
The expensive assumption going into the project was that the problem was discoverability — that members couldn't find the in-app dispute flow. The data told a different story. 80% of first-time filers successfully completed the in-app process. The problem wasn't that members couldn't find the flow. It was that the flow itself was breaking them.
The dispute flow wasn't a minor friction point — it was a trust-breaking experience at exactly the moment members needed to feel supported. A disputed transaction is already an anxiety-producing event: 75% were fraud-related, meaning members' cards were being hotlisted while they tried to navigate an intimidating form. The cost wasn't just operational; it was relational. Members who felt unsupported in a crisis don't stay members — and the data confirmed it was the second-leading driver of churn.
When a member tapped the "Dispute" button, the form flow was replaced entirely by a chat conversation.
The AI accessed the member's account data at the start of the conversation — transaction details, account history, relevant context. Members weren't asked to repeat information the system already had.
Rather than presenting a full form, the conversation collected information sequentially and conversationally. Each question was framed in plain language, in the member's context, without legal framing.
Once the AI had collected the necessary information, it filed the claim automatically, sent a confirmation email with a clear explanation of next steps and how to track the claim, and ended the conversation. The member had a record and a path forward within minutes of starting.
Most interactions were restricted to predetermined selections rather than open-ended input, with sensitivity parameters set high to reduce the risk of the model going off-script.
Within one week of launch:
The experience was sunsetted after a review with legal counsel.
In a small number of instances, the model gave incorrect advice and asked questions that fell outside the scope of the prompt — despite high sensitivity settings and restricted interaction modes. No legal action, bad press, or viral incidents resulted. The decision to sunset was precautionary.
This outcome is worth documenting honestly, because it's as instructive as the success metrics.
The failure modes weren't random. They were edge cases the prompt hadn't anticipated — situations where the member's context fell outside the scenarios the conversation had been designed for. The model, operating without a hard boundary for those cases, improvised. Improvisation in a regulated financial product is a governance problem, and we treated it as one. The right response to that risk, at that moment, was to sunset the experience and return to a more controlled flow while the model and the governance framework matured. That decision was correct.
This project demonstrated something that's easy to say and hard to prove: that the design of a service interaction is inseparable from the trustworthiness of the product it represents. A dispute flow that feels like an interrogation erodes trust. A conversation that feels like a knowledgeable, calm assistant restores it.
It also demonstrated that AI in regulated financial products requires a governance model that moves at the same pace as the product. The gap between what the model can do and what the product is legally permitted to do is a design problem — and it needs a designer at the table when those boundaries are being drawn.