All case studies

AI Banking Platform

Client
Top-3 Korean Bank
Date
2025–2026
Role
Platform Strategy Lead → Business PMO
Team
80+ people across 20+ departments ($200M program, 300+ total)

Context

The bank was rebuilding its mobile app from the ground up, with eight gen-AI services planned for the new platform. I led the platform strategy module, then moved into Business PMO through the build.

The Chief Digital Officer had a reference point: JPMorgan Chase's hyper-personalized wealth platform. The mandate was to replicate it.

The Problem

The mandate arrived already solved, and the logic held together until you asked why.

Why would personalization drive growth? And underneath that, a question nobody had asked: did these customers want wealth management at all?

The early data made it worse rather than clearer. AI recommendations were hitting high click-through rates — personalization was working by every metric we had. Retention was slipping anyway. Users were clicking and leaving.

What I Decided

1. Stop asking why customers weren't using our app.

That question makes people defensive and returns polite non-answers. I ran the interviews the other way: why do you use the competitor's app? What do you open it for? What brings you back? People explain their own behavior far more honestly than they critique yours — and stated preference about your product is the least reliable data you can collect. What came back was that everyone had been using the same word to mean different things. The bank read personalization as putting the right product in front of the right person. Users wanted to understand their own financial position well enough to decide for themselves.

2. Reframe the mandate — and take it up with evidence, not opinion.

I brought the reframe to the team and then to the CDO: not an AI wealth manager, but a system that explains the reasoning behind every recommendation. The case wasn't that we disagreed. It was funnel data showing where users dropped, paired with interviews showing why — and the observation that our best-performing metric, click-through, was measuring the wrong thing.

3. Make explanation the design principle for all eight services, not a patch on one.

The wealth product was where the insight surfaced, but the failure mode was platform-wide. Instead of "Try fractional investment," the app said: "People in your age group often target around 3% monthly returns. Want to explore that?" The same logic ran through AI product recommendations, a daily financial to-do prompt, and a rebuilt chatbot — previously a click-through decision tree, now conversational. We extended it past investing into everyday finance: bills, expenses, small proactive decisions. Across the platform, the goal moved from selling products to helping people choose well.

4. Validate without an A/B test, and be honest that it's weaker.

Rebuilding a core system means there's no stable control to test against — you can't hold the old app constant while you replace it. We ran structured comparison sessions with the bank's internal test specialists instead: same user task, competing interaction designs, observed behavior rather than stated preference. It's weaker evidence than a live experiment and we knew it going in. It was enough to kill the options that clearly failed.

5. Make disagreement visible before trying to resolve it.

Redefining the mandate took weeks. Building it took a year across twenty-plus departments, each with its own KPIs and its own idea of what the app was for. The thing that unlocked those rooms was making disagreement explicit — mapping where interests actually overlapped and where they genuinely conflicted, and putting that map in front of everyone. Most of what people wanted turned out to be shared; the real conflict was narrower than it felt. Without that picture, meetings circled.

What Didn't Work

We didn't build a way to measure the thing we were claiming would work. The reframe was right, but the evidence was qualitative on the way in and aggregate on the way out. App-level MAU can't tell which of the eight services moved it, or whether explanation was what did the work.

The Outcome

Within six months, monthly active users rose 33% and retention improved 18%. Users described the app as a financial partner rather than a provider — the original objective, reached by a different route than the one we were handed.

What I'd Do Differently

I'd build a way to measure the thing I was claiming.

The reframe was right, but my evidence for it was qualitative on the way in and aggregate on the way out. App-level MAU can't tell me which of the eight services moved it, or whether explanation was what did the work. On a core system rebuild there's no obvious place to put an experiment — but there was room to instrument the services individually, and I didn't push for it early enough to have that data when it would have mattered.

Capabilities used