BFS · Discovery / Research / Decision v0.5 · 2026-06-11
Approach · Improving the UX of Blue Chatbot

A discovery
approach for
Blu.

Arnesh Mandal · UX · Discovery proposal
Reframe the problem

Several questions, not one.

“How do we make the chatbot work harder” is several questions, not one. Each gets a different method.

Sub-questionWhat it asksMethod
DiscoverabilityDo users see the entry point?Funnel data + usability tests
ComprehensionDo they understand what it offers before clicking?Concept tests + usability tests
Intent matchAre they at the right intent step when they see it?Segmented funnel + usability tests
Conversation valueDid the conversation itself feel valuable to the user?Transcripts (where available) + usability tests
HandoffDoes the bot hand off cleanly into the product flow?Funnel analytics + usability tests
ArchitectureOne Blu or N bots — which fits user mental models?Survey preference + concept testing
TrustWhat trust signals do users need for AI in a financial context?Trust probes in usability + survey
Segment fitDo high-intent shoppers and low-intent browsers want the same thing?Segmented funnel + segmented survey + segmented usability
A methodological commitment · Before Phase 1 begins
Important

Pre-register the questions and decision rules
before Phase 1 begins.

Without pre-registration, any result can be back-rationalised, and the research stops being able to change anyone's mind. The sub-question set, the success criteria, and the decision rules are signed off before any data is gathered.

01
The Approach

Six phases.

Phases 1–3 build the initial picture in parallel. The survey deepens it. Concepts are tested against the live experience. Synthesis ties evidence back to the questions and produces the recommendation.

The six phases · click any phase to read its detail

Six phases at a glance.

Phase 1 · Funnel Analysis
1 week · Live funnel mapping

Map where the funnel leaks.

PurposeMap current behaviour on the live Blu surface; find where the funnel leaks.
MethodsFunnel construction from Google Analytics custom events; cross-check with the CleverTap funnel; segmentation by product.
InputsGoogle Analytics events; CleverTap funnel data.
OutputsAnnotated funnel with drop-offs per product; instrumentation gap list.
Time1 week

The funnel — end to end

Homepage impression Chatbot card in viewport Chatbot card click Conversation start (first user turn) Meaningful exchange (≥ N turns) Explicit product intent expressed Handoff to product flow Application started Application submitted
Phase 2 · Conversation Review
1 week · Transcript-based qualitative coding

Read what users say to the bots.

PurposeRead what users say to the bots and where the bots fail.
MethodsStratified transcript sampling; qualitative coding; intent and failure taxonomies.
InputsAnonymised chatbot transcripts with outcome tags.
OutputsIntent taxonomy across bots; ranked failure-mode catalogue; user-phrasing glossary.
Time1 week
Conditional

This phase is conditional on transcript access. The approach proceeds without it if transcripts are not accessible — Phase 5 surfaces similar insights through usability sessions, at lower scale.

Coding categories

  • Intent — what the user is actually trying to do.
  • Failure mode — misunderstanding, repeated prompts, premature form-fill.
  • Emotional friction — frustration, hesitation, drop-out cues.
  • Cross-product confusion — wrong bot, wrong product, wrong path.
Phase 3 · Secondary Research
1.5 weeks · Competitive sweep + literature review

Learn from what others have shipped.

PurposeLearn from what others have shipped and what the literature says about trust and adoption.
MethodsCompetitive teardown; structured literature review.
InputsPublic products, app store reviews, academic search (TAM / UTAUT for banking AI, conversational AI KPI frameworks).
OutputsCompetitive matrix; literature digest.
Time1.5 weeks

Competitive sweep — minimum set

  • India banks — HDFC EVA, ICICI iPal, Kotak Keya, SBI YONO assistant
  • Global banks — Bank of America Erica, Wells Fargo Fargo
  • Fintech-native — Cleo, Klarna, Revolut

For each — capture

  • Entry-point placement
  • Single-vs-many architecture
  • Branding and naming
  • Mid-funnel triggers
  • Public complaints (app store, social)
Phase 4 · Survey
2 weeks · Panel quantitative survey

Discover the need. Validate the pattern.

PurposeDiscover user need around conversational AI for financial products; validate the patterns surfaced by the funnel, transcripts, and competitive sweep.
MethodsQuantitative survey on the BFS user panel.
InputsWorking hypotheses; insights from Phases 1–3; panel access.
OutputsUser-need patterns named and ranked; magnitude-of-preference report; adoption-construct scores by segment.
Time2 weeks
The two jobs of the survey

1 · Need discovery — where users look for help today; what would make them try a conversational AI; what they expect it to do; what they would not trust it to handle.

2 · Validation — magnitude-check the patterns the funnel, transcripts, and competitive sweep already pointed at — so the concepts that follow rest on more than a hunch.

Adoption constructs measured

Perceived usefulness Perceived ease of use Perceived risk Trust Architecture preference Entry-point preference Unprompted BFS-AI recall
Phase 5 · Concept & Usability Testing
3 weeks · Prototypes tested against the live control

Build informed. Test against the live.

PurposeBuild concepts informed by everything learned so far; test them against the current live experience.
MethodsUsability testing; concept testing on paper or clickable prototypes.
InputsConcept directions informed by Phases 1–4; prototypes (one or several variants); research-panel access.
OutputsPer-prototype findings against the live control; ranked direction preferences; trust signals.
ComparisonThe live experience is included as control. Every prototype is judged against it — the A/B comparison this team runs.
Time3 weeks

Methods

Usability testing Task framed as “find a personal loan that suits you,” not “use this chatbot.” Live experience as control.
Concept testing Paper or clickable prototypes of distinct directions; concepts are stimuli to provoke preference and trust signals.
Phase 6 · Synthesis & Recommendation
1 week · Hypothesis-by-hypothesis evidence review

Walk the hypotheses. Write the memo.

PurposePull the outputs of every prior phase together against the working hypotheses; write the recommendation.
MethodsHypothesis-by-hypothesis review of evidence; prototype-vs-live-control comparison.
InputsOutputs from Phases 1–5.
OutputsRecommendation memo.
VerdictEach hypothesis marked supported, killed, or undetermined.
Time1 week

For each hypothesis — evidence walk

  • Funnel — what it showed at the relevant step.
  • Transcripts — what conversation review revealed (where it ran).
  • Secondary — what competitors and the adoption literature contributed.
  • Survey — what users said about need and preference, and at what magnitude.
  • Usability — how the prototype performed against the live control.
Memo carries forward

Supported — the prototype direction that won, with the user evidence behind it. Killed / undetermined — what was tried and what would unblock the next decision.

Blu · A discovery approach · End of part 01

A discovery approach.
Output: research-backed solutions.

Six phases of evidence — funnel, transcripts, secondary, survey, prototypes against the live — distilled into design directions that won against the live control, paired with the user evidence behind them and the questions still worth opening next.

Versionv0.4 Date2026-05-20 AuthorArnesh Mandal
02
Part 02 · The Findings · Midway readout

What the data says.

Phases 1 and 2 are executed — the funnel mapped across Gold Loan, EMI Card, and Personal Loan, and 533,195 real conversations read and coded. Six findings follow. Each carries its numbers and the users' own words.

Phases 1 + 2 · How to read the numbers

Two phases of evidence, cross-checked.

533,195 Conversations read Every conversation from two May exports — the full dataset, not a sample. PII masked.
3 Product funnels mapped Gold Loan · EMI Card · Personal Loan across its three journey variants.
15 Analysis lenses Scripted passes over the full dataset — outcomes, gates, language, trust, loops, offers, rejection.
Cross-validation

Every funnel claim is re-measured inside the chat data — of users who reach a step, how many end there. Where chat confirms a funnel point, it makes it stronger and adds the users' own words. Every quote in this part is verbatim transcript text.

The six findings · click any finding to read its evidence

Six findings at a glance.

Finding 01 · The phone gate

Users refuse the number. OTP isn't the problem.

34.5% of all conversations (183,773) end at the mobile-number / OTP ask. The funnel sees the same wall — EMI Card loses 45% at login before anything happens. Splitting the gate apart shows what actually fails:

Number ask, no number given14.1%
OTP-framed ask, no number given12.2%
No number ever entered — reluctance26.3%
Gave a number, died at OTP entry4.0%

Reluctance is 6.6× larger than OTP friction. The ask is the opening move — it lands at median turn 3, and ~90% of gate messages give only a procedural reason ("for verification").

bot👋 Big plans or sudden needs? … check your eligibility in seconds — simple, quick, and hassle-free.
userKnow More
botKindly enter your 10-digit mobile number to receive the OTP required for verification.
— 84,834 conversations end on this exact pattern: curiosity met with a number demand.
botHappy to help with your dues! To begin, I'll need your 10-digit mobile number.
— the collections bot (dms): ~55% of its chats end here.

And the number step has a correction problem too — "Change mobile number" is the single most-tapped friction action: 8,331×.

Finding 02 · Where engaged users leak

Data-entry kills the journey; consent does not.

In the funnel, drops cluster on typing — salary −57%, manual demographics −81% — while consent and CIBIL steps barely move (190 → 190). Chat confirms it, per step:

Chat step, all botsDrop rateType
Company name48.4%data-entry
Salary / income47.5%data-entry
Mobile number43.5%data-entry
PAN37.0%data-entry
Name33.6%data-entry
Consent / T&C24.1%consent
OTP entry15.4%light confirm

A 2–3× gap. Users balk at typing, not at the credit pull.

The worst trap — the name field

The bot demands a name "as on PAN" but never states the rule. 57% of "name invalid" rejections punish a real single-token name, and 39.6% of those conversations abandon at the reject.

userTanish dhimanji
botThe provided name is invalid. Provide your complete name as stated on your PAN card.
— a documented loop: Atul → Atul Atul → Atul. → ATUL → 9 straight rejects, never told to add a surname → abandons.
Finding 03 · What "conversation" means today

It's a tap-through form, not a conversation.

83.4% of all user inputs are button taps. Only 13.1% of conversations contain any typed text — and 56.4% of what is typed is identity data the bot demanded, not talking.

BotTyped inputOf that, forced data-entryReads as
dms · collections5.7%30.8%near-pure tap wizard
iemi · EMI card19.1%75.6%wizard + heavy form
pl · personal loan20.3%54.7%wizard + form
b2b41.6%34.5%genuinely conversational

b2b proves a real conversation is achievable on the same surface — real product Q&A, not form-filling.

userCTCPD3238G
botThanks for providing your PAN. Select your gender from the options provided.
userNeelam Dabas
botThanks for providing your Name. Input your date of birth exactly in the DD/MM/YYYY format.
— a "typed" turn is usually the user feeding a field the bot asked for, not talking.
Engagement pays

Conversations where the user types anything reach the business end — offer, payment, lead — 44.9% of the time, vs 33.8% for pure button-taps. +11 points for engaged users.

Finding 04 · Where conversations actually end

Almost nothing ends well.

Last bot message, all 533,195 conversations

Phone / OTP ask34.0%
"Rate Us" / session closed9.3%
Ended on an FAQ / info answer7.1%
No bot text at all6.0%
Mid-form (name / PAN / DOB / income)5.1%
Stuck on T&C checkbox4.5%
Rejected / not eligible3.8%
Payment successful / disbursed1.6%

The places people land are walls, not finishes. For most bots Blu hands off to a payment page or the app and goes dark — it rarely witnesses completion.

1.6% Confirmed success Of 533,195 conversations, the share that ends on "payment successful / disbursed".
The silent no-shows

~22.5% of conversations (~120,000) have zero meaningful engagement — 6.0% are completely empty, 16.5% are bot-only monologues. That is larger than the entire phone-gate bucket. The opening surface mis-fires before anything begins.

Finding 04 · The walls, up close

Three walls: the fee, the offer, the rejection.

The fee step

82%Quit at the finish lineFunnel: add details 817 → fee paid 150. In chat, 78.5% of EMI Card journeys that reach the joining-fee step end right there.
botPay the one-time joining fee to get the Insta EMI Card. You are just 1 step away from getting your card.
— 78.5% never continue.

The offer screen

1 in 3Vanish silentlySilent drop after the offer: personal loan 34.2%, EMI Card 30.4%. Amount objections are handled ~94%; rate objections get a scripted refusal.
userAs a loyal customer for years, can u offer a lower interest rate?
botKindly Accept terms and condition and click Confirm to proceed.
— the objection is ignored.

The rejection

~90%Dead-stop at the "no"The EMI Card bot hard-rejects ~15% of its applicants (~20,079 conversations); only 1 in 10 types anything after.
botUnable to process your application. AVAIL NOW — Get Wallet Care and secure your debit/credit card with a coverage of up to ₹2 lakh & free annual Zee5 subscription.
— an upsell in the same breath as the no. No alternative, no retry timeline.
Finding 05 · What users actually want

They come to be served; the bot can only sell.

Existing customers show up in acquisition and collections bots wanting to service an account. 5,585 conversations carry an explicit servicing question — the dominant typed intent outside forced form entry (collections 2,089 · EMI card 1,364 · personal loan 1,201).

No bot can see the user's account data or their other Bajaj products — every ask is deflected to login, the app, or an acquisition CTA.

Requirement

Bots must connect — to each other and to account data — instead of operating as isolated single-purpose silos. Users expect one assistant that knows them.

userWhat is my loan status
botTo check your loan status, please log in to your account. Just click the Check Offer button…
userI just want to know how much loan remaining in my account?
bot…you'll need to log in to your Bajaj Finserv account. Just click the Check Offer button…
— a servicing ask, answered with an acquisition CTA. More verbatim asks: "I want loan statement" · "My loan total outstanding" · "Check my application status" · "Want to settle my loan amount."

The bot isn't drowning in out-of-scope demand either — the "I can only help with product searches…" line opened 11,529 conversations but was triggered by a real user query only 14 times in 533,195. Several of those deserved answers ("what is Blu").

Finding 06 · Trust and language

The bot raises the doubt — then answers it with nothing.

The loudest trust signal is that Blu offers the worry as buttons — and users tap them:

"How will this affect my privacy?"73,472×
"Is my KYC data safe?"22,504×
"I don't want to provide my PAN?"13,966×

Spontaneous typed distrust is rare (436 conversations, 0.08%) — users aren't confused about talking to a machine. They worry where their data goes, and when the doubt fires, the answer is generic:

botYour privacy is taken seriously…
— no "we don't sell data", no specifics, no link.

And it doesn't speak the user's language

The bot mirrors the user's language only 42.6% of the time — real users type Hindi and romanized Hindi inside the "English" locale. Hindi-locale users drop at the phone gate 47.4% vs 36.0% for English — an 11-point penalty.

userकितना किस्त है
botYou can manage your EMI repayments easily with our platform…
— an English reply to a Hindi question.
03
Part 03 · Phase 4 · The survey

Next: the survey.

The funnel and the transcripts say where users stop and what they say when they do. They cannot say why — or what would make them stay. Seven questions remain. Phase 4 asks them, directly, on the BFS user panel.

Phase 4 · The survey brief · Priority order

Seven things we still need to hear from users.

#What we need to learnThe finding it grew from
1What users would need to see or feel before they're willing to share a phone number.F1 · 34.5% end at the number ask
2What they actually want an AI for on a financial site — comparing, applying, checking status, paying, support — and how big the need to service existing accounts is.F5 · 5,585 servicing asks, unmet
3Where they would naturally look for AI help on the page — and where they'd click on a homepage with a few likely spots.F4 · ~22.5% never engage at all
4Whether tap-through flows even read as "AI", whether they help — and whether users would rather type and talk in their own words.F3 · 83.4% taps, not conversation
5How much data entry users will accept from an AI, and what makes sharing PAN or salary feel safe or wary.F2 · data-entry leaks 34–48%
6What an AI should say or do when it can't help — explain itself, pass to a human, suggest something else.F4 · rejection and dead-ends
7Whether they'd trust an AI's answers on money — rates, eligibility — or want to double-check with a human, and how reachable that human must be.F6 · the bot manufactures the doubt
Blu · The findings · End of part 02

Two phases in.
The evidence points; the survey confirms.

Two weeks on the BFS user panel, against the constructs the approach defined. The output is magnitude behind every pattern above — so the concepts that follow in Phase 5 rest on validated ground, not a hunch.

Versionv0.5 Date2026-06-11 AuthorArnesh Mandal