“How do we make the chatbot work harder” is several questions, not one. Each gets a different method.
| Sub-question | What it asks | Method |
|---|---|---|
| Discoverability | Do users see the entry point? | Funnel data + usability tests |
| Comprehension | Do they understand what it offers before clicking? | Concept tests + usability tests |
| Intent match | Are they at the right intent step when they see it? | Segmented funnel + usability tests |
| Conversation value | Did the conversation itself feel valuable to the user? | Transcripts (where available) + usability tests |
| Handoff | Does the bot hand off cleanly into the product flow? | Funnel analytics + usability tests |
| Architecture | One Blu or N bots — which fits user mental models? | Survey preference + concept testing |
| Trust | What trust signals do users need for AI in a financial context? | Trust probes in usability + survey |
| Segment fit | Do high-intent shoppers and low-intent browsers want the same thing? | Segmented funnel + segmented survey + segmented usability |
Without pre-registration, any result can be back-rationalised, and the research stops being able to change anyone's mind. The sub-question set, the success criteria, and the decision rules are signed off before any data is gathered.
Phases 1–3 build the initial picture in parallel. The survey deepens it. Concepts are tested against the live experience. Synthesis ties evidence back to the questions and produces the recommendation.
| Purpose | Map current behaviour on the live Blu surface; find where the funnel leaks. |
|---|---|
| Methods | Funnel construction from Google Analytics custom events; cross-check with the CleverTap funnel; segmentation by product. |
| Inputs | Google Analytics events; CleverTap funnel data. |
| Outputs | Annotated funnel with drop-offs per product; instrumentation gap list. |
| Time | 1 week |
| Purpose | Read what users say to the bots and where the bots fail. |
|---|---|
| Methods | Stratified transcript sampling; qualitative coding; intent and failure taxonomies. |
| Inputs | Anonymised chatbot transcripts with outcome tags. |
| Outputs | Intent taxonomy across bots; ranked failure-mode catalogue; user-phrasing glossary. |
| Time | 1 week |
This phase is conditional on transcript access. The approach proceeds without it if transcripts are not accessible — Phase 5 surfaces similar insights through usability sessions, at lower scale.
| Purpose | Learn from what others have shipped and what the literature says about trust and adoption. |
|---|---|
| Methods | Competitive teardown; structured literature review. |
| Inputs | Public products, app store reviews, academic search (TAM / UTAUT for banking AI, conversational AI KPI frameworks). |
| Outputs | Competitive matrix; literature digest. |
| Time | 1.5 weeks |
| Purpose | Discover user need around conversational AI for financial products; validate the patterns surfaced by the funnel, transcripts, and competitive sweep. |
|---|---|
| Methods | Quantitative survey on the BFS user panel. |
| Inputs | Working hypotheses; insights from Phases 1–3; panel access. |
| Outputs | User-need patterns named and ranked; magnitude-of-preference report; adoption-construct scores by segment. |
| Time | 2 weeks |
1 · Need discovery — where users look for help today; what would make them try a conversational AI; what they expect it to do; what they would not trust it to handle.
2 · Validation — magnitude-check the patterns the funnel, transcripts, and competitive sweep already pointed at — so the concepts that follow rest on more than a hunch.
| Purpose | Build concepts informed by everything learned so far; test them against the current live experience. |
|---|---|
| Methods | Usability testing; concept testing on paper or clickable prototypes. |
| Inputs | Concept directions informed by Phases 1–4; prototypes (one or several variants); research-panel access. |
| Outputs | Per-prototype findings against the live control; ranked direction preferences; trust signals. |
| Comparison | The live experience is included as control. Every prototype is judged against it — the A/B comparison this team runs. |
| Time | 3 weeks |
| Purpose | Pull the outputs of every prior phase together against the working hypotheses; write the recommendation. |
|---|---|
| Methods | Hypothesis-by-hypothesis review of evidence; prototype-vs-live-control comparison. |
| Inputs | Outputs from Phases 1–5. |
| Outputs | Recommendation memo. |
| Verdict | Each hypothesis marked supported, killed, or undetermined. |
| Time | 1 week |
Supported — the prototype direction that won, with the user evidence behind it. Killed / undetermined — what was tried and what would unblock the next decision.
Six phases of evidence — funnel, transcripts, secondary, survey, prototypes against the live — distilled into design directions that won against the live control, paired with the user evidence behind them and the questions still worth opening next.
Phases 1 and 2 are executed — the funnel mapped across Gold Loan, EMI Card, and Personal Loan, and 533,195 real conversations read and coded. Six findings follow. Each carries its numbers and the users' own words.
Every funnel claim is re-measured inside the chat data — of users who reach a step, how many end there. Where chat confirms a funnel point, it makes it stronger and adds the users' own words. Every quote in this part is verbatim transcript text.
34.5% of all conversations (183,773) end at the mobile-number / OTP ask. The funnel sees the same wall — EMI Card loses 45% at login before anything happens. Splitting the gate apart shows what actually fails:
| Number ask, no number given | 14.1% |
| OTP-framed ask, no number given | 12.2% |
| No number ever entered — reluctance | 26.3% |
| Gave a number, died at OTP entry | 4.0% |
Reluctance is 6.6× larger than OTP friction. The ask is the opening move — it lands at median turn 3, and ~90% of gate messages give only a procedural reason ("for verification").
And the number step has a correction problem too — "Change mobile number" is the single most-tapped friction action: 8,331×.
In the funnel, drops cluster on typing — salary −57%, manual demographics −81% — while consent and CIBIL steps barely move (190 → 190). Chat confirms it, per step:
| Chat step, all bots | Drop rate | Type |
|---|---|---|
| Company name | 48.4% | data-entry |
| Salary / income | 47.5% | data-entry |
| Mobile number | 43.5% | data-entry |
| PAN | 37.0% | data-entry |
| Name | 33.6% | data-entry |
| Consent / T&C | 24.1% | consent |
| OTP entry | 15.4% | light confirm |
A 2–3× gap. Users balk at typing, not at the credit pull.
The bot demands a name "as on PAN" but never states the rule. 57% of "name invalid" rejections punish a real single-token name, and 39.6% of those conversations abandon at the reject.
83.4% of all user inputs are button taps. Only 13.1% of conversations contain any typed text — and 56.4% of what is typed is identity data the bot demanded, not talking.
| Bot | Typed input | Of that, forced data-entry | Reads as |
|---|---|---|---|
| dms · collections | 5.7% | 30.8% | near-pure tap wizard |
| iemi · EMI card | 19.1% | 75.6% | wizard + heavy form |
| pl · personal loan | 20.3% | 54.7% | wizard + form |
| b2b | 41.6% | 34.5% | genuinely conversational |
b2b proves a real conversation is achievable on the same surface — real product Q&A, not form-filling.
Conversations where the user types anything reach the business end — offer, payment, lead — 44.9% of the time, vs 33.8% for pure button-taps. +11 points for engaged users.
| Phone / OTP ask | 34.0% |
| "Rate Us" / session closed | 9.3% |
| Ended on an FAQ / info answer | 7.1% |
| No bot text at all | 6.0% |
| Mid-form (name / PAN / DOB / income) | 5.1% |
| Stuck on T&C checkbox | 4.5% |
| Rejected / not eligible | 3.8% |
| Payment successful / disbursed | 1.6% |
The places people land are walls, not finishes. For most bots Blu hands off to a payment page or the app and goes dark — it rarely witnesses completion.
~22.5% of conversations (~120,000) have zero meaningful engagement — 6.0% are completely empty, 16.5% are bot-only monologues. That is larger than the entire phone-gate bucket. The opening surface mis-fires before anything begins.
Existing customers show up in acquisition and collections bots wanting to service an account. 5,585 conversations carry an explicit servicing question — the dominant typed intent outside forced form entry (collections 2,089 · EMI card 1,364 · personal loan 1,201).
No bot can see the user's account data or their other Bajaj products — every ask is deflected to login, the app, or an acquisition CTA.
Bots must connect — to each other and to account data — instead of operating as isolated single-purpose silos. Users expect one assistant that knows them.
The bot isn't drowning in out-of-scope demand either — the "I can only help with product searches…" line opened 11,529 conversations but was triggered by a real user query only 14 times in 533,195. Several of those deserved answers ("what is Blu").
The loudest trust signal is that Blu offers the worry as buttons — and users tap them:
| "How will this affect my privacy?" | 73,472× |
| "Is my KYC data safe?" | 22,504× |
| "I don't want to provide my PAN?" | 13,966× |
Spontaneous typed distrust is rare (436 conversations, 0.08%) — users aren't confused about talking to a machine. They worry where their data goes, and when the doubt fires, the answer is generic:
The bot mirrors the user's language only 42.6% of the time — real users type Hindi and romanized Hindi inside the "English" locale. Hindi-locale users drop at the phone gate 47.4% vs 36.0% for English — an 11-point penalty.
The funnel and the transcripts say where users stop and what they say when they do. They cannot say why — or what would make them stay. Seven questions remain. Phase 4 asks them, directly, on the BFS user panel.
| # | What we need to learn | The finding it grew from |
|---|---|---|
| 1 | What users would need to see or feel before they're willing to share a phone number. | F1 · 34.5% end at the number ask |
| 2 | What they actually want an AI for on a financial site — comparing, applying, checking status, paying, support — and how big the need to service existing accounts is. | F5 · 5,585 servicing asks, unmet |
| 3 | Where they would naturally look for AI help on the page — and where they'd click on a homepage with a few likely spots. | F4 · ~22.5% never engage at all |
| 4 | Whether tap-through flows even read as "AI", whether they help — and whether users would rather type and talk in their own words. | F3 · 83.4% taps, not conversation |
| 5 | How much data entry users will accept from an AI, and what makes sharing PAN or salary feel safe or wary. | F2 · data-entry leaks 34–48% |
| 6 | What an AI should say or do when it can't help — explain itself, pass to a human, suggest something else. | F4 · rejection and dead-ends |
| 7 | Whether they'd trust an AI's answers on money — rates, eligibility — or want to double-check with a human, and how reachable that human must be. | F6 · the bot manufactures the doubt |
Two weeks on the BFS user panel, against the constructs the approach defined. The output is magnitude behind every pattern above — so the concepts that follow in Phase 5 rest on validated ground, not a hunch.