Open for newness

Info Icon
Info Icon
Info Icon
Info Icon

Learning where AI belongs

Learning where AI belongs

About

Schmooze is a Gen Z dating app that matches you on personality, not just looks. You can swipe on memes, search with AI, or talk to a AI matchmaker. It picks up on your vibe and finds people who get you. Grew to 2.2M+ downloads and 300K+ MAU, with 20-25% of them using it daily.

Role:

⋅ Everything design
⋅ Strategy

⋅ Everything design
⋅ Strategy

Team:

⋅ Dewanshi (Sr PD)
⋅ Asif (Lead PD)

⋅ Dewanshi (Sr PD)
⋅ Asif (Lead PD)

Timeline:

2023-26

2023-26

Desc:

Retention, Conversation, Matching. How AI helped us tackle each one.

Retention, Conversation, Matching. How AI helped us tackle each one.

The challenge

Dating apps run on human connection. But behind every match, conversation, moment of hope, there are real problems: users going quiet, chatrooms dying, people struggling to even start. We kept asking where AI could actually help, without making it feel cold.

We didn't have an AI strategy. We had a series of bets, some from data, some from gut, one straight from customer support tickets. Some worked, one didn't and some made real money. The pattern only made sense in hindsight.

How it began - AI companion

The hypothesis: What if a user with zero matches still had someone to talk to? It was our first real swing at AI, built to test engagement. We just knew that an empty dating app is a lonely place, and lonely places make people delete the app.

The approach: Three personas in the chatroom, each with a defined role: a stand-up comedian and a therapist. Roast, Rant and Sex-ed came later.

The persona work was never finished, it was maintained. We sampled and reviewed chat outputs on a regular cycle, tuned system prompts against what we found, and tightened guardrails as new failure modes surfaced. Model swaps were the highest-risk moment, because a better benchmark score doesn't mean a better. So we never cut over directly: every new model went out to a partial rollout first and was measured against the incumbent on tone consistency, guardrail violations, conversation length before it replaced anything.

The hypothesis: What if a user with zero matches still had someone to talk to? It was our first real swing at AI, built to test engagement. We just knew that an empty dating app is a lonely place, and lonely places make people delete the app.

The approach: Three personas in the chatroom, each with a defined role: a stand-up comedian and a therapist. Roast, Rant and Sex-ed came later.

The persona work was never finished, it was maintained. We sampled and reviewed chat outputs on a regular cycle, tuned system prompts against what we found, and tightened guardrails as new failure modes surfaced. Model swaps were the highest-risk moment, because a better benchmark score doesn't mean a better. So we never cut over directly: every new model went out to a partial rollout first and was measured against the incumbent on tone consistency, guardrail violations, conversation length before it replaced anything.

The hypothesis: What if a user with zero matches still had someone to talk to? It was our first real swing at AI, built to test engagement. We just knew that an empty dating app is a lonely place, and lonely places make people delete the app.

The approach: Three personas in the chatroom, each with a defined role: a stand-up comedian and a therapist. Roast, Rant and Sex-ed came later.

The persona work was never finished, it was maintained. We sampled and reviewed chat outputs on a regular cycle, tuned system prompts against what we found, and tightened guardrails as new failure modes surfaced. Model swaps were the highest-risk moment, because a better benchmark score doesn't mean a better. So we never cut over directly: every new model went out to a partial rollout first and was measured against the incumbent on tone consistency, guardrail violations, conversation length before it replaced anything.

What happened: 7% of WAU started talking to the bots, and they retained better on D0-D2 than users who didn't, though bot users were likely more engaged to begin with.

The finding I trusted more came from the split. Women were the power users, despite receiving far more matches than men. If loneliness were the driver, that's the cohort that should have adopted least. They adopted most. It reframed the feature for me: this wasn't a cold-start patch, it was a companionship product. They weren't using it to find a match. They wanted to talk without being judged.

The upgrade: Asked to improve entry into the feature, I gave the bots faces, names and personalities. Adoption went from 7% to 16% of WAU over time.

The tail was where it got interesting. Longest single chat, roughly 5,000 messages. Longest call, 82 minutes. Those are outliers rather than typical sessions, but the fact that a tail like that existed at all was the signal: a small group had stopped treating this as a feature.

What happened: 7% of WAU started talking to the bots, and they retained better on D0-D2 than users who didn't, though bot users were likely more engaged to begin with.

The finding I trusted more came from the split. Women were the power users, despite receiving far more matches than men. If loneliness were the driver, that's the cohort that should have adopted least. They adopted most. It reframed the feature for me: this wasn't a cold-start patch, it was a companionship product. They weren't using it to find a match. They wanted to talk without being judged.

The upgrade: Asked to improve entry into the feature, I gave the bots faces, names and personalities. Adoption went from 7% to 16% of WAU over time.

The tail was where it got interesting. Longest single chat, roughly 5,000 messages. Longest call, 82 minutes. Those are outliers rather than typical sessions, but the fact that a tail like that existed at all was the signal: a small group had stopped treating this as a feature.

What happened: 7% of WAU started talking to the bots, and they retained better on D0-D2 than users who didn't, though bot users were likely more engaged to begin with.

The finding I trusted more came from the split. Women were the power users, despite receiving far more matches than men. If loneliness were the driver, that's the cohort that should have adopted least. They adopted most. It reframed the feature for me: this wasn't a cold-start patch, it was a companionship product. They weren't using it to find a match. They wanted to talk without being judged.

The upgrade: Asked to improve entry into the feature, I gave the bots faces, names and personalities. Adoption went from 7% to 16% of WAU over time.

The tail was where it got interesting. Longest single chat, roughly 5,000 messages. Longest call, 82 minutes. Those are outliers rather than typical sessions, but the fact that a tail like that existed at all was the signal: a small group had stopped treating this as a feature.

The call experiment. Voice calls were exciting until we tried to monetise them. Adoption dropped, willingness to pay was lower. And unlike text, voice runs three separate models in sequence: speech-to-text, the LLM, then text-to-speech back. The compute cost per conversation was significant. Free wasn't sustainable, paid didn't convert. We pulled the feature and moved on.

Lesson: We went looking for engagement and accidentally built a retention engine. Not everything you build does what you built it for. Always look out the 2nd order effect.

The ice breaker - AI Dating Coach

The problem: Around 22% of matches never got a real conversation going. Two people matched, stared at an empty thread, panicked about the first line, and let it die.

Meanwhile support was getting messages that weren't support at all. Screenshots of chats. "How do I reply to this?" "What does this mean?" People were asking a human help desk to write their dating messages.

One fix for both. Put the help inside the chat, and take the load off support.

The problem: Around 22% of matches never got a real conversation going. Two people matched, stared at an empty thread, panicked about the first line, and let it die.

Meanwhile support was getting messages that weren't support at all. Screenshots of chats. "How do I reply to this?" "What does this mean?" People were asking a human help desk to write their dating messages.

One fix for both. Put the help inside the chat, and take the load off support.

The problem: Around 22% of matches never got a real conversation going. Two people matched, stared at an empty thread, panicked about the first line, and let it die.

Meanwhile support was getting messages that weren't support at all. Screenshots of chats. "How do I reply to this?" "What does this mean?" People were asking a human help desk to write their dating messages.

One fix for both. Put the help inside the chat, and take the load off support.

V1: We had no concrete idea on what people wanted help with, also have to ship fast. So I designed the plainest thing possible, a conversation starter and an open input box. Ask anything. Branded as Genie.

It didn't move usage much. But the input box was the real win, because every question typed into it told me what people actually needed. We shipped a feature and got a research log.

V2: Built from that filtering log. I dropped the open input and just gave people the things they kept asking for: conversation starters, profile summaries, date ideas, reply suggestions. I also moved it inside the chatroom as a tab, so help was right there with continuity.
Dead chatrooms dropped from around 22% to 16%. Support tickets for chat help fell too.

V3: where I was wrong? After a while i tasked with improving entry into the feature, I interviewed people who were still messaging support after V2. Three things came out: They don't read. Same questions, phrased differently every time. And most typed in their own vernacular language, not English.

That last one broke my earlier call. I'd removed the input because the data used to create custom prompts. What the data couldn't show me were the people whose question never fit a prompt, because they'd never have typed it in English in the first place. So I brought the input back next to the quick actions, and turned Genie into Dating Coach with persona, still retained custom prompts for flexibility.

It did improved TOFU significantly from all entry points, Support spam dropped further.

Keeping it alive: This was never ship and forget.

It reads real profile context to personalise replies: name, bio, bingelist, playlist, hometown. The AI had to sound like the user, and not everyone is fluent. System prompt is wrote with an example doc purely to stop it sounding Shakespearean. Target: average Indian Gen Z. It also picked up on the user's own bio for tone reference. In V3 memory was add, so an instruction given in one chat carried to the rest.

Same discipline as the companion bots. We sampled outputs, added guardrails as things broke, and dropped models that handled vernacular badly no matter how they benchmarked. Model and prompt changes always went out A/B and staged, measured against the old one before replacing it.

Keeping it alive: This was never ship and forget.

It reads real profile context to personalise replies: name, bio, bingelist, playlist, hometown. The AI had to sound like the user, and not everyone is fluent. System prompt is wrote with an example doc purely to stop it sounding Shakespearean. Target: average Indian Gen Z. It also picked up on the user's own bio for tone reference. In V3 memory was add, so an instruction given in one chat carried to the rest.

Same discipline as the companion bots. We sampled outputs, added guardrails as things broke, and dropped models that handled vernacular badly no matter how they benchmarked. Model and prompt changes always went out A/B and staged, measured against the old one before replacing it.

Keeping it alive: This was never ship and forget.

It reads real profile context to personalise replies: name, bio, bingelist, playlist, hometown. The AI had to sound like the user, and not everyone is fluent. System prompt is wrote with an example doc purely to stop it sounding Shakespearean. Target: average Indian Gen Z. It also picked up on the user's own bio for tone reference. In V3 memory was add, so an instruction given in one chat carried to the rest.

Same discipline as the companion bots. We sampled outputs, added guardrails as things broke, and dropped models that handled vernacular badly no matter how they benchmarked. Model and prompt changes always went out A/B and staged, measured against the old one before replacing it.

Reflection: The goal isn’t to write for people. It’s to reduce blank-cursor anxiety and help them find their words.

The accidental winner - People Finder

Where it came from: This one didn't come from a roadmap or a brainstorm. It came straight out of customer support. The same request kept landing in the queue. "I want to match with someone from my college." "From my community." "Someone tall." People were trying to search the app, and we hadn't given them a search box.

Where it came from: This one didn't come from a roadmap or a brainstorm. It came straight out of customer support. The same request kept landing in the queue. "I want to match with someone from my college." "From my community." "Someone tall." People were trying to search the app, and we hadn't given them a search box.

The approach: Natural-language search for people. Type what you're after and the AI surfaces profiles that fit. The index wasn't just bio text. Every profile photo got converted into a description, so a search like "girl with curly hair" or "guy with a cat" could actually match. Alongside that bio, interests, bingelist, playlist, all the same signals the AI companion and Dating Coach were already using. Same data, different job.

The approach: Natural-language search for people. Type what you're after and the AI surfaces profiles that fit. The index wasn't just bio text. Every profile photo got converted into a description, so a search like "girl with curly hair" or "guy with a cat" could actually match. Alongside that bio, interests, bingelist, playlist, all the same signals the AI companion and Dating Coach were already using. Same data, different job.

V1: The first version had no example prompts. Just some topic hints, career, hobby, shared liking, and a plain input. That was a deliberate call, for two reasons. One, we'd already learned from the AI companion that a static list of prompts backfires. People just pick from the list and stop thinking beyond the first try. If we were going to suggest searches, they had to be personalized to each user's own profile, and we weren't ready to build that yet. Another is time. We wanted to see how people actually searched before investing in it, so we rolled it out to a slice of users and watched.

What happened: About 70% of first-time users dropped off right at the search box. No prompt to get started, so anyone who wanted an easy entry point just left.

V1: The first version had no example prompts. Just some topic hints, career, hobby, shared liking, and a plain input. That was a deliberate call, for two reasons. One, we'd already learned from the AI companion that a static list of prompts backfires. People just pick from the list and stop thinking beyond the first try. If we were going to suggest searches, they had to be personalized to each user's own profile, and we weren't ready to build that yet. Another is time. We wanted to see how people actually searched before investing in it, so we rolled it out to a slice of users and watched.

What happened: About 70% of first-time users dropped off right at the search box. No prompt to get started, so anyone who wanted an easy entry point just left.

V1: The first version had no example prompts. Just some topic hints, career, hobby, shared liking, and a plain input. That was a deliberate call, for two reasons. One, we'd already learned from the AI companion that a static list of prompts backfires. People just pick from the list and stop thinking beyond the first try. If we were going to suggest searches, they had to be personalized to each user's own profile, and we weren't ready to build that yet. Another is time. We wanted to see how people actually searched before investing in it, so we rolled it out to a slice of users and watched.

What happened: About 70% of first-time users dropped off right at the search box. No prompt to get started, so anyone who wanted an easy entry point just left.

But the people who did search got genuinely creative with it. Those queries became the most useful thing we got out of V1. I used them to tune the search results and tighten the guardrails.

V2: give them a starting point, just not a generic one. I reworked the UI and added the personalised prompt suggestions we'd wanted from the start, built off each user's own profile. First-time drop-off at the search box went from 70% to around 35%.

Then the zero-result problem. Adoption climbed, then stalled somewhere else. 90% of users who got a zero-result search never searched again. One empty screen and they were gone. We resolved it by nearest relevant results,  Search for a specific college, you'd get people from nearby colleges instead. Ask for something too narrow, you'd get the closest thing that actually exists. No dead ends. which reduced the zero results while searching significanly.

But the people who did search got genuinely creative with it. Those queries became the most useful thing we got out of V1. I used them to tune the search results and tighten the guardrails.

V2: give them a starting point, just not a generic one. I reworked the UI and added the personalised prompt suggestions we'd wanted from the start, built off each user's own profile. First-time drop-off at the search box went from 70% to around 35%.

Then the zero-result problem. Adoption climbed, then stalled somewhere else. 90% of users who got a zero-result search never searched again. One empty screen and they were gone. We resolved it by nearest relevant results,  Search for a specific college, you'd get people from nearby colleges instead. Ask for something too narrow, you'd get the closest thing that actually exists. No dead ends. which reduced the zero results while searching significanly.

But the people who did search got genuinely creative with it. Those queries became the most useful thing we got out of V1. I used them to tune the search results and tighten the guardrails.

V2: give them a starting point, just not a generic one. I reworked the UI and added the personalised prompt suggestions we'd wanted from the start, built off each user's own profile. First-time drop-off at the search box went from 70% to around 35%.

Then the zero-result problem. Adoption climbed, then stalled somewhere else. 90% of users who got a zero-result search never searched again. One empty screen and they were gone. We resolved it by nearest relevant results,  Search for a specific college, you'd get people from nearby colleges instead. Ask for something too narrow, you'd get the closest thing that actually exists. No dead ends. which reduced the zero results while searching significanly.

The guardrails: Three hard rules: no searching by name, NSFW queries blocked, and every result still filtered through your existing match preferences. That last one mattered most. Without it, someone could reverse prompt their way into showing them people outside their settings. Blocking name search wasn't a limitation, it was the point. Most people don't want someone they know spotting them on a dating app. Letting anyone type a name and check would have made the app a spying tool for the people users are trying hardest to avoid. That one I wasn't willing to trade.

The guardrails: Three hard rules: no searching by name, NSFW queries blocked, and every result still filtered through your existing match preferences. That last one mattered most. Without it, someone could reverse prompt their way into showing them people outside their settings. Blocking name search wasn't a limitation, it was the point. Most people don't want someone they know spotting them on a dating app. Letting anyone type a name and check would have made the app a spying tool for the people users are trying hardest to avoid. That one I wasn't willing to trade.

How it made money: Search was free, but opening a full profile wasn't, user paid per profile. Good search result, yields people's willingness to pay. Which made it easy to see exactly what search was earning and keep on improving the search recommendation. Over time People Finder became 10-15% of daily revenue.

How it made money: Search was free, but opening a full profile wasn't, user paid per profile. Good search result, yields people's willingness to pay. Which made it easy to see exactly what search was earning and keep on improving the search recommendation. Over time People Finder became 10-15% of daily revenue.

Lesson: People Search isn't a decision-maker, and that's why it worked. Another AI feature that flopped tried to decide something for the user. This one just handed them a tool and got out of the way.

Matchmaker - Bet that didn't pay off

Riya Matchmaker: The most ambitious thing we built. By this point, swiping on memes and searching for people had already become the primary way users found matches on Schmooze. Riya was an experiment, meant to skip all of that entirely: just tell an AI what you want in a partner, and Riya find those profiles.

The approach: You'd talk to Riya once. She'd ask about you, then about what you wanted in a partner, and that single conversation is what generated your match recommendations. Framed like a concierge.

Riya Matchmaker: The most ambitious thing we built. By this point, swiping on memes and searching for people had already become the primary way users found matches on Schmooze. Riya was an experiment, meant to skip all of that entirely: just tell an AI what you want in a partner, and Riya find those profiles.

The approach: You'd talk to Riya once. She'd ask about you, then about what you wanted in a partner, and that single conversation is what generated your match recommendations. Framed like a concierge.

We launched voice-first. You'd call Riya and describe your ideal match out loud, Call took around 6 to 7 minutes to fully capture what we needed. Around 90% dropped off within the first minute. Talking to a bot about your romantic preferences turns out to be deeply, physically awkward. The intimacy of voice with none of the trust that earns it.

So later I added chat, as a medium to talk with Riya. Chat had a better funnel than call, only around 62% dropped off before answering half the question bank. which was decent enough to recommend some matches unlike <1min call data points.

The deeper problem: completing wasn't the same as answering well, people skipped questions or gave low-quality answers. We read some transcripts to understand why, and the pattern was clear: people gave minimal answers, moved fast, and didn't take it seriously. Thin input, thin recommendations.

Part of that was us. Some people were only comfortable answering in their own vernacular language, and the model we had at the time struggled with vernacular input. When Riya couldn't quite follow, people gave up trying rather than switch to English, Background noise derailed calls too, which made an already awkward thing worse.

We launched voice-first. You'd call Riya and describe your ideal match out loud, Call took around 6 to 7 minutes to fully capture what we needed. Around 90% dropped off within the first minute. Talking to a bot about your romantic preferences turns out to be deeply, physically awkward. The intimacy of voice with none of the trust that earns it.

So later I added chat, as a medium to talk with Riya. Chat had a better funnel than call, only around 62% dropped off before answering half the question bank. which was decent enough to recommend some matches unlike <1min call data points.

The deeper problem: completing wasn't the same as answering well, people skipped questions or gave low-quality answers. We read some transcripts to understand why, and the pattern was clear: people gave minimal answers, moved fast, and didn't take it seriously. Thin input, thin recommendations.

Part of that was us. Some people were only comfortable answering in their own vernacular language, and the model we had at the time struggled with vernacular input. When Riya couldn't quite follow, people gave up trying rather than switch to English, Background noise derailed calls too, which made an already awkward thing worse.

We launched voice-first. You'd call Riya and describe your ideal match out loud, Call took around 6 to 7 minutes to fully capture what we needed. Around 90% dropped off within the first minute. Talking to a bot about your romantic preferences turns out to be deeply, physically awkward. The intimacy of voice with none of the trust that earns it.

So later I added chat, as a medium to talk with Riya. Chat had a better funnel than call, only around 62% dropped off before answering half the question bank. which was decent enough to recommend some matches unlike <1min call data points.

The deeper problem: completing wasn't the same as answering well, people skipped questions or gave low-quality answers. We read some transcripts to understand why, and the pattern was clear: people gave minimal answers, moved fast, and didn't take it seriously. Thin input, thin recommendations.

Part of that was us. Some people were only comfortable answering in their own vernacular language, and the model we had at the time struggled with vernacular input. When Riya couldn't quite follow, people gave up trying rather than switch to English, Background noise derailed calls too, which made an already awkward thing worse.

Riya needed people to articulate what they wanted in a partner, and most people can't. The honest answer to that question is usually "I don't know." An AI that needs a clear brief can't do much with that.There's a second thing underneath it. Users don't actually want the choice taken away. When a bot hands you a shortlist, it quietly steals the part that makes a match feel earned.

Riya never got past that. With more time it might have, but the bet didn't pay off.

Reflection: AI tried to do the choosing for people, that didn't work out. Sometimes people don't always want things made easy, sometimes they want to be part of the process.