Zero setup AI voice

Creating an "it just works" experience for an AI receptionist service.

Overview
A self-service AI receptionist customers configure themselves, against the cold-start problem of a call flow they can't describe.
My role
Lead Product Designer. I defined the product experience and led the design of sign-up, configuration, and conversation design. I spoke to users, managed stakeholder engagement and feedback, and shadowed human receptionists to define call experience.
Zero setup AI voice

Impact Overview

  • Time-to-value dropped from weeks to minutes

  • The product scaled from 0 to $4M ARR in 2 years

  • People stopped being the blocker on getting customers and capturing revenue

If I were to ask you to describe the structure and flow of the last phone call that you had, I bet you can’t. If you were presented with all the options to configure all of it I bet you still couldn’t.

It’s hard to solve the cold start problem in software. It’s even harder when you don’t have a mental model for describing what you want.

The most important aspect for me was to deliver an “aha” moment (and therefore value) as soon as possible after payment. I made sure that Smith’s AI Receptionist just worked right out of the box. That took time-to-value from weeks to minutes, and the self-serve model it enabled grew the product from zero to $4M ARR in two years.

Every new customer cost days of people’s work

Onboarding a new customer started with a Typeform. Not very professional. Not very scalable. Then followed a sales call, a discovery session, and a few days of human-led configuration before they could even try anything. That assumed the teams involved got it right on the first try. This process cost several people’s time (and the business money) every time.

That is the cost product-led growth had to remove. Smith.ai had already handled tens of millions of calls with trained human agents, so the model worked; it just couldn’t scale one customer at a time. Law firms were about 75% of the customer base on the human reception service, against 25% all others. We didn’t initially target law firms specifically; we accommodated their needs.

Before and after of the sign-up flow, from an unactionable form to sign-up, payment, and dropping straight into the dashboard
Before: a form that produced no actionable items. After: sign-up, payment, and dropping directly into the dashboard.

The goal was that a customer could sign up and speak to a working AI receptionist in minutes. My vision for this went further: given the technical capabilities of the target audience (1-20 person legal practices), I decided it should also “just work” with no setup after someone signed up. Call the number. Hear your AI receptionist answer the phone as if it were you.

Not only did I need to design the configuration experience on top of the existing human receptionist system (which shipped faster but meant all of the complexity came with it) but it also needed to live in the same dashboard used for said service. Imagine lots of if → else → display:none;. It also meant elements that were reused needed to stay where they were or we were duplicating interface elements only to hide them. Engineering did not like that idea. I knew we’d need to address IA at some point but needed to validate product market fit first. No sense prematurely optimising.

I designed the first version of everything. From the v1 onwards I continued mainly on configuration while a senior product designer iterated on sign-up and onboarding.

Every call had a unique set of instructions

Most voice-AI startups were building from scratch and inventing what ‘good’ sounded like. Smith.ai had already taken the calls, but each of them had a completely unique set of instructions. There was no mass migration possible at all so automated pattern analysis wasn’t going to work.

An effort on the operations and updates team had defined some structure for those calls, and had the manual call instructions follow that pattern. That work was already done, so I leaned on it.

I sat with human agents over two weeks. Shadowing their calls to listen to how they spoke to callers, how they handled data capture, how they answered when they didn’t know the answer. I validated all of this with early customers to ensure the flow met their expectations and needs.

The structure was five call stages (Greet, Listen & Identify, Gather, Qualify, Actions) and five common caller types.

Caller Types

  • Potential new clients
  • Existing clients
  • Sales
  • Spam
  • All other calls
The distilled call flow from the human service, mapped in Figma against a previous self-serve attempt
The call flow distilled from documenting hundreds of real calls, mapped onto a remnant of a previous self-serve attempt to see what could be kept, redesigned, or was net new.

With the engineering team and access to Vapi, our orchestration tool, the CPO and I created demo accounts both using this flow (and also exploring other options), to see what gave the best results. This was before we had any sort of quality scoring or transcript sentiment evaluations, so it was a question of live demos, recordings and consensus to determine what sounded good.

The best results were for the AI when no flow was given, but a series of questions that needed to be answered in a specific order, and then operational instructions about how to request that information. Because the system needed to function both for our human receptionists and the AI, the instructions had to be written for a human, in an explicit order with explicit questions. Which is what we ended up with.

The existing receptionist service ran on a configuration system that was unique per customer and increasingly complex even for our internal teams. I couldn’t put that in front of a self-service customer at sign-up. They’d close the tab.

Research conversation analysis with the customer onboarding team
I ran research calls with three customer-onboarding team members to find the overlap between how they onboard a customer and what the system would actually allow/need to provide.

Stakeholders were excited to build a node-based flow for handling all of the complexity we had identified in the existing human receptionist service. I maintained that we needed something simple and approachable. People don’t think in node-based flows, they think in the order of a call.

The Head of Product had the senior product designer prototype both options in Figma, a simple form-style configuration and a node-based flow builder, and tested them with three small-business owners he was already interviewing for broad product-market fit. He handed me the research with a recommendation: the linear flow was easier to understand, but the node-based flow was easier to understand once someone had experienced the linear one. The simple version was how people learned the advanced one.

I reviewed all of the research calls myself. I didn’t use the work the senior designer had produced, because it was produced in a vacuum and lacked the constraints the actual system operated under. My philosophy for avoiding the cold start problem is that the product should function immediately and signpost where to change things, rather than rely on an existing configuration. I took the recommendation, built the linear flow out, and presented it to the CPO.

The call had to read top to bottom

Early concepts converged on a linear flow in the configuration UI. Any branching (qualified versus disqualified, transfer versus capture) lived inside the structure, not as a separate flow the customer had to wire up.

A caller pattern shown in the configuration UI as a single top-to-bottom linear conversation
Each caller pattern reads as one linear conversation, so a self-service customer follows it top to bottom instead of wiring up a flow chart. The trade-off: less branching power visible up front, for a setup someone can finish at sign-up.

The most complex of the five patterns was potential new clients. It mattered most to get right because it tied directly to new revenue for our customers. Get it wrong and we’d cost them business, which is a hard objection to overcome once they’re deciding whether to cancel.

The most complex caller pattern, potential new clients, with qualification and routing branches held inside one linear structure
The most complex of the five patterns. The branching (qualified vs disqualified, transfer vs capture) lives inside the linear structure, not as a separate flow the customer has to build.

Qualify sat fourth in the structure I’d inherited, after Gather. A caller answered a run of questions and then got disqualified, having spent minutes giving information that led nowhere, and they were frustrated by it. We could hear it. We reviewed real calls in an internal tool, and the product team and I reached the same conclusion: qualification was landing too late. I moved qualification to the top of the flow for prospective clients. A caller who wasn’t a fit now found that out early instead of at the end.

Watching real calls also told me what makes a call feel like a call: turn-taking. A human gap between turns is 200 to 300 milliseconds; a two to four second pause reads as reluctance or a dropped line. Retrieval and tool calls introduced exactly that kind of delay, so I treated latency as a design constraint from day one, rather than something engineering would optimise away after launch.

Humans cover the gap when they look something up: they keep talking, draw out their speech, or you hear them typing. I gave the AI the same affordance, a set of natural holding phrases to pull from while a tool call ran.

Verification was the default I got wrong in the safe direction. The AI confirmed the information it had captured back to the caller at the end of the call, which felt like the responsible default. It frustrated callers instead: capture, transfer, then re-confirm the same details. I made it a setting, then turned it off by default for new customers. Nobody has ever turned it back on. Information accuracy mattered less to the call experience than I had expected it to.

Frustrated Caller became the primary measure of product health

The whole team reached all of that by listening together in the call-review tool, and that was enough to act on. But it was manual. We had the data (call lengths, caller types) and never analysed it. If I did this again I’d run weekly cohort analyses by caller type, to back up what we were hearing or point us somewhere different.

What we did evaluate against the transcript was data capture: how many data keys were populated by the AI, how many by a human, how many corrected by a human. There was one for abrupt call ending, where the call just ended with no attribution. And there was Frustrated Caller, a basic sentiment analysis on the call.

Frustrated Caller ended up being the primary measure of product health. It only showed us if a caller was frustrated. Not whether the caller achieved their goal, not whether they achieved the customer’s goal, not whether we provided any inaccurate information or went down any wrong paths, not whether we greeted them effectively. I had many conversations with the VP of Operations about this.

Later I ran a workshop to bring alignment to Product and Design on what a successful call was. Everyone had different ideas about how to improve the system: no cohesive language, no cohesive voice. My aim was to bring them all together so everyone could have their say, and we could then group the results and prioritise them. Out of it came the first conversations about the AI quality score and the strategy around measurement that I proposed.

Clustered and prioritised sticky notes from the Product and Design workshop on what makes a successful call
The initial clustering and prioritisation from the Product & Design team for what constitutes a successful call.

The interface had to match the model in the customer’s head

The interface was simple because it matched the model in the customer’s head, not the architecture underneath. My goal: a customer signs up and can immediately call their AI, with zero configuration. No ‘set up your flow’ wizard, no configuration call, no support handoff to get them live. Sign-up captured everything I needed to pre-populate a working assistant (the business name), so the very first call wasn’t a tutorial. It was the AI representing the business, with the customer hearing exactly what their callers would.

The default AI receptionist configuration a customer gets with no setup
The default experience with no configuration required. It avoids the cold-start problem, but doesn't yet help a customer understand what they can configure, or what they should configure for the best call experience.

The principle I adhere to is simple: the product is the manual.

The configuration had a field for ‘what information do you want to capture from callers?’ I expected data keys: “date of accident”, “type of injury”, “court date”. These mirrored how we captured data on the human receptionist side. What customers actually did was write questions. “What date was the accident?” “Have you contacted another attorney?”

That looked harmless, but the AI read those entries as questions to ask verbatim, and explicit questions made the call worse than implicit capture. The call started to feel like an intake form being read aloud. Their mental model didn’t match the interface: they wanted to script the call; I wanted to capture information; two different goals landing in the same input field.

We were dogfooding the system ourselves too, on several demo lines and as the emergency call-in line for our receptionist staff, so we hit the same constraint while setting it up.

The original single capture field with a character limit, where customers typed full questions
Before: a single 'what else should we ask for?' field with a character limit. Customers wrote questions for the AI to read out verbatim ('What date was the accident?'). We needed a label for the data field.

The fix had three parts. Reduce friction for the right input. Show a real example of what works versus what people commonly try. Add validation that recognised the wrong shape and offered a path back to the right one.

Those capture fields were data keys, dropped into the system prompt verbatim, so a question in one broke the call in unpredictable ways; the AI got confused, and the summary came back with a question in it. Engineering couldn’t solve that at the time, so the surface that was the easiest to change was the interface. I wrote the regex myself: it blocks a question but allows a statement, with a message that says you only add the information you need; we ask the question for you.

The revised capture field with a worked example and validation that catches a question and redirects to information capture
After: force a key label, show an example, and catch the wrong shape. Not ideal, but Support stopped raising it and agreed it looked resolved.

The capture field had a better answer than validation and helper text, and I knew it at the time. Given how capable LLMs already were, I wanted a translation layer that took the customer’s wording and generated the data label, so a customer could write “What date was the accident?” and the system would store the key “date of accident” without their ever thinking about the difference. I built a demo and put it in the backlog. In a small team there were higher-priority features, so it stayed there. I saw the better solution, built enough to prove it, and it lost the priority call.

The same mismatch showed up in publishing. The first version of configuration had a draft state and a manual publish: you changed something, then pressed publish to put it live. Customers expected the opposite. They assumed every change saved and went live the moment they made it, and they missed the publish step completely, banner and all. Engineering found a way to save every change live, and auto-save was the first thing we changed. Build the interface around the model in the customer’s head, because the one in the system underneath is invisible to them.

The publish banner most customers missed, treated as a nag and ignored
The publish mechanism that most internal users and customers missed. My diagnosis was 'banner blindness': it resembled how advertising and upgrade nags are usually displayed, so people tuned it out.

Defaults made a complex system safe to hand over

Zero-config works because I captured the minimum and set sensible defaults and guardrails for everything else.

The greeting was the first thing a caller heard, so it was the first control I designed. I kept it prescriptive at first: three options to choose from, collected from real customers on the human service and confirmed with Product Marketing. Once I had feedback from customers, I opened it up to free text, with constraints.

The original greeting control and the first iteration allowing a custom greeting with a call-recording disclosure
The original greeting and the first iteration: more visually compelling, and allowing custom greetings with a legal acknowledgement of the call-recording disclosure.

The obvious move for the legal segment was to make the one flow conditional. I did the opposite.

The most common use case from the target legal practice size was either capture all their information and let them through, capture all their information and schedule an appointment, or capture all their information, make sure they need one of these specific services and are in a specific state, and then transfer them through. Variations of that theme. None of that required complex branching logic. All of it could be conveyed in a linear flow.

What it needed was allowing for the possibility that at any one of those junctures that could be conceived of as a branch, the answer could be no. So there’s a UI element at the end of those sections for when the answer is basically no.

The simplest mechanism was that we take a message and we pass their information on, and we tell them that we will pass the information on. That way the caller feels heard. We don’t give the caller incorrect information: the wording was closer to saying it doesn’t look like we can help in this specific case, but we’ll pass the information on in case we can, or in case we can refer them to another practice. Then we ask if there’s anything else they’d like to tell us, and end the call.

The customer got a summary sent by email with the disposition as disqualified, and a link to the dashboard where they could review all of the structured fields, listen to the full recording, and search the full transcript.

I added a single caller-type setting, so a customer who only ever gets one kind of call could switch off the existing-client question and the spam and sales branches entirely. Single caller types worked around some of the trouble I was having with the automatic caller-type detection based on “How can I help you?”, but they solved a real customer need in the meantime.

The single caller-type setting in the configuration modal
Single caller type: a lot of customers asked for this. My diagnosis was issues with the 'How can we help you?' detection and routing. We hard-coded a temporary 'Are you a new or existing client?' question as a stopgap.

Because the publish step was dropped, the configuration had to be self-contained, which meant almost everything became a modal. That is where the caller-type setting lived, and there was only one place to put it that made sense: I broke a UX convention and put the checkbox in the footer of the modal. It was the only spot, outside the configuration itself, that kept the whole decision inside one self-contained view. The configuration still had to be represented in the dashboard; the product is the manual.

The strongest idea I had is one I never shipped

The strongest idea I had for the product is one I never shipped. The core system prompt, the instruction set the AI ran every call against, was too restrictive. As the underlying models improved, the caller experience stayed capped by a prompt written for weaker ones. I argued for a full re-evaluation, repeatedly, and I built demos to show what a looser prompt made possible.

The bet was never realised. The prompt was owned by an engineer, and when it was eventually revisited it was for cost and brevity, not the caller experience I was making the case for. I had the demo and the conviction; what I didn’t have then were the influencing skills to turn a working prototype into a priority.

The strongest reaction we ever got came from an early demo line, before the guardrails. A caller in the US asked about California trade law, whether they had to disclose something on their packaging. The AI answered; they asked a more detailed follow-up; it answered again. A completely natural conversation, where the caller forgot they were talking to an AI.

“Holy ****, whoever made this is a genius!” — a caller to our early demo line

I don’t think we ever captured that wow again. It came from the model answering a real question off its own knowledge, not from anything the customer had configured; the guardrails were what kept it from happening once the prompt took over.

What I’d do differently is instrument the case. The argument rested on demos and my own read of the calls; it needed caller-experience data sitting next to the cost-and-brevity numbers the decision was actually made on, so it wasn’t my judgement against someone else’s.

People were no longer the blocker

All of the human cost survived. It just meant we could tailor it to higher-value customers. Self-serve captured the lower end of the funnel, and anybody that needed higher volumes of calls was given the white-glove treatment. Sales still had a job. They would still be selling the service, and they would then ask customers to sign up using the self-serve. Larger accounts had an onboarding call to familiarise themselves with the software.

The important thing is that people were no longer the blocker for us getting customers and capturing revenue. A payment link had to be sent out by email previously. Now payment was in flow.

The voice assistant launched in late 2023, and the self-serve configuration was what let it grow product-led instead of sales-assisted, from zero to $4 million ARR in two years. Time-to-value went from weeks to minutes. Sign up, configure inside the flow, take real calls the same day.

Let's talk

Interested in working together?

Get in touch