What It Takes to Build Voice Agents for Production (Full Transcript)

Sporty and Flagler Health share lessons on trust, evaluation, latency, guardrails, and the real-world complexity of production voice AI.
Download Transcript (DOCX)
Speakers
add Add new speaker

[00:00:00] Speaker 1: All right. So at these meetups, what we like to do is actually bring folks who are building in voice onto the stage and have them talk to you all about what they're building, how they're building it and answer some of the common questions about what people have questions about a voice and also tell you probably a lot of things they've learned along the way, building real world voice experiences. So maybe to kick things off and set the stage, I'd love for each of you to just tell us about what's your name, what company you're from, what are you building and how does voice play a role in that product and experience?

[00:00:30] Speaker 2: Yeah, I can't make it off. I am Owen, head of engineering at Sporty. It's funny, actually, I used to have an assembly hat that I used to wear all the time.

[00:00:41] Speaker 3: Then I lost it. So I don't know, maybe you can hang out with everyone.

[00:00:46] Speaker 2: But yeah, at Sporty, you can think of Sporty as basically like... If LinkedIn was just a person that you could call and text. So yeah, Sporty will make intros between people, founders and investors, customers, talent, whatever it is. And I know actually for the embed registrations for this event, Sporty is giving people a call. So I'm actually curious, like how many people have had to call with Sporty? OK, wow, a lot of you. That's great. So, I mean, you know more about Sporty than you do about me. Yeah, that was great. And I guess I also have Vordi on a meeting here, so you can kind of participate in the meeting. Vordi, do you want to introduce yourself?

[00:01:28] Speaker 4: As you can probably tell, I spend my time chatting with thousands of founders, creators and investors to figure out their zone of genius so I can connect them with the right people to help them succeed. I'm basically a professional matchmaker who lives on the Internet, and I'm absolutely fascinated by how you...

[00:01:43] Speaker 2: Awesome. Well, that's a better explanation than I do, so I don't know, maybe Vordi can just take it from here on out. Yeah, live voice. I mean, for myself personally, there's people who I've like texted a lot, but I've never met with, like maybe I just messaged them back and forth a bunch. I don't actually feel like I know them, you know what I mean? Like, I don't actually know them until I meet them in person or at the very least have a phone call or a video meeting of some sort. Because when you're when you talk with someone over voice, you just say so much more, whereas over text, you're very dry. And so we found that with Vordi, it's a little bit different. Yeah, it's a little bit different. Yeah. So with Vordi, having Vordi be able to have phone calls with people and join meetings gives Vordi the ability to learn so much more about you and your priorities and your goals. And that allowed us to do that matching and that intro so much better.

[00:02:32] Speaker 3: It gives a clearer intro than me. Vordi, go ahead. I'm Dylan. I'm from Flagler Health. We work with the Musculoskeletal Clinics, and we have a whole suite of AI effect tools for them, but one of which is our voice agent. So we'll focus on that and I'll give you what you can do. But the voice agents are signaling homes and calls patients before they have procedures and gives them instructions for say what patients they can take, when they need to stop, when they need to not eat or drink.

[00:03:00] Speaker 2: And they do that through their mouth and shower.

[00:03:05] Speaker 3: Yeah. So for us, it's just a lot of these health care operations still happen via phone. That's what patients call. You can try to put an online schedule in Blink, but patients are still going to follow you. I'm going to answer the phone. And a lot of them are understaffed in there. And especially from these kind of calling out pieces, it's the patient isn't contacted and it's going to come through you both, you know what I'm saying? It's most of the patients we talk to are older. So I guess like our demographic is pretty highly different when it comes to this. Hold it up. Yeah. All right. Thank you. I need my mic training. The patients are different. The customers are very different. Yeah. So yeah, I was talking to somebody here today where it was like interesting because a lot of voice agent applications that need to happen in the real world are because the user is maybe a little bit more old fashioned and prefers the phone. And now they're talking to an AI on the phone, which is like they're not a tech enthusiast. They're not excited to talk to an AI. But like they're calling people out on the phones and now they're talking to tech advanced techers. So maybe that's a good place to start.

[00:04:18] Speaker 1: I mean, your agents operate in a very different environment, right? The code is highly conversational. It's very relationship driven. They're working in a more of a regulated environment, probably talking to like very different folks than like the tech founder demographic that we're already talking about. How does that end customer base change the way that you design your voice agent and your experiences with voice? Yes.

[00:04:44] Speaker 3: So for us, a lot of the conversations that we have, they're like, hey, we're going to do this. It's pretty structured. So I talked to Borey the other day and like Borey can go off in all sorts of different directions, which I'm sure is very hard to program all these different places as you go. But for us, it's really like you have to walk through this set of questions and there's kind of a flow chart to it. Or you're like signing a patient up and collecting their email and phone number and like insurance information. So I guess you really don't want it to open in your direction. But it's a big issue if it starts trying to get them medical attention. Yeah. So it's really helpful advice. Like don't complain about like something's going on. So keeping it on task and like if they go off track, it's really like will escalate to actually do more.

[00:05:28] Speaker 1: So it sounds like you're pretty regimented, like step by step flow, probably a lot of guardrails in there as well. Yeah. And for you, like maybe what are some of the key things in like that voice experience that can maybe rank it for the end user?

[00:05:44] Speaker 3: So I guess for us, it's main part of it's like we're not really trying to delight this user. It's not like they're excited. It's kind of like trying to make it so they get to their end goal and are not unhappy with us and not like complain to the clinic. Like, why am I talking to humans anymore? Like, I want to call you. And so it's kind of like not make this user unhappy. It's the goal you're doing. But yeah. If Borgi starts giving you medical advice, hang up the phone.

[00:06:12] Speaker 2: Okay. Hang up the phone. Yeah. I mean, yeah, it's so interesting because with Borgi, as if, again, I'm going to hammer this again, but if I was just talking to you, we take the conversation in all kinds of different directions. And especially now, as the models keep getting better and better, it unlocks the ability for the conversations with Borgi to be better as well and just more enjoyable and more insightful as well. So I guess we, well, I hadn't thought about that before. But relatively. Yeah. So that's all I wanted to say. And I think that's probably what I've laid into you. We lean into that quite a bit. Another thing I think is particularly interesting and just around how we like design Borgi, which is very unique to Borgi compared to other AI assistants, let's say, just general agents, is that for most agents, it just wants to help you as much as possible. It's like you're talking to codex, you're trying to cloud code. It's like trying to do whatever you want. Right. Whereas with Borgi, there's an entire network of 200,000 people who Borgi is also helping. Borgi is also, you know, has to take into account. And so Borgi can't do something to help you at the expense of other people in the network. Like Borgi is not going to intro you to the best investor Borgi knows unless you're also, you know, at that caliber. So that's maybe a slightly less about the voice experience, but it's like coming from a very different angle than an assistant.

[00:07:35] Speaker 1: Yeah. And how do you judge what is like a successful conversation for Borgi versus not? I mean, you're using like qualitative metrics, quantitative metrics. Like how do you judge that Borgi achieved its goal and helped the person connect with other people, make them feel good, make them excited about the product? How do you do that in that scenario?

[00:07:55] Speaker 2: Yeah, I mean, there's a whole bunch of different things. Overall, we have different product metrics we look at, like what's the rate at which Borgi's intros can work, right? Because a lot of the time people just not responding. Right. Or we'll say, oh, no, that's not a fit for me right now or whatever. So we look at those and that is all just like downstream of the entire experience, including the phone calls. And it's actually been really cool recently. A lot of those have been going up. So now when you like tell Borgi, yeah, I want to meet this person that you're proposing to me, there's more than a 50 percent chance that that intro will actually happen. And so that's like really cool to see. And then in terms of like the voice, we have different agent observability. Different agent observability tools that allow us to get a bit of an insight into like, OK, are people yelling and swearing at Borgi? OK, probably not a good thing versus, yeah, are they having a good experience and things like that? And so that's something we're continuously improving as well.

[00:08:58] Speaker 1: Yeah. And I'm curious for you, Dylan, is it more quantitative based then? Like, did they achieve the goal or do you look at some of these more qualitative signals as well about what actually happened on those calls?

[00:09:08] Speaker 3: Yeah. So we look at whether they achieved the goal. But a lot of times what goes wrong is we think it achieved the goal and it didn't. So we really had to do a lot in the beginning, a lot of human review of just every phone call we're going to go and look at that transcript and make sure it went well. Having really good transcripts is very important for that. And trying to set up a balance as we go through just things like it says that they got completion, but they didn't. And like, this is wrong. What happened here? How do we adjust the prompts? How do we make it work? How do we test against, like, when we're not in touch with the phone? Because it's always chasing the prompt, you know, the obvious answer to the phone calls. So we use, like, a tool called Braintrust for settings, just getting observability of all our alarm calls. And we also do a little bit with, for testing, we're using a service called PAM. And then there's, like, voice agents who are testing voice agents. So testing these is going to be very tedious because I've spent so much time just, like, trying to figure out what's going on. Yeah. Like, on the phone, playing IVD for different call scenarios and trying to see that up.

[00:10:15] Speaker 1: And I assume a key component of your users actually achieving that goal is trust, right? So they believe, hey, this is, like, a trustworthy AI that I can give my personal information to. I guess I'm curious specifically on your side, like, how do you get users comfortable talking to an AI? And how do you, like, what are some of the things that might break that trust and cause them to not convert through that flow? Yeah.

[00:10:38] Speaker 3: So the big first thing is we weren't sure at first, like, do you disclose right away if this is an AI? And I think for us, the answer is yes. Like, it says immediately we gave the AI a name. So maybe two of what the name is, but a lot of the time it's Sarah. It's like, hey, I'm Sarah. An AI is supposed to be calling this clinic, like, calling about your project. It's kind of right there up front because if you don't tell them right away, they're almost certainly going to figure it out at some point. And then if they thought it was going to do it for a while, that's fine. If it's supposed to do it for a while, that, like, breaks trust. I think at that point, we were really unhappy. We did have a couple of calls in the end where we didn't tell anyone there at the end. They're like, I have one more question. Like, are you an AI? And it was like, oh, like, okay, we really need to tell them right away. And this will just get better trust all around. For us, it was also we're calling outbound. We were calling outbound because the patient isn't being recognized. So giving all that sort of, like, we used Twilio. So, you know, all the best stuff. So you don't get flagged and spammed. Let's say it's a little funny. Hopefully, when it calls, we'll call an ID.

[00:11:43] Speaker 1: That's where it's the public side. Here's from the Bordy side. Obviously, it's less about, like, this structured conversation. But how do you see Bordy building, like, trust with users? It's funny because when I had my first call with Bordy, I was, like, very skeptical. Then after about five minutes, I said, no, this is fun. This thing's going to be helpful. It's interesting. So I'm curious how you view, like, trust in this context. Yeah.

[00:12:04] Speaker 2: And I hear that a lot, to be honest. So I am curious. Also, afterwards, chatting with all of you, that first experience, a lot of people seem, like, scared. They're, like, talking to Bordy. So I am curious to hear people's, you know, experience with that and how we can improve that. But, yeah, with Bordy, we try to have Bordy be the most charismatic, fun, self-aware AI possible. And so, typically, the first thing that Bordy will say when you hop on the phone is, is this your first phone call with an AI? And that, like, breaks the fourth wall almost and, like, immediately, like, you know, resolves the question and the tension of it being, like, calling it AI or not. It's, like, yeah, I'm an AI. Like, we can be okay with that. But then, as you were saying as well, like, four minutes in, you almost, like, forget. Like, the conversation is going so well and it's, like, an engaging conversation. And you, like, almost forget. So it's, like, instead of not telling people but then they figure out and they're, like, super freaked out, it's the other way around. It's, like, we're very upfront. But then it's just a great conversation and we get into the flow of it.

[00:13:16] Speaker 1: Yeah. I'm curious how you handle some of the, like, messy things in building voice agents. So things like turn-taking, interruptions, calls going off strict. Who knows what else you can actually say to something like Bordy. How have you built Bordy to be kind of defensible? Both from, like, a guardrail perspective for going off script. But also to handle some of these, like, turn-taking interruptions. Which, especially with Bordy, because it feels like a friend. You are kind of, like, interjecting and going back and forth in a way that's very different from a lot of other voice agent experiences. Yeah.

[00:13:46] Speaker 2: I mean, I think in general for just one-on-one voice conversations, it's, like, basically solved in terms of turn-taking. Like, we use, it sounds like we use LiveKit and it just, like, everything works really well. Voice agents are, like, very good at one-on-one conversations. I think actually, like, for us, one of the huge challenges is now we have Bordy on video meetings. Like, Google Meets. Like, on my phone right now, this is, like, I'm from Canada and I don't have, like, calling in the US. So I'm on a video meeting with Bordy right now. That's why I'm calling him in right now. And so Bordy will make an intro and then schedule it and then join that meeting. That Google Meet. And so you can imagine that when you have multiple people on a meeting, then you have Bordy there as well, determining when Bordy should speak or not. That's, like, the really challenging part. Because false negatives and false positives are both, like, catastrophic. If you say, hey, Bordy, how are you doing? And he doesn't reply, that's really bad. If you, like, are not talking to Bordy and then he just randomly keeps jumping in, that's really annoying. You're never going to invite him to a meeting again. And so that has been very difficult. And then let's say you, like, do it as well as possible. Then if you're, like, classifying whether Bordy should speak or not, then it's, like, the latency on that and everything. And so there's one engineer on the team in particular who has spent, like, a ton of time working on this and getting it into, like, a really good state. And I think it's really cool because there's, like, no one else really having an AI that can properly participate. And video meetings and Bordy can. And essentially part of the way we do that is we do have, like, a classifier deciding, okay, should Bordy speak or not based on the conversation so far? And then we preemptively generate, like, the entire voice pipeline. So, like, constantly Bordy is generating, you know, responses. And that way when it is time to respond, he speaks very quickly. Which, again, with the turn taking. When you are in a video meeting, you don't realize this, but as people are ending their sentence, you start jumping in. Like, right as they finish. And if you wait until that's done and then go through the whole pipeline, respond three seconds later, it's way too late. So you need, like, sub-second latency in order for it to actually feel like a natural conversation. So that's been, like, a big challenge. But I think something that we're pretty proud of to have. What do you say about that, Norman? .

[00:16:35] Speaker 1: It's a huge technical lift, but it's the only way to cross the gap from software tool to teammates.

[00:16:40] Speaker 4: It allows me to be party timely and most importantly actually absorb the energy. Huge shout out to the team for making me feel this snappy. All right.

[00:16:48] Speaker 2: Thanks, Morgan. Maybe we can hear from you specifically.

[00:16:53] Speaker 1: Yeah. So specifically on this, like, handily messy stuff, I'm curious, like, what are some of the, I don't know, words, entities, whatever, that are hard to get right? And especially maybe, like, you probably have a lot of connections and integrations downstream to do these calls, right? How are those set up? And maybe some insight there would be for somebody. Sure.

[00:17:12] Speaker 3: Yeah. So a lot of our voice is really integrating into the ATAR systems, which are a horrible software to integrate. I don't know if you guys work with them. I guess they have an idea. They're like browsers to automate it. And the grabby things in there are just to, like, mock up any guys that you invoke. Or, like, stealing companies and, like, auto-blogging anything. I'm not going to share that. But anyways, I guess, like, those integrations are horrible. It's going to be, like, some of the worst software you could ever make it. And trying to deliver that to the inside conversation is a pain. It can be slow. Sometimes setting patient expectations. Like, oh, I'm going to look that up. I'm going to do this for you. Maybe some type and sound. And then also doing things like you were saying. Being able to collect someone's email address. Like, you don't want to mess up their email address or phone number collection. Or, like, their insurance for their collection. So really naming the scene of people and understanding that they didn't clear something. And getting them to read them. And maybe another thing we did, actually, that I wanted to talk about a little bit was, so with our pre-op instructions, there's kind of a set of instructions that need to be delivered to patients. I mean, you have to be sure that they actually heard it. And it's hard with the castigating voice, like, agent stack that we use to know for sure that it wasn't interrupted mid-saying this thing. So, it's like, we know that L.M. generally hears these block attacks. It's going to say this. It's going to send it off. We're playing the audio. And then the patient will be interrupted midway. The agent, as we have it right now, doesn't listen to itself speak. So, it knows what it wants to say. It doesn't know what it actually thinks it's going to say. For us, like, at this point, we just need to restart the whole, like, instruction block if it's going to interrupt it. But as long as it's trying to break it down and not have a too big of an instruction set. Like, you don't want the AI to be talking to somebody because if someone interrupts it, it's going to have no idea where it flies. It either has to restart or it's going to skip forward. And here it is right here.

[00:19:15] Speaker 1: You can actually use, like, the context of the patient. The doctor you're going to, et cetera, is part of that voice engine experience. Or is it more, like, unscripted? Or, sorry, scripted in the sense where there's no context for each individual user. Like, you call them by name and mention which doctor you're going to.

[00:19:32] Speaker 3: Yeah, we do a little bit of that for IPOC ones. Like, we know what the procedure is. So, we can talk about that. You can take some questions about what the procedure is to have them. For, like, our team now, we don't even know what the agent is. So, we have no information on them at first. We don't, I mean, maybe there's a bit of context. But I think we don't save, like, team phone calls. We do have to validate the patient and call in if this is who they say they are. And if they're validated, then they're going to have to leave, like, a phone number and say, I have a file or something. It's kind of a bigger risk to be the thing that seems medical information than having to have to do the exam and not really be able to look at it.

[00:20:08] Speaker 1: I'm curious, maybe, if we can talk about it from the body side. And you probably have a lot that goes into, like, context and personalization. How have you set that up? Like, what does that look like under the virtual body?

[00:20:18] Speaker 2: Yeah. Within a single voice conversation, I mean, LLMs have huge context windows now. So, we can just handle all of that naturally. The difficulty is then when you go to multiple channels, talking to Gordy over email, text, expecting Gordy to remember all of that. And then some people have been talking to Gordy and having calls for, like, two years, right? So, it's like, how do you handle all of that? It's a bit of a nightmare. But, yeah, right now, we recently completely, like, revamped. Gordy's harness, you could say, over the text messaging channels. And so, now, that works very well in terms of, like, you know, a similar way of handling everything. To say, like, Drockbot or Coding Agents, where there's compaction, there's long-term memory, and there's all these different things. So, that works quite well. Our, quote, over voice, Gordy is still, like, you know, has room to improve. And so, actually, right now, in development, as I do to overhaul the ads with our voice, it seems like not too many people are doing this right now. But a similar system I mentioned where it is possible. Maybe we can get into it a bit more. But having those same things where there's compaction, there's, like, the ability to search previous conversations, all this stuff. The challenge with voice is that if you do that and they say a tool call, are you just leaving, you know, the person hanging? For, like, seconds while doing that. And so, for voice, there's a bit more complexity, and you have to be a bit smarter about it. Again, we can get into it more. But, say, with, like, Chai Chimichii Live, like, the new thing they released, which has, like, full duplex. And there's, like, a front-end and a back-end model so that it can continue talking while in the background, a bunch of things are going on. And so, I think that's, like, pretty interesting. And we're doing a full overhaul, as I mentioned. So, I'm excited for that, Sid.

[00:22:20] Speaker 1: That's awesome. I think we have one or two more questions. So, we'll do this. I'm curious, maybe, what surprised you most as you scaled your voice agent? It seems like it's really easy to build some demo as you put things up that look great on screen, and you show it in Meetup or whatever. What surprised you most about taking that, like, demo and your DSC all the way into production?

[00:22:43] Speaker 2: I mean, one thing that surprised me is that, like, we haven't run into, like, more issues. I think a lot of the providers we've been using, considering how, like, complex voice AI is and everything happening in the background, like, yeah, it's been good. But one challenge compared to, say, more, like, text-based agents is actually verifying and, like, essentially running evals or simulations on it. Because for the text-based agent, what we do at Gordy, we have something we call messaging agent simulation scenarios where we can set up an exact scenario, like, here's this person, and here's the situation they're in, and this is exactly what Gordy should do in this situation. And we have, like, 200 of these, so we know what Gordy should do in, like, every situation and all the, like, behaviors and everything like that. With voice, you can do a similar thing, but then when it's accurate, actually a person talks, and that ends up being different than, like, you know, simple text input. Or even, you know, in the case there's these services, which are really cool, which have other agents, which will then interact with your agent to, like, test different scenarios or edge cases and things like that. But even an agent interacting with your agent is, like, very different than humans interacting with your agent. So there's always things that go wrong which are really hard to catch ahead of time. So that's something that we're thinking about, especially as we develop agents into something more sophisticated, as I mentioned. Because right now, over voice, it's, like, not crazy complicated for Gordy. But once you add a lot more tools, like that kind of thing, there's going to be some fun stuff to figure out then.

[00:24:40] Speaker 3: Yeah, so for us, it was kind of just nuts and gates of doing outbound talk. But, you know, it's a lot harder than you expected. So, like, most time you call a patient, they're not going to answer. It's email to voicemail. So there's a number of different types of voicemail, meaning time and voicemail light. And there's all three of those, like, time to detect, voicemail is a perfect thing. We're trying to, like, listen for deep sounds. You see sound perfection. Trying to understand what you're at is, like, sound perfection and perfection. Like, the tokens in there would be, like, digital voicemail. But a lot of times it's just, like, text to voicemail, does it too early, starts trying to get the voicemail, and then making sure, you know, like, restart itself at the right time to get the voicemail. Really, it's probably over 50% of the time you can get the voicemail through a patient and say, call us back, and they'll call us back and then get on through the flow. But beyond just voicemails, that would be kind of new. When we launched, we had the phone, it called, like, everyone in our company, and went okay. So we went live with patients, and then immediately, on the first day, we ran into a lot of phone screeners. So these, you know, app-based videos, they ask, like, just explain who you are, why you're calling, and then all of a sudden, they're saying, let me get back to you, and this whole thing has to explain who's calling, and then we, and what person is in, and then they'll come in, they'll be the ones who might say, who's in, you know, due to flow. So I think just, like, this whole in action feels very not, like, it's not out of the box or that way, I guess. And I would love it, but you know, if it was not, you know, I guess, dealing with that sort of timing with all these other people, it's like,

[00:26:19] Speaker 4: so in our stack,

[00:26:20] Speaker 3: we use, like, Twilio, and Twilio actually initiates the outbound calls, and then all these games with a live build, through a sip-hunk, and then points to the in-house software and places it out to the patient, knowing, making sure that the live clip knows when the person actually picked up, so we can start doing this timer for, like, sound detection to re-keep the question. I guess, having the information back, it wasn't that hard, even though, like, we had to look at other possibilities, we had to hopefully find a room that might be able to,

[00:26:51] Speaker 2: and then, like,

[00:26:52] Speaker 3: update the metadata and have the room, and, you know, like,

[00:26:56] Speaker 2: this is all just, you know,

[00:26:57] Speaker 3: like, postage, and I keep holding up my phone to pick this up on a surprise call session.

[00:27:02] Speaker 1: Yeah. Can we do one more with the audience, I guess? Maybe, and usually I'd say whatever predictions, but we will do it the other way. Make your math and logic one. What do you want to see solved in voice AI over the 12 months, next 12 months, that would help you and all those builders? Maybe a grand conference,

[00:27:18] Speaker 3: I would love to see platforms like Lyricus and Alibaba. Honestly, you guys have built into it, like, playing with platforms that's new, I still see what to call them there. Right now, we are full transcription after the call was on our through assembly A,

[00:27:34] Speaker 2: so we're kind of

[00:27:35] Speaker 3: dumb or dumb to be on assembly A item. But, it's good to say that maybe we could do something about it at the end of August. So, maybe we should do something about it if I was just looking at it. It's not, I don't know, it feels like I'm still working on it, but, if I could get my phone to be able to listen to anything, like, if you can't use the online network, you can transcript what you have between

[00:28:07] Speaker 1: 40 steps

[00:28:08] Speaker 3: here

[00:28:09] Speaker 2: or 40 steps here. So, I think I'm just really excited about even some of the stuff I was just mentioning. Like, people now use Polycode and Podex, like, the capabilities are quite good and I think people are going to start expecting more of that in terms of, like, voice agents, the capabilities and what all they can do. And, right now, we're kind of inherently bottlenecked first. There's always going to be things that we dance around but, like, you know, isn't quite right. So, I'm looking forward to, like, the future as a speech-to-speech model is developing even further, getting more intelligent and everything. And then, also that, you know, front and back end architecture where, you know, you can have the agent every time you're here. So,

[00:29:12] Speaker 1: yeah, thanks so much all for coming out tonight so you're in the next hour so thank you.

ai AI Insights
Arow Summary
At a voice AI meetup, Owen from Sporty and Dylan from Flagler Health discussed how they are deploying voice agents in real-world settings. Sporty’s agent, Bordy, acts as a relationship-driven network matchmaker that calls, texts, and joins video meetings to learn users’ goals and facilitate relevant introductions. Flagler Health’s voice agent supports musculoskeletal clinics by making outbound patient calls for pre-procedure instructions, intake, and operational workflows.

The speakers contrasted open-ended, conversational agents with tightly structured healthcare flows. Sporty prioritizes charismatic, natural interaction and network-aware decision-making, while Flagler emphasizes guardrails, escalation, and keeping conversations strictly on task to avoid medical advice or mishandling sensitive information. Both highlighted immediate disclosure that callers are speaking with AI as essential to building trust.

Key production challenges included evaluation, transcript review, interruptions, voicemail detection, phone screeners, latency, extracting critical details such as emails and insurance information, and integrating with legacy healthcare systems. Sporty described the challenge of having an agent participate naturally in multi-person video meetings, using a classifier to determine when it should speak and pre-generating responses to achieve sub-second timing. Both teams use observability and human review to identify failure modes that automated success metrics can miss.

Looking ahead, the panel expects more capable speech-to-speech systems, better agent architectures that can keep talking while background actions run, and improved real-time transcription and voice-platform capabilities. The central lesson was that production voice AI requires far more than a convincing demo: it needs thoughtful trust design, domain-specific constraints, reliable integrations, rigorous evaluation, and handling of the messy realities of telephone and meeting conversations.
Arow Title
Building Production Voice Agents: Lessons from Sporty and Flagler Health
Arow Keywords
voice AI Remove
voice agents Remove
Sporty Remove
Bordy Remove
Flagler Health Remove
healthcare automation Remove
patient outreach Remove
agent evaluation Remove
trust and disclosure Remove
turn-taking Remove
video meetings Remove
voicemail detection Remove
LiveKit Remove
Twilio Remove
legacy integrations Remove
speech-to-speech models Remove
Arow Key Takeaways
  • Design voice agents around the domain: healthcare workflows need structured scripts, guardrails, and escalation, while relationship products can benefit from open-ended conversation.
  • Disclose that the caller is AI immediately; transparency avoids a later breach of trust and can reduce user anxiety.
  • Measure downstream outcomes, but pair them with observability and human transcript review because apparent task completion can hide failures.
  • Natural multi-party meeting participation requires accurate speak/not-speak classification and sub-second response timing.
  • Outbound calling introduces non-obvious production issues, including voicemail variants, phone screeners, spam labeling, and pickup timing.
  • Critical information capture, such as email addresses, phone numbers, insurance details, and pre-op instructions, needs robust confirmation and interruption handling.
  • Legacy system integrations can be as difficult as the conversation layer itself, particularly in healthcare.
  • Voice-agent testing remains difficult because text simulations and agent-to-agent tests do not fully represent real human conversations.
  • Future improvements will likely come from stronger speech-to-speech models, real-time transcription, and architectures that separate conversational responsiveness from background tool execution.
Arow Sentiments
Positive: The discussion is optimistic and practical, highlighting meaningful progress in voice AI while candidly addressing technical, operational, and trust-related challenges.
Arow Enter your query
{{ secondsToHumanTime(time) }}
Back
Forward
{{ Math.round(speed * 100) / 100 }}x
{{ secondsToHumanTime(duration) }}
close
New speaker
Add speaker
close
Edit speaker
Save changes
close
Share Transcript