[00:00:00] Speaker 1: I will pass it over to you, Craig. Thanks.
[00:00:04] Speaker 2: Awesome, hi everyone, thanks so much, Natalie. Yeah, quick introduction on myself. I joined the team in June, so just four months ago. And this is probably the number one question I receive, is I'm building a X, which model should I use? And our product suite has been growing quick. We've been shipping quick new models. I feel like it's been like two launches a month now. So yeah, with how fast the space is moving, we all thought it would be great to host a session and kind of introduce when to pick which model. So that's the goal is that you leave here today, understanding which API fits what you want to build, whether it's real time, async, sync, dictation, how we compare against the market, the mistakes we see most, and all in between. So really the core theme is you want to pick the API based on when the audio exists and who is waiting for the text. So it really depends on the shape of what you want the latency to look like. Then you determine what the cost you're willing to pay for the different models. And so yeah, I'll just jump in right here. But first, I want to take a quick poll to see what folks are building, whether it's a voice agent, a live note taker, or captions, something on recorded audio, voice input, or you're still figuring it out. And if you are still figuring it out, no worries. This is exactly what today is for. So you should see the poll pop up there. I'll give it just another second. Cool. OK, so it looks like 29% are building a voice agent. Others, voice inputs, the most popular. We have some folks just still figuring it out, which is no worries at all. So I'm going to go ahead and take a look at that. It's no worries at all, so cool. Here's really the framework I'd like to look at, really based on asking these three fundamental questions. Question one, is someone waiting on the text right now? If the answer is no, then async is where you want to be. This means you don't really care about latency too much. You have a prerecorded audio file, and you want to turn that into a transcript. If the answer to that question is yes, someone is waiting on that text right now, then there's really three options you want to look at after that. If you need the transcript back live in real time, then you'll want to use our real-time model. If it's a voice agent, you'll want to use 3.6 Pro. I'll get into the specifics there a little bit later of 3.6 Pro versus our base model. This is actually going to launch shortly, I think at the end of October. We do have it live in a preview, but I'll get to that more in just a bit. But our real-time model, which essentially streams back, you open up a WebSocket and see that in real time. Then you look at, is the person still talking? If the answer is no, you already have that prerecorded snippet of whatever was just said, then you want to look at our sync or dictation APIs. So I'll get into those in just a bit here. Here are four real questions I received just this month. And I think these are important because it really determines the shape. It shows what the customer's looking for and helps guide them in picking the API. So the first one, should we transcribe live or after the call? If nobody needs to text during the call, async is cheaper, it sees the entire file and it gives higher accuracy as well. Someone has asked, what is our P99 latency? Building a voice agent, you need real-time responses really quick. So if that's a requirement, then that points you to our 3.6 Pro model. Other questions we've got, our old vendor speaker labels lag about five seconds. That's too slow for the live call. So this is like a note taker use case here. You want live speaker labels. And then the last one, this is a funny one. Someone asked, what is the difference between real-time and streaming? The answer is that's a trick question. It's the same thing there. So I'll jump in. This is the four models back to back at a glance. I'll start with async. I'll start with how you call each of them and we'll go like column by column here. Async, you submit a job and you get a webhook where you can pull. You could submit, or sorry, I'll stay with columns. In terms of sync, you send one request and you get the transcript back super, super quick. Real-time, you open up that web socket, like I said, and you stream words back, chunks at a time. And then dictation, you send one request and you get clean text back to you. This shows the amount of audio that each endpoint accepts. Async can process the largest files. Real-time stays open for three hours. That's the second longest one there. And then sync and dictation are meant for smaller chopped audio clips. In terms of latency, you could see sync is like blazing fast as well as real-time. Real-time works in chunks of turn-taking. So it's about 500 milliseconds to the end of the turn. And then async is about 10 to 32 seconds in latency to process a full file. But obviously that depends on what is the size of that file. And then languages, we're going to be launching a new async model pretty soon that will keep up with this 32. But for now, we're at 18 languages. If you use Universal 2 as a fallback, that will increase this up to 99 languages and all the rest are 32. And then you can see here the list price. It goes up, down the list here. So diving in a little bit deeper to each of these. Async, the audio is already recorded and nobody's waiting on it. So this is going to be your 21 cents an hour. Max file size is going to be 10 hours. And this is best for like call recordings, QA, coaching. The meeting ended, you already have the file. The file could be an MP3, an MP4, a WAV file or others. We accept a ton of different formats there. You could add speaker labels and we have other add-on features you could do on this. It's about two cents for speaker diarization. Basically you want to know who said what during a call. So it will label speaker A, speaker B, et cetera. And here's like a tiny code snippet. If you were to copy and paste this, you could get your first transcript running in just seconds here. So this is actually a live, you could call this and you'll get back a snippet there.
[00:07:51] Speaker 1: Hey Craig, we actually have two quick questions. The first one is how is dictation different from streaming with real time and cleaning up the text within LLM ourselves is the first one.
[00:08:07] Speaker 2: Cool. Yeah, I'm happy to touch on it in a little bit. I'll give you a sneak peek here in terms of dictation. Basically dictation is a way you hit one endpoint and in the response, you'll get the verbatim text as well as the cleaned up version. So I'll just jump a little bit out of order and we'll cover this slide and then get back to the previous there. So for example, if someone is dictating and they say, so can you send me the Q3 numbers before the Thursday meeting? No, wait, the Friday meeting, thanks. It will clean it up and send back both like the verbatim and the LLM response. So the difference is really just, do you want one request where we handle it for you or do you wanna do multiple requests? It's designed for short clips, just like sync. But yeah, hopefully that answers the question.
[00:09:06] Speaker 1: Perfect, and then, sorry, we can actually go back to that other slide. The next question is for a voice agent, is real-time the only choice or does sync make sense for short turn-by-turn exchanges?
[00:09:20] Speaker 2: Yeah, really good question. I will go ahead and jump into sync in real-time. So sync is designed for shorter audio files. So two minute is the maximum per request. So basically, if you have an audio file that's larger than two minutes, one thing you could do is you could chunk it down yourself. And there's one pitfall that I wanna make really clear that you'd like to avoid is you don't wanna chunk in like 30 second clippings or five second clippings because what happens, let's say you're a radiologist and you're examining an X-ray and you chop the audio right in the middle of the sentence, you're gonna get a higher word error rate and the down-funnel effects of that could be very damaging if you misdiagnose or misprescribe. So when it comes to chunking, you wanna ensure that you're chunking in utterances. So an utterance is a spoken phrase. So you want the chunk to be aware of when the sentence is starting and ending and switching between speakers. Basically, the answer to that is Realtime handles all of this for you. And you open up a WebSocket where it could be a three-hour long conversation and we detect all of the turns, we detect the utterances, we could label the speakers, all of that. If you wanna bring that yourselves, you're more than welcome to, and you could leverage a sync API and you'll get one HTTP request. And in the response, super, super quick, you'll get that transcript with no pulling, no WebSocket. So this absolves some of that complexity when it comes to creating a WebSocket, worrying about dropping packets if the internet's a little bit slow and things like that. But on the flip side, there's a little bit more complexity in tracking those turns. So yeah, I hope that answers the question there. So Realtime, I basically just covered that. We just launched our state-of-the-art model, Universal 3.6 Pro. And we are going to be launching a second model on this series that's going to be called Universal 3.6 Base at the end of the month. And the distinction I wanna draw here is that, when you're picking out which one to use, I think the question to ask is, does something talk back? Is it a voice agent? Is there like an automated phone line? Is it something that requires the turns, passing contacts, things like that? If you're building a voice agent, you'll wanna go with our state-of-the-art Universal 3 Pro. You can fine tune all of the parameters in terms of the turn detection settings. You can pass the contacts that the voice agent said back into the prompt for the next turn. So you'll get more accurate results and a lower word error rate. And things like conversation memory is related to that and voice focus to cut out a lot of the background noise. So if someone is talking to a voice agent to reschedule their flight, and they happen to be at their 13-year-old's birthday party and it's super, super loud, they won't hear all the background noise of the kids screaming and things like that. Or I love this example, if you're a drive-through voice agent and you have a kid in the backseat that yells like, I want the Frosty or like whatever, it won't pick up that background noise. It'll focus on the speaker who's talking and it will force the end of a turn based on whatever configuration is set. So you could set like a minimum delay, a maximum delay. And yeah, in terms of our base model, this is if you require, if you're building something that only listens to you, like a meeting note taker, you're building live captions or like call recording media, this is for you. This is a little bit simpler. It doesn't have all the fancy knobs, bells and whistles that our pro version has, but it comes at a cheaper cost and it's built for just pure transcription and speaker labels. So yeah, if you want the voice agent features, you're gonna be looking at the pro. If you're building a note taker, the base is gonna be for you. But awesome, we already covered dictation. So yeah, the only difference between this and sync is really the fact that you'll get the LLM generated cleanup in addition to that. So it's great for, I know in the medical space, we've had a ton of demand here, but yeah.
[00:14:31] Speaker 1: Craig, I got two questions.
[00:14:33] Speaker 2: Cool.
[00:14:35] Speaker 1: One, a quick one of our Arabic language is part of real-time models.
[00:14:41] Speaker 2: Yeah, I believe we just launched Arabic and I don't have the benchmarks off the top of my head, but I believe we just launched app for U36 Pro.
[00:14:51] Speaker 1: Amazing. And then we have another one of when is the base model that you were just talking about being released?
[00:14:58] Speaker 2: Yeah, good question. So it's already live right now in a preview state. So you could go to like the playground and you can select real-time model and you could build in production with it. It's all ready to go. Our official GA date, I'd have to confirm with the team, but I think it's late October.
[00:15:18] Speaker 1: Awesome, thanks.
[00:15:21] Speaker 2: Cool. So here's just a side-by-side of the four different ones here the four different endpoints, async, sync, real-time and dictation. And here's how long it will take to get a transcript back for each of them. So I ran these all just two days ago and I recorded, you know, three to seven runs on each. So async for a four minute clip with speaker labels on, it took 10 to 32 seconds to, or from the submit to the finish request. Sync took about one second round trip from a 27 second clip. Real-time first words after one second after connecting. The final text took about 0.3 to 0.4 seconds after the audio ended. So super quick there. And then dictation is about one second round trip. So pretty negligible, you know, the added latency for dictation is not really, you know, not too large there. So I've told you all about our different models, but I didn't tell you why assembly AI. So here are some benchmarks of how we compare to other models in the market. So really focused on async and real-time. And the reason for that is our sync and dictation models are built off of the same real-time model that we have trained. So looking at our async here, I have pulled our average word error rate against the price. So this is going to be 21 cents an hour. You could see we're in the bottom left here against a few others in the space. We came first of 11. In terms of missed entity rate, I think this one is a really important one that the headline numbers don't always focus on. We define a missed entity or an entity as a phone number, an email, a location, really anything that will determine down-funnel success for, you know, if there's going to be tool calling on top of the transcription. There's no point in building a voice agent and getting perfect word error rates on everything except for the email, because if you miss the email and you send it to the wrong address, then you won't land that customer and you won't have success. So we take a lot of time looking at our missed entity rate and how they compare to the rest of the market. This is specifically looking at medical just because it's a really big part of the market and it's one place we really thrive. So you could see the missed entity rate compared to others, other ASR models on Async. And again, we show up in like that bottom left box. In terms of our real-time model, this one is based on Universal 3.6 Pro. We came in first of 30 ASR models in terms of average word error rate. Coval, they run these independent every single day. They get fresh benchmarks. And yeah, you could see at the very top of that, there's Universal 3.5 Pro and then even better than that, 3.6 Pro. And then looking at semantic word error rate. So we define this as, if someone has a word doctor, right? And one ASR model spells it as DR, and the other spells it as like the full word D-O-C-T-O-R. They both have the semantic meaning correct, but depending on how you're normalizing it, one might count it as a word error. So when we look at the semantic word error rate, this was pulled from another independent benchmark ran in Pipecat. You could see we landed in the top three of 24. Then looking at missed entity rates on voice agent calls, we landed first of 22 at just about 14.41%. And these can be optimized further. Maybe that deserves its own session because there's a lot to dive in there. But these are benchmarking our models out of the box, but then there's a lot of parameters and things you could configure to work best with the shape of your data. So happy to dive into that on a later date. Here, I just want to highlight some of the differences we've seen in Universal 3.6 Pro against 3.5 on a real-time model. So really the three or what is it? Yeah, four double digit gains and fewer errors happen and the word error rates and heavy background noise. So that's going to be at your kid's birthday party or the kid in the backseat of the car yelling, their order to a voice agent. Code switching. So code switching is going back and forth between different languages. So if someone speaks a sentence half English, half Spanish, our models will be able to pick that up pretty accurately. And we've seen a lot of gains there where even mid-turn, it will be able to pick it up much better. Medical entity error has also improved quite a bit. It doesn't seem like a lot when you look at the numbers, but then you look at the Delta and you see, oh, that's a 12% difference. And then English short form word error rate. So if you are running a voice agent and someone says no, or like short utterances, it'll be able to pick that up much better. Beyond that, I just wanted to go over really quick, like some pitfalls to avoid. The number one that I see is a customer that says, that wants to use async when someone is waiting. And what a lot of folks do, cause they want to use the model with the cheapest price, they'll try to fit a round peg into a square hole and they'll chop it into five second clip utterances. And we'll have like brutal clippings, like I was saying before, if it cuts off mid sentences, that could be quite bad for the downfunnel success of what you want to do with those transcripts. So if you want a real-time flow, but you want to chop, I would say use a VAD, voice activity detector. Solero is an open source one that's pretty popular. And using that, you could chop your own and maintain your own chunking. And instead of sending that to async, you should be sending that to our sync API, which will allow you to get those requests back virtually in real-time, blazing fast latency. The second pitfall that I see a lot is streaming a recorded file through real-time. So real-time, you're built on how long that web socket is open. But on top of that, if you want to send like a pre-recorded file through real-time and it's a 30 minute clip, it'll take 30 minutes to complete. So you'd be much better off sending that through the async path. The third one is chopping long recordings to fit sync or dictation. Both take clips up to two minutes. So similar to what I was saying here, you want to chop it into the utterances rather than the brutal 30 second clippings. And then just three other common mistakes that I hope everyone here avoids. Assuming every feature is on every API. So speaker labels and medical mode, they're not yet on sync. Although I think we're making quick progress to launch that on there. So yeah, good to always check the docs. If you run cloud code or another agentic code editor, run it through there, make sure connected to our docs MCP to pick up all the nuances and differences. The next one is picking the model by price. Oftentimes when you do that, you'll be fitting those round pegs into a square hole and it's better to pick the one that works with the shape of what you're going for. And then you can optimize from there. And then the sixth one is leaving web sockets open. Actually, I see this one pretty common where your build on how long the connection is open, not on how much speech is spoken. So the pitfall here is if you leave every connection open, it'll time out after three hours and you'll be billed for the three hours there. So you could terminate when the call ends and you could set inactivity timeout as a backstop to make sure that doesn't happen to you. Real quick with the $20 credits, here's exactly what that buys at list price. It's 95 hours of async, 66 hours of sync, 44 hours of real time and 32 hours of dictation. So it's enough to really get started, play around with these APIs and learn like which one will work best for you. Yeah, so this afternoon, go ahead and claim your $20 in credits and you could call your first API in just seconds. We have some quick start links here. I could go ahead and send in the chat.
[00:25:06] Speaker 1: And then Craig, another quick question. So regarding the upcoming text-to-speech API, will it be real time and configurable for emotional tone and what's the expected general admission date if you know?
[00:25:22] Speaker 2: Sorry, can you repeat that?
[00:25:24] Speaker 1: Yes, regarding the text-to-speech API, will it be real time and configurable for emotional tone and what is the expected GA date?
[00:25:34] Speaker 2: Yeah, great question. I believe GA for that will be October 22nd. So yeah, that's the one API I did not cover. Actually that and our voice agent API, I think those each probably deserve their own session. But as far as text-to-speech, I am not sure, have to check with the engineering team what the shape of that request will look like. I don't believe it has that configurable in the parameters, but I'm not 100% sure.
[00:26:05] Speaker 1: Yes, amazing. We will also have a text-to-speech webinar preview. So definitely look forward to that on October 22nd. So we'll get information for that as well sent.
[00:26:21] Speaker 2: Sweet, awesome. Yeah, that's really all for me. Any other questions?
[00:26:30] Speaker 1: Let's see, I'll give it a quick minute if anyone has any lingering questions.
[00:26:36] Speaker 2: Cool, and I'll go ahead and copy these docs in here.
[00:26:59] Speaker 1: Okay, it looks like no other questions left, but feel free, we'll send a follow-up after this as well if you weren't able to access the link to get credits and then I'll make sure to send those resources that Craig has as well. Oh, there we go. Okay, quick question, is language attached to a people voice like Gradium for instance, for iStands?
[00:27:31] Speaker 2: I'm sorry, is it attached to, what was the question?
[00:27:34] Speaker 1: Is language attached to a people voice is the question.
[00:27:41] Speaker 2: I'm not sure I understand the question clearly, but yeah, language can be set. You could set one or multiple in a single API request. I'm not sure if this is related to text-to-speech, but I know for text-to-speech, we're gonna launch I think like close to 100 voices right off the gate and I know voice cloning as well as on our radar, so hopefully that answers the question.
[00:28:16] Speaker 1: Oh, for in Fradium for instance, an agent speaks French with Sarah and English with Mark.
[00:28:24] Speaker 2: Yeah, so we're launching different languages and different voices and each of them correlate to different language or accents. So a lot of that, we'll have more details to come.
[00:28:49] Speaker 1: Great, any other last minute questions feel free to put in the Q and A. Do you have an idea on the model speech-to-text to French, which do the streaming?
[00:29:11] Speaker 2: In terms of our real-time model, we do have French launched. I think that's like one of our main like tier one languages that we have. So you could go ahead and that's just a parameter on the API request. You could set your languages there or we also have something called ALD, automatic language detection. So you don't even need to configure that and the model will automatically pick up that you are speaking French.
[00:29:59] Speaker 1: Okay, I don't think I'm seeing any other questions. So happy to wrap it up here. Thanks so much, Craig.