New Dictation API Delivers Clean, Fast Voice Input (Full Transcript)

A new API converts speech into polished, customizable text in 19 languages, with fast responses and context-aware recognition.
Download Transcript (DOCX)
Speakers
add Add new speaker

[00:00:00] Speaker 1: You've probably seen that little microphone icon at this point on every app you use. That's dictation showing up everywhere right now. And everyone building it hits the same wall which is transcription models are verbatim. They are trained to capture every sound that comes out of your mouth that's okay for a transcript but not okay for dictation. So that's why we built our first dictation API. You send audio you get back text that you'll actually use in 19 languages and shape the way your product needs it. Let's take a look. Um so so I think for the retry logic we should we should back off exponentially and cap it to like 30 seconds. Here's what I said and here's what my app gets. The dictation API gives you the original as well as the cleaned up results and you decide which one your product shows. And the number on the corner right here is not a marketing figure. It's actually how long it took to give you the response. Now let me take a few more takes on different domains on purpose. Can you make our one-on-one move it to Thursday afternoon? Anything after two words. Note for the ticket um the crash only happens after cold starts. So just fix it only for the iOS. Remind me to send Priya the updated pricing deck right before Friday. We need almond milk, coffee filter, and um good bread if they have it. There are two things. The first is speed. The whole point of dictation is to type 3x faster than a keyboard. Second is consistency. The user needs to have the same speed as the keyboard. Okay, so fix is in Lipsodium. Can you can you ping Anjali on the SDK team? A key part of dictation is capturing words accurately. So this runs on our flagship Universal 3.5 Pro model. The same one behind our async and real-time APIs. So let's take a look at the code. So the first thing I'm going to do is I'm going to show you how to do the dictation. The same one behind our async and real-time APIs. This also means the dictation API supports more than just English. And it holds up on how people actually speak. Remind me to ask about my offer letter right before my 4 o'clock. And that's because we want it to work no matter where you are. On a noisy street or in an office. Um, so 45 year old guy, lower back pain, um, going down the left leg. It's been about two weeks. We're gonna put him on a physio twice a week and see him in a month. Out of the box, it cleans up the mess. Out of the box, it cleans up the mess. But if you want something different, you can customize the instructions as well. So the instruction cleans up the output you receive. But prompting shapes how the model hears it. Dictation runs on Universal 3.5 Pro. So the prompt runs the same way like anywhere else on our API. You can pass context about the conversation or the key terms you expect. And the recognition improves before the cleanup ever runs. With the right prompt, you can see the right terms and entities captured correctly. So that's it. With one call, you can quickly build dictation to any product. Fast, customizable, and reliable. Our dictation API is available today. And we can't wait to see what you build with it.

ai AI Insights
Arow Summary
The speaker introduces a new dictation API designed to turn speech into usable text rather than a verbatim transcript. Built on the Universal 3.5 Pro model, it returns both original and cleaned-up output across 19 languages, with low-latency responses suited to keyboard-like dictation. Demonstrations cover scheduling, ticket notes, reminders, shopping lists, technical terms, multilingual usage, noisy environments, and medical notes. Developers can customize cleanup instructions and provide prompts with conversational context or expected terminology to improve recognition before cleanup. The API is available now through a single call.
Arow Title
New Dictation API for Fast, Clean Voice Input
Arow Keywords
dictation API Remove
speech-to-text Remove
Universal 3.5 Pro Remove
transcription Remove
voice input Remove
low latency Remove
19 languages Remove
custom instructions Remove
prompting Remove
entity recognition Remove
real-time API Remove
audio cleanup Remove
Arow Key Takeaways
  • The API is intended for dictation, producing polished usable text instead of purely verbatim transcripts.
  • It returns both the original transcription and a cleaned-up version, allowing product teams to choose what to display.
  • It is powered by the Universal 3.5 Pro model and supports 19 languages.
  • Speed is positioned as essential for dictation, with response latency demonstrated live.
  • Custom instructions control output cleanup, while prompts provide context and expected terms to improve recognition.
  • The API is designed to work across domains, accents, noisy conditions, and specialized vocabulary.
  • It is available now and can be integrated with a single API call.
Arow Sentiments
Positive: The tone is upbeat and product-focused, emphasizing speed, reliability, customization, multilingual support, and immediate availability.
Arow Enter your query
{{ secondsToHumanTime(time) }}
Back
Forward
{{ Math.round(speed * 100) / 100 }}x
{{ secondsToHumanTime(duration) }}
close
New speaker
Add speaker
close
Edit speaker
Save changes
close
Share Transcript