On October 1, Microsoft introduced three audio models: MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. They are building blocks for developers of voice services. Despite similar names, recognizing speech and speaking a response are separate jobs.
Text arrives while a person is speaking
Microsoft lists 60 languages with automatic language detection for MAI-Transcribe-2-Streaming, at an introductory rate of $0.54 per hour of audio through the end of 2026.
The streaming transcription documentation distinguishes intermediate and final results. Intermediate text can change as more words arrive. Integration is available through the Realtime API or Azure Speech SDK. In Azure, this is a public preview without an SLA, which Microsoft does not recommend for production workloads.
The practical implication is that words already displayed on a screen may not be final. An order form, for example, should distinguish a provisional address from a confirmed transcript segment. That is our interface recommendation, rather than a finding from an accuracy test of this model.
Speech generation: two versions, a separate language list
The MAI-Voice-2.1 model page lists $22 per million characters for the standard model and $15 for Flash. Both support 23 languages. Flash targets responsive voice interactions, while the standard version emphasizes fidelity and longer narration. Billing is based on characters, unlike the audio-hour rate for transcription.
Check the speech generation documentation for supported languages and voices: recognizing 60 languages does not mean speaking all 60. Using a custom reference voice requires consent and gated access. The Azure voice models are also marked as public preview without an SLA.
Where to try them and what to consider
Microsoft lists Microsoft Foundry and MAI Playground as access points for all three models; the voice versions are also available through OpenRouter. An API announcement does not mean these capabilities have already appeared in every recorder or Windows application.
For someone who needs a finished interview or meeting transcript, the whole service matters: uploading recordings, editing, exporting and correcting mistakes. Our guide to choosing an AI transcription tool can help with those decisions. New models give developers more options, but do not replace the criteria for choosing a finished product.

Join the conversation
Stay on topic and respect other readers. Your first comment may appear after editorial review.