Latest Technology News, Gadget Reviews & Tech Updates | Gadgets About

collapse
...
Home / AI Technology / Microsoft Launches New AI Models for Faster Real-Time Voice Conversations

Microsoft Launches New AI Models for Faster Real-Time Voice Conversations

Oct 02, 2026  Chaudhry Arslan  13 views

Microsoft has expanded its in-house AI portfolio with three new speech models designed to support faster and more natural real-time voice interactions.

The new models are MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Microsoft says they are designed to work together for low-latency conversational AI applications, including voice agents and other real-time services.

 

Microsoft Launches New AI Models for Faster Real-Time Voice Conversations

The biggest addition is MAI-Transcribe-2-Streaming, which is built for applications that need speech to be converted into text while a person is still talking.

Unlike Microsoft’s earlier MAI-Transcribe-2 model, which works with recorded audio, the new streaming version starts producing partial transcription results just over 100 milliseconds after receiving audio. It continues updating the text as additional context becomes available before producing the final transcript.

 

Microsoft says the model supports 60 languages and can automatically detect language changes continuously. The company also claims that it currently ranks first on Artificial Analysis for the accuracy of both partial and final transcripts.

In Microsoft’s internal testing for dictation and subtitles, words appeared in the transcript about twice as quickly as with the closest competing model.

 

MAI-Transcribe-2-Streaming is priced at $0.54 per hour of audio through the end of 2026. This is higher than the promotional $0.10-per-hour price of the non-streaming MAI-Transcribe-2.

New multilingual voice model

Microsoft has also introduced MAI-Voice-2.1, a new text-to-speech model designed for multilingual applications.

The model supports 23 languages across 26 locales. One of its features allows a single generated voice to switch between languages while keeping the same speaker identity. It can also use a native accent appropriate for each supported language.

Microsoft is pricing MAI-Voice-2.1 at $22 per one million characters.

Faster MAI-Voice-2.1-Flash

For applications where response speed is particularly important, Microsoft has launched MAI-Voice-2.1-Flash.

The company says the model can generate around 45 seconds of audio with approximately 150 milliseconds of end-to-end latency. Microsoft also claims that it provides 55% faster inference and costs about 60% less than comparable models.

MAI-Voice-2.1-Flash costs $15 per one million characters.

Both new voice models also support voice cloning in their supported languages using only a few seconds of reference audio.

Availability

Microsoft has made all three models available through Microsoft Foundry, MAI Playground and Vercel. The two voice models are also available through OpenRouter, while LiveKit support is expected to arrive soon.

The models can also be used through Azure Voice Live, giving developers several options for building real-time conversational applications.

With these launches, Microsoft is expanding its focus on speech AI alongside its existing work in areas such as reasoning, coding, image generation and voice technology.


Share:

Leave a comment

Your email address will not be published. Required fields are marked *

Your experience on this site will be improved by allowing cookies Cookie Policy