Meta Launches Muse Voice Transcribe With Support for Five Indian Languages

New Delhi — Meta has launched Muse Voice Transcribe, its first real-time audio perception model, with native support for five major Indian languages and a range of streaming transcription features.
Developed by Meta Superintelligence Labs, the model is designed to provide real-time speech-to-text transcription, speaker separation and multilingual code-switching without requiring a separate post-processing step, the company said.
Muse Voice Transcribe has been trained on more than 70 languages, with 25 languages validated at launch.
Meta said the model ranked first on the Artificial Analysis streaming speech-to-text leaderboard as of Sept. 1, 2026.
The model is available through the Meta Model API and is already being used for dictation in Meta AI for Mac and Muse Code.
Among its key features are automatic speech recognition, speaker diarization for more than 20 speakers in recordings lasting more than an hour, endpoint detection and support for seamless switching between languages.
Meta said transcription accuracy can also be improved by providing the model with language, keyword and contextual guidance.
The company said Muse Voice Transcribe uses an “adaptive delay” system that changes how long the model waits before producing each word. More difficult words can receive additional processing time, while easier words can be transcribed more quickly.
“The longer the model waits to predict, the more accurate the transcript, but the higher the latency,” Meta said.
The adaptive delay capability is powered by reinforcement learning, which balances word error rates against transcription delay.
Meta described Muse Voice Transcribe as an autoregressive multimodal model from its Muse Spark family.
Audio is processed in 80-millisecond segments, with each segment converted into a single soft token. At each step, the model determines whether to continue listening to additional audio or generate a text token.
Meta said the system is designed to improve the trade-off between transcription speed and accuracy, allowing the model to respond quickly while maintaining strong performance. (Source: IANS)



