top of page
Search

Microsoft Launches MAI-Transcribe-2 With 60-Language Speech Recognition at $0.10 Per Hour

  • Veronika
  • 4 hours ago
  • 5 min read

September 6, 2026

Microsoft AI has launched MAI-Transcribe-2, a new speech-recognition model built for fast, accurate transcription across 60 languages. The model combines speaker identification, word-level timestamps and configurable output styles with a limited-time launch price of $0.10 per hour of audio.

The release targets practical workloads including clinical documentation, legal records, accessibility, captions, call analysis and media production. Microsoft says the model improves both accuracy and processing speed, especially for long recordings and noisy real-world audio.

MAI-Transcribe-2: key features

  • Speech recognition across 60 languages.

  • Average word error rate of 5.2% on the multilingual FLEURS benchmark.

  • Speaker diarization that assigns words to individual speakers.

  • Word-level timestamps, keyword biasing and automatic language detection.

  • Limited-time launch pricing of $0.10 per hour through the end of 2026.

What is speaker diarization?

Speaker diarization identifies when different people are speaking and labels their portions of a recording. This is essential for meeting transcripts, interviews, medical conversations and customer-service calls, where a block of undifferentiated text is much less useful.

When combined with word-level timestamps, diarization makes audio searchable and easier to edit. A user can jump directly to a statement, create synchronized captions or trace who made a particular decision during a meeting.

Designed for multilingual and mixed-language audio

Microsoft says MAI-Transcribe-2 works across 60 languages and supports code switching, where speakers naturally move between languages such as English and Spanish or Hindi and English. Automatic language identification reduces the need to configure the model before processing each file.

The model ranked first in Microsoft’s evaluation of the FLEURS multilingual benchmark, with an average word error rate of 5.2%. Benchmark performance is useful for comparison, but customers should also test regional accents, specialist vocabulary and their own recording conditions.

Verbatim and clean transcription styles

Developers can choose between different transcription styles. A verbatim mode retains filler words and false starts, which may be important for compliance, research or detailed analysis. A clean mode removes verbal clutter to create more readable notes, captions and publishable transcripts.

Keyword biasing helps the system recognize company names, technical language, abbreviations and domain-specific terminology. This can be valuable in medicine, engineering and law, where a small transcription error may materially change meaning.

Speed, accuracy and price

Microsoft reports that MAI-Transcribe-2 can process audio up to 10 times faster than leading competitors in its comparisons. Based on Artificial Analysis testing cited by the company, it was 10 times faster than OpenAI’s GPT-Transcribe, seven times faster than ElevenLabs Scribe v2 and five times faster than Gemini 3.5 Transcribe while maintaining stronger accuracy.

The launch price is $0.10 per hour of audio through December 31, 2026. Developers should note that this is a limited-time rate and calculate future operating costs before making long-term commitments.

Business use cases

Healthcare organizations could use speech AI to draft clinical notes, subject to privacy controls and clinician review. Legal teams may search depositions and hearings, while media companies can create captions and searchable archives. Contact centers can analyze calls for customer needs, quality and compliance.

Accessibility is another major use case. Faster transcription can improve live captions and make audio or video content easier to use for people who are deaf or hard of hearing.

Privacy and accuracy considerations

Audio often contains personal, medical, financial or confidential information. Organizations should evaluate data retention, encryption, access controls, regional processing and contractual requirements before deployment.

Transcripts should also be treated as model output, not an unquestionable record. Names, numbers and technical terms deserve review, particularly when the result will support medical, legal or financial decisions.

Availability

MAI-Transcribe-2 is available to try through Microsoft Foundry, the MAI Playground and OpenRouter. Production teams should confirm service-region availability, throughput limits and final pricing for their deployment path.

The bottom line

MAI-Transcribe-2 shows how speech AI is becoming a specialized infrastructure service rather than a simple dictation feature. Its multilingual coverage, diarization, timestamps and low launch price could make large audio archives more useful. Real business value will depend on performance with local accents, specialist vocabulary and responsible handling of sensitive recordings.

Source: Microsoft AI’s MAI-Transcribe-2 announcementWhy word error rate is only one metric

Word error rate measures substitutions, deletions and insertions, but it does not capture every business requirement. Misidentifying a speaker, medication, legal term or number can be more damaging than several harmless errors. Teams should create weighted evaluations that reflect the consequences of mistakes in their domain.

Punctuation, formatting and latency also affect usability. A technically accurate transcript may still require significant editing if paragraphs, names or speaker turns are wrong.

Designing a production transcription pipeline

Audio should be validated before processing, with checks for duration, format and signal quality. Long recordings may need chunking, but boundaries must preserve speaker and sentence context. After transcription, automated rules can flag low-confidence names, dates and numbers for review.

Search indexes and summaries should link back to precise timestamps. That lets users verify a statement by listening to the original audio instead of trusting derived text.

Clinical and legal safeguards

Medical and legal recordings contain highly sensitive information. Organizations need encryption, strict role-based access, retention limits and contracts appropriate to the data. A transcription service should never receive more information than the task requires.

Human review remains essential when text enters a patient record, contract, court filing or compliance process. The model can accelerate documentation, but responsibility stays with qualified professionals.

Accessibility and live-caption design

Low latency can improve live captions, but speed should not come at the expense of stability. Rapidly changing text is difficult to follow. Applications need readable line breaks, appropriate delay and clear indication when a caption is provisional.

Multilingual events add complexity. Automatic language detection and code switching can reduce setup, while organizers should still test accents, names and specialist terminology before a high-profile broadcast.

Cost modeling beyond the launch price

At $0.10 per hour, processing a large archive may appear inexpensive. Total cost also includes storage, data transfer, indexing, review and downstream AI summaries. The announced rate is temporary, so procurement models should include a higher-price scenario after 2026.

Throughput limits matter as much as unit price. A newsroom or call center may need thousands of hours processed during a short window, which requires capacity planning and retry logic.

Comparing speech-recognition providers

A fair comparison uses the same audio set and scoring rules. Teams should include quiet studio speech, background noise, overlapping speakers, phone audio and regional accents. They should also test diarization, timestamps and vocabulary controls separately.

Portability is valuable because speech models improve quickly. Storing original audio, standardized transcripts and evaluation labels makes it easier to switch providers without rebuilding the entire workflow.

What to watch next

Independent testing will show whether Microsoft’s speed and accuracy claims hold across industries and languages. Developers should watch post-launch pricing, regional availability, streaming support and data-governance options. These factors will determine whether MAI-Transcribe-2 becomes a specialized tool or a broad speech infrastructure layer.

 
 
 

Recent Posts

See All

Comments


bottom of page