Author: Aadarsh Patel | EQMint
Humyn Labs has launched the second edition of BRIDGE, a global benchmark designed to evaluate the real-world performance of voice AI and automatic speech recognition (ASR) models across different languages and conversational conditions.
The Bengaluru-based Physical AI research lab evaluated 23 voice AI models across 23 languages, testing how the systems perform in noisy, real-world conversations involving accents, dialects, overlapping speech, pauses and conversational density. The models evaluated include Sarvam v3, Gemini 3 Pro and ElevenLabs, among others.
Voice AI Faces Challenges With Overlapping Speech and Dialects
According to the BRIDGE benchmark, overlapping speech increased average error rates from 41.2% to 45.2%. The study also identified significant differences in performance across regional dialects.
For instance, standard Bengali recorded a 42.4% error rate, while a regional Bengali dialect outside Kolkata recorded 51.0%. Similar differences were observed in Spanish, with Argentinian Spanish recording a 7.85% error rate compared with 16.04% for Venezuelan Spanish.
The benchmark also found substantial differences between models when processing identical audio. ElevenLabs recorded an average error rate of 5.8%, while GPT-4o-mini-transcribe recorded 24.6% in the tested dataset.
Different Models Show Different Types of Errors
BRIDGE found that voice AI errors vary depending on the model. Substitution errors, where the system produces the wrong word, were the dominant error type across 19 of the 23 models evaluated.
Omission errors, involving dropped words, were observed in models from OpenAI, Speechmatics and Gnani Vachana. The study also identified fabrication errors in Gemini Flash, where the model generated content that was not present in the original speech.
The benchmark also examined the impact of silence and pauses. In Brazilian Portuguese conversations, error rates reached 18.8% for pauses longer than 150 seconds, compared with 12.4% for shorter gaps.
Benchmark Built on 200+ Hours of Real-World Audio
The BRIDGE benchmark uses seven core evaluation metrics covering factors such as overlapping speech, conversational density, code-switching, pauses and dialect variation.
The dataset includes Indic languages, Latin American Spanish, Brazilian Portuguese and Vietnamese. Humyn Labs said the benchmark was built using more than 200 hours of human-verified audio, collected across two to three districts per language.
The methodology also distinguishes script mismatches from genuine transcription errors, providing additional insight into how voice AI systems could perform in commercial workflows.
Implications for Physical AI and Voice-Based Systems
Humyn Labs said reliable voice understanding is particularly important as AI moves beyond software environments and into physical systems such as robots.
Manish Agarwal, Co-Founder of Humyn Labs, said voice is a critical interface for Physical AI and that systems need to handle interruptions, overlapping speech, code-switching, pauses and linguistic diversity to support customer experience, automation and trust.
Ishank Gupta, Co-Founder of Humyn Labs, said sound provides information about people, actions, distance and intent, making speech evaluation an important component of Physical AI development.
The benchmark’s commercial analysis found that the best single model achieved a 10.7% loanword-adjusted error rate, while the best result at the language level reached 9.7%. The theoretical best result per call was 8.9%. One model won 78.6% of files outright, highlighting the potential importance of model selection and routing in voice AI deployments.
Humyn Labs describes BRIDGE as an independent ASR benchmark covering commercial and open-source models using field-collected, human-verified conversational audio. The company operates across India, Southeast Asia, Latin America and the Middle East, with a broader focus on data collection and enrichment across voice, vision, motion and touch for Physical AI applications.
Source: Startup Success Stories
Disclaimer: This article is for informational purposes only and does not constitute investment, financial, legal or professional advice. Readers are advised to independently verify the information and consult qualified professionals before making any investment or financial decisions.
For more such information, visit EQMint
Join our WhatsApp channel for timely updates: Whatsapp






