Multi-microphone Speech Recognition via Concurrent Utterance Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face challenges in far-field environments due to additive noise and reverberation, leading to high word error rates and unnatural user interactions, especially when users are distant from microphones, as they require users to stand near a single receiver and speak directly into it.
Innovation Solution
The development of acoustically distributed yet functionally centralized speech recognition systems that use multiple microphones to concurrently process and combine different versions of an utterance, employing adaptive weighting and fusion techniques to generate a highest-probability transcription, allowing for natural interaction with devices regardless of user position or orientation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a single microphone receiver is used for speech recognition, then the device complexity is reduced, but the speech recognition reliability deteriorates in far-field environments due to additive noise and reverberation
Solution Approach 1:
The speech recognition system is segmented into multiple independent microphone receivers distributed throughout the room, each capturing the utterance independently. This segmentation allows the system to process multiple versions of the same utterance from different spatial positions, improving reliability in far-field environments by reducing the impact of additive noise and reverberation at any single location.
Solution Approach 2:
Multiple versions of the utterance captured by different microphones are merged and processed concurrently by a speech recognition engine. The system combines the information from multiple independent streams to determine the most likely transcription, thereby improving speech recognition reliability through ensemble processing while managing system complexity through centralized recognition logic.
2Ease of operation
If multiple microphones are distributed throughout the room, then the ease of operation improves by allowing natural user movement, but the device complexity increases due to multiple receivers and processing requirements
Solution Approach 1:
Multiple microphones are distributed throughout the room to create a universal speech recognition system that can handle user interactions from any position or orientation. This multi-functional arrangement allows users to naturally move and speak without needing to position themselves specifically, as any microphone can capture the utterance effectively.
Solution Approach 2:
A centralized speech recognition engine acts as an intermediary that receives and processes versions of utterances from multiple distributed microphones. This intermediary consolidates the complexity of handling multiple independent receivers, managing the concurrent processing and selection of the most likely transcription, thereby reducing the operational complexity exposed to the user.
3Measurement precision
If conventional speech recognition systems are used in far-field environments, then the device complexity remains low, but the measurement precision deteriorates with word error rates increasing to 60-80%
Solution Approach 1:
The speech recognition system dynamically processes multiple concurrent versions of the utterance from different microphones, adapting to the varying quality and characteristics of each stream. The recognition engine evaluates multiple transcription candidates and dynamically selects the most likely transcription based on the combined information, thereby improving measurement precision in far-field environments.
Solution Approach 2:
The system changes the parameter of processing by concurrently analyzing multiple independent streams of the same utterance rather than processing a single stream. This parameter change allows the system to leverage the diverse information from multiple microphones, improving transcription accuracy by reducing the impact of noise and reverberation that affect individual streams.
Data Source
AI summary
A speech recognition system for resolving impaired utterances can have a speech recognition engine configured to receive a plurality of representations of an utterance and concurrently to determine a plurality of highest-likelihood transcription candidates corresponding to each respective representation of the utterance. The recognition system can also have a selector configured to determine a most-likely accurate transcription from among the transcription candidates. As but one example, the plurality of representations of the utterance can be acquired by a microphone array, and beamforming techniques can generate independent streams of the utterance across various look directions using output from the microphone array.


