Multi-microphone Speech Recognition with Beamforming and Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face challenges in far-field environments due to additive noise and reverberation, leading to high word error rates (WER) and unnatural user interactions, as they require users to stand near a single microphone, limiting movement and causing errors in device control.
Innovation Solution
The implementation of acoustically distributed yet functionally centralized speech recognition systems that use multiple microphones to concurrently process and combine different versions of an utterance, employing adaptive weighting and deep-neural-networks to improve transcription accuracy and reduce WER, allowing users to move freely while maintaining accurate device control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If multiple microphones are used to capture far-field speech, then the coverage area and user mobility are improved, but the speech recognition accuracy deteriorates due to additive noise and reverberation
Solution Approach 1:
The patent segments the speech signal processing by creating multiple independent recognition streams from different microphone recordings. Each stream processes the speech signal separately through its own acoustic feature extraction and recognition engine, allowing the system to handle far-field conditions more effectively by dividing the processing task across multiple parallel pathways rather than relying on a single processed signal.
Solution Approach 2:
The patent merges the results from multiple independent recognition streams through a selection mechanism. The selector component combines the transcription candidates from different streams and selects the best transcription, effectively merging the strengths of multiple microphones and processing paths to overcome the weaknesses of individual far-field recordings.
2Reliability
If users speak towards a single near-field receiver, then the speech recognition accuracy is improved, but the user mobility and natural interaction are restricted
Solution Approach 1:
The patent creates a universal speech recognition system that functions effectively in both near-field and far-field conditions. By deploying multiple microphones throughout the environment and processing each recording independently, the system becomes universally responsive to speech from various positions and orientations, eliminating the need for users to adopt specific speaking postures or positions.
Solution Approach 2:
The patent transitions from a single-point (0D) or line (1D) recognition approach to a distributed spatial (3D) arrangement of microphones. This dimensional expansion allows the system to capture speech from any location in the room, converting the limitation of single-position recognition into a multi-position spatial recognition capability.
3Adaptability or versatility
If independent devices operate autonomously with their own speech recognition, then the device autonomy is improved, but the misinterpretation of user intent increases
Solution Approach 1:
The patent introduces a central hub as an intermediary that coordinates speech recognition across multiple devices. The hub receives recordings from various devices, manages the independent recognition streams, and selects the appropriate transcription and target device. This intermediary layer maintains device autonomy while preventing misinterpretation by centralizing the decision-making process for intent recognition.
Data Source
AI summary
A speech recognition system for resolving impaired utterances can have a speech recognition engine configured to receive a plurality of representations of an utterance and concurrently to determine a plurality of highest-likelihood transcription candidates corresponding to each respective representation of the utterance. The recognition system can also have a selector configured to determine a most-likely accurate transcription from among the transcription candidates. As but one example, the plurality of representations of the utterance can be acquired by a microphone array, and beamforming techniques can generate independent streams of the utterance across various look directions using output from the microphone array.


