Microphone Array Source Location Estimation for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems in noisy environments struggle to accurately distinguish between intended and unintended voice signals, particularly in video game applications, due to unreliable voice volume estimation and sub-optimal performance of far-field microphones for close-talk scenarios.
Innovation Solution
A method using two or more microphones to estimate the distance and direction of a sound source by comparing volume and time delay properties of signals, allowing for reliable identification of the intended voice signal and rejection of background noise, employing a sound source discriminator with modules for voice segment detection, source location estimation, and decision-making to enable precise voice recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If far-field microphones are used for speech recognition, then the coverage area is increased, but the performance for close-talk scenarios deteriorates
Solution Approach 1:
The patent segments the microphone array into multiple groups, with some microphones optimized for far-field capture and others for close-talk scenarios. This segmentation allows the system to handle both distant and nearby speech sources effectively by processing signals from different microphone groups according to their respective optimization characteristics.
Solution Approach 2:
The patent introduces spatial dimensionality by using an array of microphones arranged in specific geometries (linear, planar, or volumetric configurations). By leveraging the spatial distribution of microphones and calculating arrival direction and distance in three-dimensional space, the system can distinguish between far-field and close-talk sources, resolving the contradiction between coverage area and close-talk performance.
2Device complexity
If voice volume is used for source distance estimation, then the estimation process is simplified, but the reliability of distance estimation deteriorates due to unknown real voice volume
Solution Approach 1:
The patent introduces an intermediary approach by using the array of microphones to measure the arrival direction and calculate distance based on signal propagation characteristics rather than relying solely on voice volume. This intermediary measurement method (using multiple microphones as mediators) provides more reliable distance estimation by accounting for the unknown real voice volume through comparative analysis of signals from multiple microphones.
Solution Approach 2:
The patent replaces the simple volume-based estimation mechanism with a more sophisticated acoustic field analysis mechanism. Instead of relying on the mechanical simplicity of volume measurement, the system uses the physical principles of sound wave propagation, arrival time differences, and intensity variations across multiple microphones to calculate distance, thereby improving measurement precision.
3Device complexity
If a single microphone is used for voice detection, then the device complexity is reduced, but the ability to distinguish between intended and unintended voice signals deteriorates in noisy environments
Solution Approach 1:
The patent segments the voice detection function across multiple microphones arranged in an array, with each microphone capturing a portion of the acoustic field. By segmenting the detection task and combining results from multiple microphones with different spatial positions, the system achieves reliable voice signal discrimination while maintaining manageable device complexity through modular processing.
Solution Approach 2:
The patent makes the microphone array universally applicable to multiple functions: far-field speech recognition, close-talk detection, voice signal discrimination, and spatial localization. The same array configuration and processing algorithms serve multiple purposes, allowing the system to distinguish between intended and unintended voice signals while maintaining cost-effectiveness and avoiding the need for separate specialized hardware for each function.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach effectively filters out unwanted voice signals, improving the accuracy of voice recognition in noisy environments and enhancing the reliability of voice-activated commands in applications like video games by accurately determining the source of speech, thereby reducing false triggers and improving user interaction.
Implementation Method 1
estimates a distance and direction to a speaker based on a comparison of a volume and a time delay of the signals from the two or more microphones
Data Source
AI summary
Computer implemented speech processing is disclosed. First and second voice segments are extracted from first and second microphone signals originating from first and second microphones. The first and second voice segments correspond to a voice sound originating from a common source. An estimated source location is generated based on a relative energy of the first and second voice segments and/or a correlation of the first and second voice segments. A determination whether the voice segment is desired or undesired may be made based on the estimated source location.


