Microphone Selection for Speech Processing in Noisy Environments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Hands-free devices in noisy environments, such as motor vehicles, face challenges in distinguishing speech from ambient noise, especially with a single microphone, where beamforming techniques require multiple microphones and assume constant speaker positions, leading to suboptimal signal-to-noise ratios.
Innovation Solution
A method that automatically selects the microphone with the least noise by calculating a speech-presence confidence index for each channel and applying a decision rule based on this index, allowing for robust microphone selection in varying environments, even with two microphones spaced apart or close together.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If beamforming techniques are used to improve signal-to-noise ratio, then speech quality is improved, but the number of microphones required increases to at least five
Solution Approach 1:
The patent applies dynamic microphone selection by continuously evaluating speech-presence confidence indices from multiple microphones and switching between them based on current acoustic conditions. This dynamic approach replaces static beamforming configurations, allowing the system to adapt to varying speaker positions and noise environments without requiring a large fixed array of microphones.
Solution Approach 2:
The system changes operational parameters by calculating speech-presence confidence indices and using these to dynamically adjust which microphone is active. This parameter-based selection method allows effective speech processing with fewer microphones by optimizing the use of available sensors based on real-time acoustic analysis.
2Reliability
If a single unidirectional microphone is used to improve signal-to-noise ratio, then speech quality is improved, but the system can only handle one speaker position
Solution Approach 1:
The patent implements multi-functionality by using multiple microphones that can each serve as the primary pickup device depending on speaker position. The system universally handles different speaker locations (driver, passenger, rear seats) by selecting the appropriate microphone based on speech-presence confidence indices, making the system adaptable to various configurations without requiring separate dedicated microphones for each position.
Solution Approach 2:
The system dynamically switches between microphones based on real-time speech-presence confidence evaluation. This dynamic adaptation allows the system to maintain optimal signal-to-noise ratio regardless of which speaker position is active, transforming a static single-position system into a versatile multi-position solution.
3Adaptability or versatility
If multiple microphones are used with beamforming to handle varying speaker positions, then adaptability is improved, but device complexity increases
Solution Approach 1:
The patent extracts the essential function of speaker position detection from complex beamforming algorithms by using speech-presence confidence indices. This extraction simplifies the system by focusing on the key parameter (speech presence confidence) rather than implementing full beamforming complexity, reducing software burden while maintaining adaptability.
Solution Approach 2:
The system uses computationally lightweight speech-presence confidence calculations instead of heavy beamforming computations. This approach trades complex long-term processing for simpler, faster evaluations that can be performed continuously with minimal computational resources, reducing overall system complexity.
4Adaptability or versatility
If automatic microphone selection is implemented to improve adaptability, then speaker position flexibility is improved, but processing complexity increases
Solution Approach 1:
The patent applies partial action by calculating speech-presence confidence indices only for the purpose of microphone selection rather than performing complete acoustic analysis. This partial processing approach provides sufficient information for effective microphone switching without the excessive computational burden of full speech recognition or detailed acoustic modeling.
Data Source
AI summary
The method comprises the steps of: digitizing sound signals picked up simultaneously by two microphones (N, M); executing a short-term Fourier transform on the signals (xn(t), xm(t)) picked up on the two channels so as to produce a succession of frames in a series of frequency bands; applying an algorithm for calculating a speech-presence confidence index on each channel, in particular a probability a speech that is present; selecting one of the two microphones by applying a decision rule to the successive frames of each of the channels, which rule is a function both of a channel selection criterion and of a speech-presence confidence index; and implementing speech processing on the sound signal picked up by the one microphone that is selected.


