Near-Field Speech Detection Using Spatial Statistics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice activity detectors fail to reliably distinguish desired speech from interfering sounds, especially when the microphone array is positioned off-axis or in the presence of diffuse-field sounds, leading to ineffective speech enhancement and noise reduction.
Innovation Solution
A telephone system with at least two microphones and a circuit that processes audio signals using statistics such as maximum normalized cross-correlation, inter-microphone level difference, and direction of arrival to enhance near-field speech detection, combining these statistics to improve reliability and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional voice activity detectors are used, then the system is simple, but they fail to reliably distinguish desired speech from interfering sounds
Solution Approach 1:
The patent combines multiple spatial statistics (inter-microphone level difference, normalized cross-correlation, direction of arrival, and diffuse-field gain) into a unified near-field detection system. This merging of multiple detection mechanisms enables reliable distinction between near-field speech and far-field interfering sounds, resolving the contradiction between detection reliability and system simplicity.
Solution Approach 2:
The system changes the detection parameters by using multiple spatial statistics instead of a single voice activity detection parameter. By computing and combining multiple spatial parameters (level difference, cross-correlation, direction of arrival), the system achieves reliable near-field speech detection in complex acoustic environments.
2Adaptability or versatility
If the microphone array is positioned off-axis, then the device is more adaptable to user positioning, but speech enhancement effectiveness deteriorates
Solution Approach 1:
The system dynamically adapts to different microphone array orientations by computing direction of arrival and using it to control speech enhancement algorithms. This dynamic adjustment allows the system to maintain effective speech enhancement regardless of whether the microphone array is positioned on-axis or off-axis relative to the user's mouth.
Solution Approach 2:
The system changes operational parameters based on detected sound field characteristics. By computing spatial statistics and using them to control the speech enhancement algorithm, the system adapts its processing parameters to maintain effectiveness across different microphone positioning scenarios.
3Object-affected harmful factors
If speech enhancement algorithms are applied, then noise reduction is improved, but interference with desired speech increases
Solution Approach 1:
The patent applies local quality by using near-field detection to selectively control speech enhancement only when near-field speech is present. The spatial statistics enable the system to identify the local acoustic environment (near-field vs. far-field) and adjust enhancement intensity accordingly, reducing harmful interference with desired speech while maintaining noise reduction effectiveness.
Solution Approach 2:
The system uses feedback from spatial statistics computation to control the speech enhancement algorithm. By continuously monitoring inter-microphone level difference, cross-correlation, and direction of arrival, the system provides feedback control that adjusts enhancement parameters to prevent over-enhancement and preserve desired speech quality.
4Measurement precision
If multiple microphones are used, then spatial statistics accuracy is improved, but device complexity increases
Solution Approach 1:
The patent segments the spatial statistics computation into distinct, manageable components: inter-microphone level difference calculation, normalized cross-correlation computation, direction of arrival estimation, and diffuse-field gain calculation. This segmentation of the measurement process enables accurate spatial statistics computation using multiple microphones while maintaining manageable system complexity through modular processing.
Data Source
AI summary
A telephone includes at least two microphones and a circuit for processing audio signals coupled to the microphones. The circuit processes the signals, in part, by providing at least one statistic representing maximum normalized cross-correlation of the signals from the microphones, doaEst, dirGain, or diffGain and comparing the at least one statistic with a threshold for that statistic. At least one of noise reduction and speech enhancement is controlled by an indication of near-field sounds in accordance with the comparison. Indication of near-field speech can be further enhanced by combining statistics, including a statistic representing inter-microphone level difference, each of which have their own threshold. dirGain and diffGain are derived from signals incident upon the microphones such that the desired near-field signal is not suppressed.


