Acoustic Echo Detection Using Microphone Array Cross-Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Audio conferencing systems face challenges in distinguishing between a single talker with acoustic reflections and two talkers due to the barrel effect, leading to incorrect audio localization and camera steering issues, especially when the direct signal is weaker than reflections.
Innovation Solution
A method involving cross-correlation between pairs of average power signals from beamformers to differentiate between a single talker with reflections and multiple talkers, using normalized power signals to determine the maximum cross-correlation and lag, which can be implemented in real-time with reduced computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If microphone arrays and beamforming techniques are used to capture sound from a desired direction and attenuate sounds from other directions, then audio quality is improved by reducing reverberations, but the system cannot discriminate between a single talker with strong acoustic reflection and two different talkers
Solution Approach 1:
The patent segments the acoustic signal analysis by separating direct path signals from reflected signals using cross-correlation techniques. By analyzing the correlation between microphone signals and identifying time delays, the system segments the signal into direct components (from talker to array) and reflected components (from talker to wall to array), enabling accurate discrimination between single talker with reflection and multiple talkers.
Solution Approach 2:
The patent introduces cross-correlation analysis as an intermediary technique between the raw microphone signals and the final talker localization decision. This intermediary process computes correlation functions and analyzes time delays to determine whether strong signals from certain directions represent direct talker signals or reflected signals, thereby resolving the ambiguity in talker localization.
2Ease of operation
If the microphone array mistakenly interprets a reflection as a second talker, then camera steering is directed to the wrong location (wall, post, or column), but this mislocalization occurs when talker looks to another participant resulting in reflected audio signal stronger than direct path signal
Solution Approach 1:
The patent employs feedback mechanisms by continuously monitoring the cross-correlation results and time delay patterns from the microphone array signals. This feedback loop allows the system to dynamically adjust its interpretation of strong signals, using the correlation analysis to determine whether a strong signal from a particular direction represents a direct talker or a reflection, thereby preventing incorrect camera steering decisions.
3Measurement precision
If cross-correlation is performed on raw microphone signals to distinguish between single talker with reflection and multiple talkers, then discrimination accuracy is achieved, but computational complexity is very high
Solution Approach 1:
The patent applies preliminary action by performing beamforming and spatial filtering on the microphone signals before conducting cross-correlation analysis. By pre-processing the signals to enhance direct path components and suppress reflections through beamforming techniques, the system reduces the computational burden of the subsequent cross-correlation operation while maintaining accurate talker discrimination capability.
Solution Approach 2:
The patent implements partial action by focusing the cross-correlation analysis only on specific frequency ranges and time delay windows that are most relevant for distinguishing direct and reflected signals. Rather than performing full-spectrum, full-time cross-correlation on all microphone signals, the system selectively applies correlation analysis to pertinent signal portions, significantly reducing computational complexity while preserving discrimination accuracy.
Data Source
AI summary
A method is provided for discriminating between the case of a single talker with an acoustic reflection and the case of two talkers, regardless of their power levels. The method is implemented in real time by performing a cross-correlation between pairs of average power signals originating from pairs of beamformers. A detection decision is then made based on the value of the cross correlation and its lag.


