Dual-Microphone Reverberation Mitigation Using Coherence Masking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Reverberant components in audio signals recorded by external microphones significantly hinder speech intelligibility in enclosed spaces, especially in environments with reflective surfaces, and existing methods either require prior knowledge of the clean signal or degrade overall speech quality.
Innovation Solution
A dual microphone signal processing system that converts time-domain signals to the time-frequency domain, applies a binary masking module to determine frequency-specific energy ratios and a sigmoid gain function based on inter-microphone coherence, reducing reverberation components through an inverse time-frequency transformation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single microphone is used with linear prediction analysis to estimate the binary mask, then the method requires no prior knowledge of the clean signal, but overall speech quality is degraded and large amounts of reverberant energy are not entirely suppressed
Solution Approach 1:
The patent divides the reverberation suppression task into two distinct stages: a binary masking stage that provides coarse suppression by comparing energy ratios, and a soft masking stage that provides fine-grained suppression using inter-microphone coherence. This segmentation allows each stage to specialize in different aspects of reverberation reduction, achieving both adaptability and effectiveness.
Solution Approach 2:
The patent combines the outputs of two separate masking approaches (binary masking and soft masking) into a unified reverberation suppression system. The binary mask provides robust suppression of dominant reverberation, while the soft mask refines the suppression by exploiting coherence information, together achieving superior speech quality and reverberation reduction compared to either method alone.
2Quantity of substance
If external microphones are placed at a distance from sound sources to capture audio signals, then the microphones detect both direct audio signals and reflected signals with time delay, but speech intelligibility is hindered due to smeared acoustic spectrum over time
Solution Approach 1:
The patent extracts the direct sound components from the mixed signal by identifying and retaining time-frequency units where the direct path energy dominates over reflected energy. By taking out only the relevant direct sound portions and suppressing the reflected components, the system preserves speech intelligibility while maintaining the benefits of distant microphone placement.
Solution Approach 2:
The patent transforms the signal processing from the time domain to the time-frequency domain, adding a frequency dimension to the analysis. This allows the system to separate direct and reflected sounds based on their different temporal and spectral characteristics, enabling effective reverberation suppression while preserving speech content.
3Loss of information
If the binary mask is generated by comparing energy of clean signal with reverberant signal in time-frequency domain, then speech intelligibility is improved, but the algorithm requires prior knowledge of the clean signal which cannot be guaranteed in realistic scenarios
Solution Approach 1:
The patent makes the system self-sufficient by deriving all necessary information from the reverberant microphone signals themselves. The binary mask is generated by comparing energy ratios within the reverberant signal, and the soft mask is generated by computing inter-microphone coherence, eliminating the need for external clean signal references or complex prior knowledge.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A dual microphone signal processing arrangement for reducing reverberation is described. Time domain microphone signals are developed from a pair of sensing microphones. These are converted to the time-frequency domain to produce complex value spectra signals. A binary gain function applies frequency-specific energy ratios between the spectra signals to produce transformed spectra signals. A sigmoid gain function based on an inter-microphone coherence value between the transformed spectra signals is applied to the transformed spectra signals to produce coherence adapted spectra signals. And an inverse time-frequency transformation is applied to the coherence adjusted spectra signals to produce time-domain reverberation-compensated microphone signals with reduced reverberation components.