Multi-Microphone Noise Suppression via Dynamic Masking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current noise suppression systems in audio processing fail to effectively reduce non-stationary noise and echo components while maintaining optimal speech quality, especially at low signal-to-noise ratios, as they either suppress noise conservatively to avoid distortion or fail to account for noise characteristics.
Innovation Solution
A robust noise suppression system that transforms acoustic signals into cochlea domain sub-band signals, subtracts noise and echo components, and applies a multiplicative mask to these signals, reconstructing them in the time domain to achieve flexible noise reduction while limiting speech distortion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If stationary noise suppression is applied by fixed or varying dB levels, then stationary noise is suppressed, but non-stationary noise is not suppressed and speech distortion occurs at low SNR
Solution Approach 1:
The system dynamically adapts noise suppression parameters based on real-time analysis of noise characteristics and speech presence. The noise suppression amount is adjusted frame-by-frame according to the detected noise type (stationary or non-stationary) and signal-to-noise ratio conditions, enabling effective suppression of both stationary and non-stationary noise while preserving speech quality.
Solution Approach 2:
The system changes suppression parameters (dB levels, filter characteristics) based on the detected noise characteristics. Different parameter sets are applied for stationary noise versus non-stationary noise conditions, and parameters are adjusted according to the estimated SNR to prevent speech distortion while maximizing noise suppression effectiveness.
2Object-affected harmful factors
If SNR-based dynamic noise suppression is applied, then overall noise level is reduced, but speech distortion occurs because SNR averaging masks different noise characteristics
Solution Approach 1:
The system segments the noise analysis into distinct components: stationary noise characteristics and non-stationary noise characteristics are analyzed separately rather than averaged together. This segmentation allows the system to apply appropriate suppression strategies for each noise type while preserving speech components, avoiding the speech distortion caused by uniform SNR-based suppression.
Solution Approach 2:
The system applies different noise suppression characteristics to different frequency regions and time frames based on local noise conditions. Rather than applying a uniform suppression level across the entire signal, the system adapts suppression parameters locally to match the specific noise characteristics present in each segment, thereby preserving speech quality while reducing noise.
3Manufacturing precision
If conservative noise suppression is applied to avoid speech distortion, then speech quality is maintained, but noise suppression effectiveness is reduced
Solution Approach 1:
The system employs feedback mechanisms where the output of noise suppression is monitored and used to adjust subsequent suppression parameters. Speech presence detection and distortion monitoring provide feedback that prevents excessive suppression, allowing the system to aggressively suppress noise when speech is absent or SNR is high, while automatically backing off when speech components are detected, thus achieving both effective noise suppression and speech quality preservation.
Data Source
AI summary
A robust noise reduction system may concurrently reduce noise and echo components in an acoustic signal while limiting the level of speech distortion. The system may receive acoustic signals from two or more microphones in a close-talk, hand-held or other configuration. The received acoustic signals are transformed to frequency domain sub-band signals and echo and noise components may be subtracted from the sub-band signals. Features in the acoustic sub-band signals are identified and used to generate a multiplicative mask. The multiplicative mask is applied to the noise subtracted sub-band signals and the sub-band signals are reconstructed in the time domain.


