Speech Noise Suppression Using Spectral Correlation for Babble Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing noise reduction systems struggle to effectively suppress 'babble noise' in speech due to its speech-like nature, particularly in far-field microphone applications where speakers can be at varying distances, leading to inefficiencies in voice activity detection.
Innovation Solution
A noise suppression method that transforms the input signal into a spectrum, applies smoothing and spectral correlation factors to estimate suppression filter coefficients, and filters the signal to enhance speech detection, using dynamic and adaptive scaling to account for varying noise environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional voice activity detectors are used to detect speech in noisy environments, then the system is simple to implement, but the detection accuracy deteriorates when babble noise is present
Solution Approach 1:
The patent segments the noise spectrum into multiple frequency bands and applies different suppression strategies to each band. The spectral correlation factor computation divides the frequency spectrum into bands, allowing selective suppression of babble noise while preserving speech components in different frequency regions
Solution Approach 2:
The patent employs dynamic adaptation by iteratively computing spectral correlation factors and updating suppression coefficients based on changing noise conditions. The system dynamically adjusts the scaling factor and suppression parameters to track non-stationary babble noise characteristics in real-time
2Object-affected harmful factors
If noise suppression filtering is applied to remove babble noise, then the noise level decreases, but speech quality may be degraded due to over-suppression
Solution Approach 1:
The patent applies local quality by computing spectral correlation factors for different frequency bands independently and applying targeted suppression only where babble noise is detected. The suppression filter coefficients are adjusted locally in the frequency domain to preserve speech components while removing noise in specific spectral regions
Solution Approach 2:
The patent implements feedback through iterative computation where the output spectrum from one iteration becomes the input for the next iteration. The spectral correlation factors are continuously updated based on the filtered output, allowing the system to adapt and refine suppression while monitoring speech quality preservation
3Adaptability or versatility
If the system adapts to non-stationary noise environments through iterative scaling, then the adaptability improves, but the computational complexity increases
Solution Approach 1:
The patent applies partial action by computing spectral correlation factors for selected frequency bands rather than processing the entire spectrum uniformly. The iterative scaling focuses computational effort on bands where babble noise is most prominent, reducing overall computational complexity while maintaining adaptability to non-stationary environments
Data Source
AI summary
A noise suppression method includes transforming a time-domain input signal into an input spectrum that is the spectrum of the input signal, the input signal comprising speech components and noise components, and the input spectrum comprising a speech spectrum that is the spectrum of the speech components and a noise spectrum that is the spectrum of the noise components, smoothing magnitudes of the input spectrum to provide a smoothed-magnitude input spectrum, and estimating basic suppression filter coefficients from the input spectrum and the smoothed input spectrum. The method further includes determining noise suppression filter coefficients from the estimated basic suppression filter coefficients and a spectral correlation factor, the spectral correlation factor indicating whether speech is present in the input signal or not, filtering the input spectrum based on the noise suppression filter coefficients to generate an output spectrum; and transforming the output spectrum into a time-domain output signal.


