Noise Suppression in Mel-Filtered Spectral Domain
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional noise suppression techniques for speech recognition are computationally complex, making them infeasible for resource-constrained applications, particularly when performing spectral analysis at high resolutions.
Innovation Solution
A system and method for noise suppression in the Mel-filtered spectral domain, involving a windowing module, conversion module, and Mel noise suppressor, which applies a window to a speech signal in the time domain, converts it to the frequency domain, and then to the Mel-filtered spectral domain for noise suppression, reducing computational complexity by filtering fewer channels than in the linear frequency domain.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional spectral enhancement techniques are used, then noise suppression performance is improved, but computational complexity increases
Solution Approach 1:
The patent segments the frequency spectrum into Mel-frequency bands, transforming the continuous spectrum into discrete, non-uniform frequency channels. This segmentation reduces the number of channels that need to be processed compared to linear frequency domain methods, thereby lowering computational complexity while maintaining noise suppression effectiveness.
Solution Approach 2:
The patent changes the frequency domain representation from linear frequency to Mel-frequency scale. This parameter transformation warps the frequency axis to better match human auditory perception, allowing for more efficient noise suppression with reduced computational requirements by focusing on perceptually relevant frequency regions.
2Measurement precision
If spectral analysis is performed at high resolutions, then measurement precision is improved, but computational complexity increases
Solution Approach 1:
The patent divides the spectrum into a manageable number of Mel-frequency bins that provide sufficient resolution for speech analysis. This segmentation achieves the necessary measurement precision for distinguishing speech from noise while avoiding the computational burden of ultra-fine frequency resolution.
Solution Approach 2:
The patent applies different processing characteristics to different frequency regions through the Mel-scale transformation. By concentrating more bins in lower frequencies where speech energy is typically concentrated and using fewer bins in higher frequencies, it achieves local optimization of resolution where it matters most while reducing overall computational complexity.
Data Source
AI summary
Techniques are described herein that suppress noise in a Mel-filtered spectral domain. For example, a window may be applied to a representation of a speech signal in a time domain. The windowed representation in the time domain may be converted to a subsequent representation of the speech signal in the Mel-filtered spectral domain. A noise suppression operation may be performed with respect to the subsequent representation to provide noise-suppressed Mel coefficients.


