Noise Suppression Gain Optimization for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automatic speech recognition systems in noisy environments suffer from degraded accuracy due to noise suppression strategies that corrupt the speech signal, making them unusable.
Innovation Solution
The use of noise suppression information to optimize speech recognition by selecting a gain value for noise suppression based on features of the sub-band and time frame, providing noise suppression information to the speech recognition module, and adjusting resources such as bit rate based on the Signal-to-Noise Ratio (SNR), which includes voice activity detection and speech-to-noise ratio.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If strong attenuation of noise-dominated spectrum portions is applied, then noise reduction and perceived output signal quality are improved, but speech recognition accuracy deteriorates due to corruption of speech features
Solution Approach 1:
The patent applies different processing strategies to different spectral regions based on their content. Speech-dominated regions are preserved with minimal attenuation, while noise-dominated regions are strongly attenuated. This local differentiation resolves the contradiction by ensuring that noise reduction is applied only where necessary without corrupting speech features in speech-dominated regions.
Solution Approach 2:
The system dynamically adjusts attenuation parameters based on the estimated speech-to-noise ratio in each spectral region. By changing the attenuation parameter according to the local SNR condition, the system achieves strong noise reduction in noise-dominated regions while maintaining speech recognition accuracy in speech-dominated regions.
2Object-affected harmful factors
If noise suppression is applied to reduce noise corruption, then output signal quality is improved, but speech features are corrupted more than by the original noise
Solution Approach 1:
The system performs preliminary classification of spectral regions into speech-dominated and noise-dominated categories before applying attenuation. This preliminary action allows the system to identify which regions contain important speech features and protect them from strong attenuation, thereby preventing speech feature corruption while still reducing noise corruption in noise-dominated regions.
Solution Approach 2:
The system uses feedback from speech recognition performance to adjust noise suppression parameters. When speech features are corrupted, the system reduces attenuation in affected regions. This feedback mechanism ensures that noise suppression does not exceed the threshold that would corrupt speech features more than the original noise.
3Adaptability or versatility
If speech recognition is performed in noisy environments, then system versatility is improved, but recognition accuracy deteriorates significantly
Solution Approach 1:
The system dynamically adapts its noise suppression strategy based on the characteristics of the noisy environment and the specific speech signal. By continuously adjusting attenuation parameters in response to changing environmental conditions, the system maintains speech recognition accuracy across various noise environments, achieving both versatility and accuracy.
Data Source
AI summary
Noise suppression information is used to optimize or improve automatic speech recognition performed for a signal. Noise suppression can be performed on a noisy speech signal using a gain value. The gain to apply to the noisy speech signal is selected to optimize speech recognition analysis of the resulting signal. The gain may be selected based on one or more features for a current sub band and time frame, as well as one or more features for other sub bands and/or time frames. Noise suppression information can be provided to a speech recognition module to improve the robustness of the speech recognition analysis. Noise suppression information can also be used to encode and identify speech.


