Speech Recognition Noise Suppression and Model Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face performance deterioration due to noise interference, with current noise suppression and acoustic model adaptation methods limited in handling various noise types effectively, resulting in reduced speech recognition accuracy.
Innovation Solution
A speech recognition device that combines noise suppression and acoustic model adaptation methods by storing suppression and adaptation coefficients, estimating noise, suppressing noise using these coefficients, and adapting acoustic models to improve recognition accuracy across different noise types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If noise suppression method is used to approximate input signal distribution to acoustic model, then speech recognition performance is improved, but performance deteriorates when noise type is not handled well
Solution Approach 1:
The patent combines noise suppression method and acoustic model adaptation method into a unified approach. The noise suppression unit processes the input signal while the acoustic model adaptation unit simultaneously adapts the acoustic model to the noise environment, merging these two separate methods to achieve both noise reduction and model adaptation benefits, thereby improving speech recognition performance across various noise types
Solution Approach 2:
The patent introduces dynamic adaptation where the acoustic model is continuously adjusted based on the estimated noise characteristics. The adaptation coefficient and suppression coefficient are dynamically determined based on the noise environment, allowing the system to adapt to changing noise conditions in real-time, enhancing both reliability and adaptability
2Reliability
If acoustic model adaptation method is used to approximate acoustic model to input signal distribution, then speech recognition performance is improved, but performance deteriorates when noise type is not handled well
Solution Approach 1:
The patent merges acoustic model adaptation with noise suppression by processing the input signal through both the noise suppression unit and the acoustic model adaptation unit. This combined approach ensures that the acoustic model is adapted to the specific noise environment while also suppressing noise components, thereby improving performance across diverse noise types
Solution Approach 2:
The patent changes the parameters of the acoustic model dynamically based on the estimated noise characteristics. The adaptation coefficient is determined based on noise estimates, allowing the model parameters to be adjusted to match the current noise environment, which improves the system's ability to handle various noise types effectively
3Object-affected harmful factors
If noise suppression coefficient is increased to suppress more noise, then noise reduction is improved, but speech recognition accuracy decreases due to over-suppression
Solution Approach 1:
The patent dynamically adjusts the suppression coefficient based on the estimated noise characteristics and signal properties. The coefficient is determined by analyzing the input signal and noise estimates, allowing the system to optimize the balance between noise suppression and speech preservation, preventing over-suppression while maximizing noise reduction
Solution Approach 2:
The patent implements feedback mechanisms where the noise estimation unit continuously monitors the input signal and provides feedback to the noise suppression unit. This feedback loop allows the suppression coefficient to be adjusted in real-time based on the actual noise conditions, ensuring optimal noise reduction without compromising speech recognition accuracy
Data Source
AI summary
The present invention can increase the types of noises that can be dealt with enough to enable speech recognition with a speech recognition rate of high accuracy.A speech recognition device of the present invention performs processes of: storing, in a manner to relate them to each other, a suppression coefficient representing a noise suppression amount and an adaptation coefficient representing an adaptation amount of a noise model, where the noise model is generated on the basis of a predetermined noise and is to be compounded (synthesized) to a clean acoustic model generated on the basis of a voice including no noise; estimating noise from an input signal; suppressing from the input signal a portion of the estimated noise of an amount specified by a suppression amount specified on the basis of the suppression coefficient; generating an adapted acoustic model which is noise-adapted, by compounding (synthesizing) the clean acoustic model with a noise model generated on the basis of the estimated noise in accordance with an adaptation amount specified on the basis of the adaptation coefficient; and recognizing voice on the basis of the noise-suppressed input signal and the generated adapted acoustic model.


