Speech Recognition Using Bottleneck Features and Noise Elimination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The accuracy of obtaining hidden Markov model (HMM)-state level information in general Gaussian mixture model HMM (GMM-HMM) based speech recognition systems is not guaranteed, which restricts the performance of context-dependent deep neural network hidden Markov model (CD-DNN-HMM) algorithm based large vocabulary continuous speech recognition (LVCSR).
Innovation Solution
A speech recognition apparatus that extracts acoustic model-state level information using feature vectors based on gammatone filterbank signal analysis and bottleneck algorithms, and preprocesses training data to eliminate background noise, enhancing the accuracy of HMM-state level information extraction by classifying noise types and applying appropriate noise reduction techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If general GMM-HMM algorithm based speech recognition is used to obtain HMM-state level information, then the CD-DNN-HMM algorithm based LVCSR can be implemented, but the accuracy in obtaining HMM-state level information cannot be guaranteed
Solution Approach 1:
The patent changes the parameter representation from traditional GMM-HMM features to bottleneck features extracted through deep neural networks. This parameter transformation enables more accurate HMM-state level information extraction by learning optimal feature representations that capture essential speech characteristics while reducing dimensionality and noise.
Solution Approach 2:
The patent substitutes the traditional GMM-HMM algorithmic approach with a deep neural network based bottleneck feature extraction system. This replacement transitions from conventional statistical modeling to learned representations, achieving superior accuracy in obtaining HMM-state level information while maintaining system reliability.
2Adaptability or versatility
If training data containing background noise is used, then more diverse training data is available, but the accuracy of HMM-state level information extraction deteriorates
Solution Approach 1:
The patent extracts and removes background noise from training data through dedicated noise elimination processing. By separating and removing the harmful noise components while preserving the useful speech signals, the system maintains diversity in training data sources while ensuring high accuracy in HMM-state level information extraction.
Solution Approach 2:
The patent converts the presence of background noise in training data from a harmful factor into a beneficial preprocessing opportunity. By intentionally processing noisy data through noise elimination algorithms, the system transforms challenging real-world conditions into refined training samples that improve both model robustness and extraction accuracy.
Data Source
AI summary
Provided is an apparatus for large vocabulary continuous speech recognition (LVCSR) based on a context-dependent deep neural network hidden Markov model (CD-DNN-HMM) algorithm. The apparatus may include an extractor configured to extract acoustic model-state level information corresponding to an input speech signal from a training data model set using at least one of a first feature vector based on a gammatone filterbank signal analysis algorithm and a second feature vector based on a bottleneck algorithm, and a speech recognizer configured to provide a result of recognizing the input speech signal based on the extracted acoustic model-state level information.


