Speech Recognition Using Bottleneck Features and Noise Elimination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The accuracy of obtaining hidden Markov model (HMM)-state level information in general Gaussian mixture model HMM (GMM-HMM) based speech recognition systems is not guaranteed, which restricts the performance of context-dependent deep neural network hidden Markov model (CD-DNN-HMM) algorithm based large vocabulary continuous speech recognition (LVCSR).

Innovation Solution

A speech recognition apparatus that extracts acoustic model-state level information using feature vectors based on gammatone filterbank signal analysis and bottleneck algorithms, and preprocesses training data to eliminate background noise, enhancing the accuracy of HMM-state level information extraction by classifying noise types and applying appropriate noise reduction techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If general GMM-HMM algorithm based speech recognition is used to obtain HMM-state level information, then the CD-DNN-HMM algorithm based LVCSR can be implemented, but the accuracy in obtaining HMM-state level information cannot be guaranteed

Engineering Contradiction:
Improveaccuracy in obtaining HMM-state level informationVSAvoidperformance stability of CD-DNN-HMM algorithm based LVCSR
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent changes the parameter representation from traditional GMM-HMM features to bottleneck features extracted through deep neural networks. This parameter transformation enables more accurate HMM-state level information extraction by learning optimal feature representations that capture essential speech characteristics while reducing dimensionality and noise.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent substitutes the traditional GMM-HMM algorithmic approach with a deep neural network based bottleneck feature extraction system. This replacement transitions from conventional statistical modeling to learned representations, achieving superior accuracy in obtaining HMM-state level information while maintaining system reliability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If training data containing background noise is used, then more diverse training data is available, but the accuracy of HMM-state level information extraction deteriorates

Engineering Contradiction:
Improvediversity of training dataVSAvoidaccuracy of HMM-state level information extraction
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent extracts and removes background noise from training data through dedicated noise elimination processing. By separating and removing the harmful noise components while preserving the useful speech signals, the system maintains diversity in training data sources while ensuring high accuracy in HMM-state level information extraction.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent converts the presence of background noise in training data from a harmful factor into a beneficial preprocessing opportunity. By intentionally processing noisy data through noise elimination algorithms, the system transforms challenging real-world conditions into refined training samples that improve both model robustness and extraction accuracy.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS9805716B2Apparatus and method for large vocabulary continuous speech recognition
Publication Date: 2017.10.31 ELECTRONICS & TELECOMM RES INST
  • US9805716B2 patent drawing
  • US9805716B2 patent drawing
  • US9805716B2 patent drawing

AI summary

Provided is an apparatus for large vocabulary continuous speech recognition (LVCSR) based on a context-dependent deep neural network hidden Markov model (CD-DNN-HMM) algorithm. The apparatus may include an extractor configured to extract acoustic model-state level information corresponding to an input speech signal from a training data model set using at least one of a first feature vector based on a gammatone filterbank signal analysis algorithm and a second feature vector based on a bottleneck algorithm, and a speech recognizer configured to provide a result of recognizing the input speech signal based on the extracted acoustic model-state level information.