Multi-Stage Speech Recognition Rescoring Temporal Posterior Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies face instability due to varying environmental noise, which affects recognition performance across different applications, and existing methods for converting feature vectors do not adequately address temporal features.
Innovation Solution
A multi-stage speech recognition apparatus and method that uses a first speech recognition unit to generate candidate words from input speech signals and a second unit to rescore these candidates using a temporal posterior feature vector, extracted through feature extractors like ASAT, TRAP, and STC-TRAP, which reflect time-varying voice characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional feature vector conversion methods (MFCC, LDA, PCA) are used, then the speech recognition system operates with standard processing complexity, but recognition performance becomes unstable in varying environmental noise conditions
Solution Approach 1:
The patent transforms the feature representation by computing posterior probabilities of phonemes given the observed feature vectors using HMM. This parameter transformation from raw feature vectors to posterior probability distributions enables the system to better handle environmental variations and noise, directly improving reliability under varying conditions
Solution Approach 2:
The patent introduces an intermediate processing stage that rescors candidate words using temporal posterior feature vectors derived from HMM observations. This intermediary rescoring mechanism acts as a mediator between initial recognition and final output, filtering out noise-affected candidates and improving overall recognition stability
2Reliability
If temporal feature conversion algorithms (RASTA, histogram normalization, delta features) are applied, then noise robustness is improved, but the processing complexity and computational load increase
Solution Approach 1:
The patent performs preliminary computation of posterior probabilities during the feature extraction and HMM observation phase. By pre-computing these temporal posterior features before the recognition decision, the system achieves noise robustness without adding complexity to the final recognition stage, as the heavy computational work is already done in the feature preparation phase
Solution Approach 2:
The HMM-based posterior probability computation automatically adapts to temporal patterns and noise characteristics in the speech signal without requiring manual feature engineering. The system uses the observed feature vectors to self-generate robust posterior probability features, eliminating the need for separate complex noise-robust feature conversion algorithms
3Measurement precision
If a multi-stage recognition system with rescoring is implemented, then recognition accuracy improves by 29.0% relatively, but the processing time and system complexity increase
Solution Approach 1:
The patent implements partial rescoring by focusing computational resources only on the top N candidate words from initial recognition rather than reprocessing all possible word sequences. This selective rescoring of high-probability candidates achieves significant accuracy improvement (29.0% relative) while minimizing additional processing time by avoiding exhaustive computation
Data Source
AI summary
Provided are a multi-stage speech recognition apparatus and method. The multi-stage speech recognition apparatus includes a first speech recognition unit performing initial speech recognition on a feature vector, which is extracted from an input speech signal, and generating a plurality of candidate words; and a second speech recognition unit rescoring the candidate words, which are provided by the first speech recognition unit, using a temporal posterior feature vector extracted from the speech signal.


