Maximum Likelihood Channel Normalization for ASR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automated speech recognition systems face challenges in achieving accurate transcriptions due to the resource-intensive nature of constrained maximum likelihood linear regression (cMLLR) and the need for preliminary transcripts, which limits its application in channel normalization.
Innovation Solution
Applying maximum likelihood methods to calculate offset and scaling factors for channel normalization, allowing feature vectors to be transformed to maximize speech recognition criteria without requiring a preliminary transcript, and using Gaussian mixture models to update statistics for improved alignment and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If constrained maximum likelihood linear regression (cMLLR) is used for channel normalization, then speech recognition accuracy is improved, but computational resources and processing time increase significantly
Solution Approach 1:
The patent changes the parameters of channel normalization by using maximum likelihood estimation to calculate offset and scaling factors instead of traditional fixed parameters. This allows the system to adapt normalization parameters to maximize speech recognition likelihood while reducing computational complexity compared to full cMLLR
Solution Approach 2:
The patent substitutes the complex mechanical system of full cMLLR with a simplified statistical approach using maximum likelihood estimation. This replacement maintains accuracy benefits while significantly reducing computational resource requirements through more efficient parameter estimation
2Measurement precision
If constrained maximum likelihood linear regression (cMLLR) is used for channel normalization, then speech recognition accuracy is improved, but processing time increases
Solution Approach 1:
The patent changes the normalization parameters using maximum likelihood estimation, which allows for faster computation compared to full cMLLR. The offset and scaling factors are estimated directly from the data distribution without requiring iterative optimization, significantly reducing processing time while maintaining accuracy
3Use of energy by moving object
If traditional channel normalization is used, then computational resources are reduced, but speech recognition accuracy deteriorates
Solution Approach 1:
The patent changes traditional channel normalization parameters to be data-driven through maximum likelihood estimation. This allows the system to maintain computational efficiency while improving accuracy by adapting normalization to the specific characteristics of the speech data distribution
4Measurement precision
If constrained maximum likelihood linear regression (cMLLR) is used, then speech recognition accuracy is improved, but the need for preliminary transcripts increases device complexity
Solution Approach 1:
The patent extracts the essential function of cMLLR (maximizing likelihood through parameter optimization) while removing the dependency on preliminary transcripts. This separation allows the system to achieve accuracy improvements without the complexity of multi-pass processing with transcript alignment
Data Source
AI summary
Features are disclosed for applying maximum likelihood methods to channel normalization in automatic speech recognition (“ASR”). Feature vectors computed from an audio input of a user utterance can be compared to a Gaussian mixture model. The Gaussian that corresponds to each feature vector can be determined, and statistics (e.g., constrained maximum likelihood linear regression statistics) can then be accumulated for each feature vector. Using these statistics, or some subset thereof, offsets and/or a diagonal transform matrix can be computed for each feature vector. The offsets and/or diagonal transform matrix can be applied to the corresponding feature vector to generate a feature vector normalized based on maximum likelihood methods. The ASR process can then proceed using the transformed feature vectors.


