RNNLM Training Using Minimum Word Error Criterion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for training recurrent neural network language models (RNNLMs) do not effectively minimize the word error rate (WER) in speech recognition, as they consider only one competitor hypothesis and do not account for inter-dependence of word errors, leading to insufficient discriminative training.
Innovation Solution
The proposed method employs Minimum Word Error (MWE) training for RNNLMs, using N-best lists and back-propagation through time (BPTT) on graphics processing units (GPUs) to minimize the expected word error rate, allowing for parallelization of computations and consideration of multiple competitor hypotheses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If cross entropy training is used for RNNLMs, then the model can be trained efficiently with maximum likelihood criterion, but the word error rate is not minimized explicitly
Solution Approach 1:
The patent changes the training objective parameter from cross entropy (maximum likelihood) to minimum word error rate criterion. This parameter change allows the model to be trained directly to minimize WER, resolving the contradiction between training efficiency and word error rate minimization by providing a discriminative training target that directly optimizes the performance metric.
Solution Approach 2:
The patent implements feedback by using the actual word error rate computed from N-best hypotheses as the training signal. The gradient of the WER criterion is computed and fed back to update the model parameters, creating a closed-loop system where the performance metric directly guides the optimization process, thereby minimizing WER explicitly while maintaining training feasibility.
2Device complexity
If only one competitor hypothesis is considered in training, then the training process is simpler, but the discriminative performance is insufficient
Solution Approach 1:
The patent applies partial action by considering a limited number of top-N competitor hypotheses from the N-best list rather than all possible hypotheses. This partial consideration strikes a balance between training complexity and discriminative performance, providing sufficient competitive signals to improve the model without overwhelming computational burden.
Solution Approach 2:
The training data is segmented into multiple competitor hypotheses drawn from the N-best list. Instead of using a single hypothesis, the patent segments the hypothesis space into multiple competing alternatives, allowing the model to learn discriminative boundaries between several plausible options, thereby improving discriminative performance while managing complexity through controlled segmentation.
3Ease of operation
If word-to-word alignment is fixed during training, then the training process is more straightforward, but inter-dependence of word errors is ignored
Solution Approach 1:
The patent introduces dynamics into the alignment process by allowing word-to-word alignment to vary during training based on the competitor hypotheses. Instead of fixing alignment, the model dynamically determines the most likely alignment between reference and hypothesis words for each training example, enabling it to capture inter-dependences among word errors while maintaining training feasibility through probabilistic alignment estimation.
4Measurement precision
If N-best lists are used for MWE training with BPTT on GPUs, then the expected word error rate is minimized, but the computation requires parallelization to be practical
Solution Approach 1:
The patent substitutes mechanical sequential computation with parallel computation on GPUs. By formulating the MWE training with BPTT in a way that can be parallelized across multiple hypotheses and time steps, the patent replaces the sequential mechanical processing bottleneck with concurrent computational operations, enabling practical implementation of the computationally intensive WER minimization approach.
Data Source
AI summary
A method and for training a language model to reduce recognition errors, wherein the language model is a recurrent neural network language model (RNNLM) by first acquiring training samples. An automatic speech recognition system (ASR) is applied to the training samples to produce recognized words and probabilites of the recognized words, and an N-best list is selected from the recognized words based on the probabilities. determining word errors using reference data for hypotheses in the N-best list. The hypotheses are rescored using the RNNLM. Then, we determine gradients for the hypotheses using the word errors and gradients for words in the hypotheses. Lastly, parameters of the RNNLM are updated using a sum of the gradients.


