RNNLM Training Using Minimum Word Error Criterion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for training recurrent neural network language models (RNNLMs) do not effectively minimize the word error rate (WER) in speech recognition, as they consider only one competitor hypothesis and do not account for inter-dependence of word errors, leading to insufficient discriminative training.

Innovation Solution

The proposed method employs Minimum Word Error (MWE) training for RNNLMs, using N-best lists and back-propagation through time (BPTT) on graphics processing units (GPUs) to minimize the expected word error rate, allowing for parallelization of computations and consideration of multiple competitor hypotheses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If cross entropy training is used for RNNLMs, then the model can be trained efficiently with maximum likelihood criterion, but the word error rate is not minimized explicitly

Engineering Contradiction:
Improvetraining efficiencyVSAvoidword error rate
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the training objective parameter from cross entropy (maximum likelihood) to minimum word error rate criterion. This parameter change allows the model to be trained directly to minimize WER, resolving the contradiction between training efficiency and word error rate minimization by providing a discriminative training target that directly optimizes the performance metric.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback by using the actual word error rate computed from N-best hypotheses as the training signal. The gradient of the WER criterion is computed and fed back to update the model parameters, creating a closed-loop system where the performance metric directly guides the optimization process, thereby minimizing WER explicitly while maintaining training feasibility.

Inventive Principle:
Principle #23Feedback

2Device complexity

If only one competitor hypothesis is considered in training, then the training process is simpler, but the discriminative performance is insufficient

Engineering Contradiction:
Improvetraining complexityVSAvoiddiscriminative performance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent applies partial action by considering a limited number of top-N competitor hypotheses from the N-best list rather than all possible hypotheses. This partial consideration strikes a balance between training complexity and discriminative performance, providing sufficient competitive signals to improve the model without overwhelming computational burden.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The training data is segmented into multiple competitor hypotheses drawn from the N-best list. Instead of using a single hypothesis, the patent segments the hypothesis space into multiple competing alternatives, allowing the model to learn discriminative boundaries between several plausible options, thereby improving discriminative performance while managing complexity through controlled segmentation.

Inventive Principle:
Principle #1Segmentation

3Ease of operation

If word-to-word alignment is fixed during training, then the training process is more straightforward, but inter-dependence of word errors is ignored

Engineering Contradiction:
Improvetraining straightforwardnessVSAvoidinter-dependence of word errors
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent introduces dynamics into the alignment process by allowing word-to-word alignment to vary during training based on the competitor hypotheses. Instead of fixing alignment, the model dynamically determines the most likely alignment between reference and hypothesis words for each training example, enabling it to capture inter-dependences among word errors while maintaining training feasibility through probabilistic alignment estimation.

Inventive Principle:
Principle #15Dynamics

4Measurement precision

If N-best lists are used for MWE training with BPTT on GPUs, then the expected word error rate is minimized, but the computation requires parallelization to be practical

Engineering Contradiction:
Improveexpected word error rate minimizationVSAvoidcomputation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent substitutes mechanical sequential computation with parallel computation on GPUs. By formulating the MWE training with BPTT in a way that can be parallelized across multiple hypotheses and time steps, the patent replaces the sequential mechanical processing bottleneck with concurrent computational operations, enabling practical implementation of the computationally intensive WER minimization approach.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10176799B2Method and system for training language models to reduce recognition errors
Publication Date: 2019.01.08 MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC
  • US10176799B2 patent drawing
  • US10176799B2 patent drawing
  • US10176799B2 patent drawing

AI summary

A method and for training a language model to reduce recognition errors, wherein the language model is a recurrent neural network language model (RNNLM) by first acquiring training samples. An automatic speech recognition system (ASR) is applied to the training samples to produce recognized words and probabilites of the recognized words, and an N-best list is selected from the recognized words based on the probabilities. determining word errors using reference data for hypotheses in the N-best list. The hypotheses are rescored using the RNNLM. Then, we determine gradients for the hypotheses using the word errors and gradients for words in the hypotheses. Lastly, parameters of the RNNLM are updated using a sum of the gradients.