RNN-T Probability Distribution Softening for External Language Model Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Recurrent Neural Network Transducer (RNN-T) architectures in speech recognition lack explicit language models, leading to a peaky probability distribution over output symbols, making effective fusion with external language models difficult.

Innovation Solution

Transforming the probability distribution of RNN-T outputs using a non-linear function to relax sharpness, allowing for fusion with external language models by distorting the probability distribution and lowering the amplitudes of higher probabilities, and using the max function to search for the best output sequence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If RNN-T architecture is used for end-to-end speech recognition, then decoding speed is improved, but the probability distribution becomes peaky making fusion with external language model difficult

Engineering Contradiction:
Improvedecoding speedVSAvoidfusability with external language model
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The patent applies parameter changes by transforming the probability distribution parameters through a temperature parameter (T) and a sharpness control parameter (alpha). The transformation modifies the peaky probability distribution from RNN-T into a softer distribution that can be effectively fused with external language models, while preserving the speed advantage of RNN-T architecture.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary transformation layer between RNN-T and the external language model. This intermediary applies a softening function to the RNN-T probability distribution, creating a bridge that enables effective fusion with the external language model without losing the benefits of either component.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If probability distribution is transformed to relax sharpness, then fusability with external language model is improved, but decoding accuracy may be affected

Engineering Contradiction:
Improvefusability with external language modelVSAvoiddecoding accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent employs dynamic parameters (temperature T and sharpness control alpha) that can be adjusted during decoding. This allows the system to dynamically balance between maintaining the sharpness needed for accurate RNN-T predictions and softening the distribution for effective external language model fusion, thereby preserving decoding accuracy while improving fusability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback mechanisms to adjust the transformation parameters based on the interaction between the transformed RNN-T distribution and the external language model. This feedback loop ensures that the softening transformation maintains decoding accuracy by adapting to the specific characteristics of the external language model being fused.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20230069628A1External language model fusing method for speech recognition
Publication Date: 2023.03.02 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20230069628A1 patent drawing
  • US20230069628A1 patent drawing
  • US20230069628A1 patent drawing

AI summary

A computer-implemented method for fusing an end-to-end speech recognition model with an external language model (ExternalLM) is provided. The method includes obtaining an output of the end-to-end speech recognition model. The output is a probability distribution. The method further includes transforming, by a hardware processor, the probability distribution into a transformed probability distribution to relax a sharpness of the probability distribution. The method also includes fusing the transformed probability distribution and a probability distribution of the ExternalLM for decoding speech.