RNN-T Probability Distribution Softening for External Language Model Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Recurrent Neural Network Transducer (RNN-T) architectures in speech recognition lack explicit language models, leading to a peaky probability distribution over output symbols, making effective fusion with external language models difficult.
Innovation Solution
Transforming the probability distribution of RNN-T outputs using a non-linear function to relax sharpness, allowing for fusion with external language models by distorting the probability distribution and lowering the amplitudes of higher probabilities, and using the max function to search for the best output sequence.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If RNN-T architecture is used for end-to-end speech recognition, then decoding speed is improved, but the probability distribution becomes peaky making fusion with external language model difficult
Solution Approach 1:
The patent applies parameter changes by transforming the probability distribution parameters through a temperature parameter (T) and a sharpness control parameter (alpha). The transformation modifies the peaky probability distribution from RNN-T into a softer distribution that can be effectively fused with external language models, while preserving the speed advantage of RNN-T architecture.
Solution Approach 2:
The patent introduces an intermediary transformation layer between RNN-T and the external language model. This intermediary applies a softening function to the RNN-T probability distribution, creating a bridge that enables effective fusion with the external language model without losing the benefits of either component.
2Adaptability or versatility
If probability distribution is transformed to relax sharpness, then fusability with external language model is improved, but decoding accuracy may be affected
Solution Approach 1:
The patent employs dynamic parameters (temperature T and sharpness control alpha) that can be adjusted during decoding. This allows the system to dynamically balance between maintaining the sharpness needed for accurate RNN-T predictions and softening the distribution for effective external language model fusion, thereby preserving decoding accuracy while improving fusability.
Solution Approach 2:
The system uses feedback mechanisms to adjust the transformation parameters based on the interaction between the transformed RNN-T distribution and the external language model. This feedback loop ensures that the softening transformation maintains decoding accuracy by adapting to the specific characteristics of the external language model being fused.
Data Source
AI summary
A computer-implemented method for fusing an end-to-end speech recognition model with an external language model (ExternalLM) is provided. The method includes obtaining an output of the end-to-end speech recognition model. The output is a probability distribution. The method further includes transforming, by a hardware processor, the probability distribution into a transformed probability distribution to relax a sharpness of the probability distribution. The method also includes fusing the transformed probability distribution and a probability distribution of the ExternalLM for decoding speech.


