E2E Speech Recognition Calibration for Acoustic-Language Model Coupling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing E2E speech recognition systems face challenges in ensuring compatibility and reliability between acoustic and language models, leading to suboptimal performance due to independent learning and propagated errors.
Innovation Solution
A method and apparatus for generating an E2E speech recognition model using calibration correction, which adjusts entropy outputs of acoustic and language models through expected calibration error (ECE) to minimize gaps and improve compatibility, involving the generation of correction parameters to optimize the coupling probability distribution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If independent module structure with separate acoustic model and language model is used, then each module can be learned separately and errors can be corrected intuitively, but compatibility between modules cannot be ensured and errors are propagated
Solution Approach 1:
The patent segments the speech recognition model into distinct acoustic model and language model components, each trained separately on their respective datasets. This allows independent optimization and error correction for each module while maintaining overall system functionality through separate training pipelines and modular architecture.
Solution Approach 2:
The patent combines the separately trained acoustic model and language model into an integrated speech recognition system where both models work together. The acoustic model processes speech signals while the language model processes text sequences, and their outputs are combined to produce final recognition results, achieving both modularity and integration.
2Extent of automation
If E2E speech recognition technology is used, then the entire process is learned through deep learning without human intervention, but compatibility with each module and reliability of output results cannot be ensured
Solution Approach 1:
The patent performs preliminary training of the acoustic model and language model separately before combining them into the E2E system. This preliminary action allows each component to be optimized independently with appropriate loss functions and training data, ensuring reliability before full integration and automation.
Solution Approach 2:
The patent implements feedback mechanisms where the output of one model is fed back into the training process of the other model. The acoustic model's text outputs are used to train the language model, and vice versa, creating a feedback loop that improves overall system reliability and compatibility through iterative optimization.
Data Source
AI summary
A speech recognition model generating device for generating an E2E speech recognition model using calibration correction comprising an acoustic model including a first artificial neural network module using a speech information as input information and using a first text information corresponding to the speech information as output information, a language model comprising a second artificial neural network module using the first text information as input information and outputting a second text information corresponding to the first text information as output information based on the characteristics of the language model, and a E2E speech model generating unit generating a coupling probability distribution based on a first probability distribution information of the acoustic model output by the acoustic model and a second probability distribution information of the language model output by the language model, and generating a E2E speech model based on the coupling probability distribution, wherein the E2E speech model generating unit is generating the E2E speech model based on the corrected acoustic model and the corrected language model after each calibration is performed on the acoustic model and the language model.


