Joint Speech Recognition Training With Robust Intermediate Representations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition methods suffer from low accuracy due to the mismatch between human-oriented speech separation and enhancement tasks and machine-oriented speech recognition tasks, leading to suboptimal performance in converting speech signals into text sequences.
Innovation Solution
A joint training method is employed to bridge the gap between speech separation and enhancement models and speech recognition models using a robust representation model, where an intermediate model is trained with a fused loss function, enabling end-to-end learning and improving the accuracy of speech recognition by aligning human-oriented and machine-oriented tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech separation and enhancement are performed using conventional separate models, then the processing pipeline is simple, but speech recognition accuracy is low
Solution Approach 1:
The patent merges speech separation, speech enhancement, and speech recognition models into a unified joint training framework. The combined model shares acoustic feature extraction components and uses a fused loss function that integrates separation loss, enhancement loss, and recognition loss, allowing the system to achieve high recognition accuracy while maintaining architectural elegance through shared representations.
Solution Approach 2:
The patent introduces an intermediate robust representation model that bridges the gap between speech separation/enhancement and speech recognition. This intermediate layer transforms separated and enhanced speech features into representations optimized for recognition, acting as a mediator that resolves the mismatch between front-end processing and back-end recognition objectives.
2Adaptability or versatility
If human-oriented speech separation and enhancement tasks are used, then the processing is straightforward, but there is a mismatch with machine-oriented speech recognition tasks
Solution Approach 1:
The patent creates a universal model framework that simultaneously performs speech separation, speech enhancement, and speech recognition. The shared acoustic feature extraction and fused loss function enable the system to handle multiple tasks with a single unified architecture, making the model adaptable to different processing objectives while maintaining optimal performance for speech recognition.
Solution Approach 2:
The patent changes the optimization parameters by introducing a fused loss function that combines separation loss, enhancement loss, and recognition loss with weighted contributions. This parameter fusion aligns the optimization objectives across different tasks, ensuring that the model learns representations that satisfy both human-oriented separation/enhancement requirements and machine-oriented recognition accuracy.
3Reliability
If separate training of speech separation and speech recognition models is performed, then training is simpler, but overall performance is suboptimal
Solution Approach 1:
The patent combines separate training processes into a unified joint training framework. The model uses a fused loss function that integrates separation loss, enhancement loss, and recognition loss, allowing all components to be trained simultaneously with coordinated optimization. This merging ensures that feature extraction, separation, enhancement, and recognition are optimized together, improving overall system reliability and performance.
Solution Approach 2:
The joint training framework implements feedback mechanisms where the recognition loss provides gradient information back to the separation and enhancement components. This feedback loop ensures that the front-end processing (separation and enhancement) is optimized according to the actual recognition performance, creating a closed-loop system that continuously improves overall effectiveness.
Data Source
AI summary
This application relates to a speech recognition method and apparatus, and a computer-readable storage medium, and the method includes: obtaining a first loss function of a speech separation and enhancement model and a second loss function of a speech recognition model; performing back propagation based on the second loss function to train an intermediate model bridged between the speech separation and enhancement model and the speech recognition model, to obtain a representation model; fusing the first loss function and the second loss function, to obtain a target loss function; and jointly training the speech separation and enhancement model, the representation model, and the speech recognition model based on the target loss function, and ending the training when a preset convergence condition is met.


