Joint Speech Recognition Training With Robust Intermediate Representations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition methods suffer from low accuracy due to the mismatch between human-oriented speech separation and enhancement tasks and machine-oriented speech recognition tasks, leading to suboptimal performance in converting speech signals into text sequences.

Innovation Solution

A joint training method is employed to bridge the gap between speech separation and enhancement models and speech recognition models using a robust representation model, where an intermediate model is trained with a fused loss function, enabling end-to-end learning and improving the accuracy of speech recognition by aligning human-oriented and machine-oriented tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech separation and enhancement are performed using conventional separate models, then the processing pipeline is simple, but speech recognition accuracy is low

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidmodel architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges speech separation, speech enhancement, and speech recognition models into a unified joint training framework. The combined model shares acoustic feature extraction components and uses a fused loss function that integrates separation loss, enhancement loss, and recognition loss, allowing the system to achieve high recognition accuracy while maintaining architectural elegance through shared representations.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediate robust representation model that bridges the gap between speech separation/enhancement and speech recognition. This intermediate layer transforms separated and enhanced speech features into representations optimized for recognition, acting as a mediator that resolves the mismatch between front-end processing and back-end recognition objectives.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If human-oriented speech separation and enhancement tasks are used, then the processing is straightforward, but there is a mismatch with machine-oriented speech recognition tasks

Engineering Contradiction:
Improvetask compatibilityVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent creates a universal model framework that simultaneously performs speech separation, speech enhancement, and speech recognition. The shared acoustic feature extraction and fused loss function enable the system to handle multiple tasks with a single unified architecture, making the model adaptable to different processing objectives while maintaining optimal performance for speech recognition.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes the optimization parameters by introducing a fused loss function that combines separation loss, enhancement loss, and recognition loss with weighted contributions. This parameter fusion aligns the optimization objectives across different tasks, ensuring that the model learns representations that satisfy both human-oriented separation/enhancement requirements and machine-oriented recognition accuracy.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If separate training of speech separation and speech recognition models is performed, then training is simpler, but overall performance is suboptimal

Engineering Contradiction:
Improveoverall system performanceVSAvoidtraining process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines separate training processes into a unified joint training framework. The model uses a fused loss function that integrates separation loss, enhancement loss, and recognition loss, allowing all components to be trained simultaneously with coordinated optimization. This merging ensures that feature extraction, separation, enhancement, and recognition are optimized together, improving overall system reliability and performance.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The joint training framework implements feedback mechanisms where the recognition loss provides gradient information back to the separation and enhancement components. This feedback loop ensures that the front-end processing (separation and enhancement) is optimized according to the actual recognition performance, creating a closed-loop system that continuously improves overall effectiveness.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12609109B2Speech recognition method and apparatus, and computer-readable storage medium
Publication Date: 2026.04.21 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12609109B2 patent drawing
  • US12609109B2 patent drawing
  • US12609109B2 patent drawing

AI summary

This application relates to a speech recognition method and apparatus, and a computer-readable storage medium, and the method includes: obtaining a first loss function of a speech separation and enhancement model and a second loss function of a speech recognition model; performing back propagation based on the second loss function to train an intermediate model bridged between the speech separation and enhancement model and the speech recognition model, to obtain a representation model; fusing the first loss function and the second loss function, to obtain a target loss function; and jointly training the speech separation and enhancement model, the representation model, and the speech recognition model based on the target loss function, and ending the training when a preset convergence condition is met.