Speech Translation Model Dual-Task Training for Bias Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech translation technologies face challenges with low translation quality and weak coherence of translation text, leading to low accuracy and inconvenience in user communication.

Innovation Solution

A training method for a speech translation model that includes executing a speech translation training task and an auxiliary training task simultaneously, with the auxiliary task aimed at reducing translation bias caused by speech recognition errors, to optimize robustness and generalization of speech translation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If speech translation model is trained using conventional methods, then training simplicity is maintained, but translation quality and coherence are poor

Engineering Contradiction:
Improvetranslation qualityVSAvoidtraining complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The training process is segmented into two distinct tasks: the speech translation training task for learning translation capabilities, and the auxiliary training task for reducing speech recognition bias. This segmentation allows each task to be optimized independently while working together to improve overall translation quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The auxiliary training task is performed in parallel during the training phase to preemptively reduce translation bias before it affects the final translation output. This preliminary action addresses the bias issue proactively rather than correcting it post-training.

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If speech recognition is performed on speech input, then speech to text conversion is achieved, but speech recognition bias is introduced

Engineering Contradiction:
Improvespeech input capabilityVSAvoidtext accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The auxiliary training task converts the harmful speech recognition bias into a beneficial training signal by using biased speech recognition texts as training data. The model learns to recognize and compensate for the bias patterns, transforming the previously harmful effect into a useful preprocessing step that actually helps the translation task.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The auxiliary training task acts as an intermediary between speech recognition and translation by processing the biased speech recognition output and preparing it for translation. This intermediate processing step reduces the bias before the text enters the translation model.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If domain data such as ASR texts are introduced to enhance translation, then translation coverage is improved, but translation coherence deteriorates

Engineering Contradiction:
Improvetranslation coverageVSAvoidtranslation coherence
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The model's parameters are adjusted through dual-task training to change its behavior when processing domain-specific texts. The auxiliary training task modifies the model's internal representations to maintain coherence even when handling specialized terminology and structures from domain data.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12308015B2Training method and apparatus for speech translation model, speech translation method and apparatus, and device
Publication Date: 2025.05.20 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US12308015B2 patent drawing
  • US12308015B2 patent drawing
  • US12308015B2 patent drawing

AI summary

A training method for a speech translation model, a speech translation method, an apparatus, and a device are provided. The training method includes: after entering a model training phase, controlling the speech translation model to execute a speech translation training task; controlling the speech translation model to simultaneously execute an auxiliary training task of the speech translation training task; and adjusting network parameters of the speech translation model according to the speech translation training task and the auxiliary training task, so as to obtain an updated speech translation model after training.