Multimodal Translation Model Training with Mixed Data Types

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multimodal machine translation models suffer from low translation accuracy due to the scarcity and difficulty in labeling data, particularly in scenarios involving text and image data, leading to inefficiencies in e-commerce and conversation applications.

Innovation Solution

A translation method and apparatus that utilize a target multimodal translation model trained on a diverse set of sample data including multimodal multilingual, unimodal multilingual, and multimodal monolingual data to enhance the training process, incorporating image and text features to improve translation accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multimodal machine translation models are used to improve translation accuracy, then translation quality is enhanced, but data scarcity and labeling difficulty worsen the training process

Engineering Contradiction:
Improvetranslation accuracyVSAvoidtraining data availability
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The training data is segmented into three distinct types: multimodal multilingual data (source text, target text, and image), unimodal multilingual data (source text and target text without images), and multimodal monolingual data (target text and image). This segmentation allows the model to learn from different data characteristics and overcome the limitation of scarce labeled multimodal data by leveraging abundant unimodal and monolingual data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extends the training from traditional unimodal (text-only) to multimodal (text + image) dimensionality. By incorporating image modality alongside text, the model gains additional semantic information that improves translation accuracy, particularly for disambiguation in e-commerce and conversation scenarios where visual context is available.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If multimodal data is used for training, then translation accuracy is improved, but data labeling difficulty increases

Engineering Contradiction:
Improvetranslation accuracyVSAvoiddata labeling ease
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent employs partial action by utilizing unimodal multilingual data (without images) and multimodal monolingual data (with images but single language) as supplementary training sources. These partial data types require less or no complex multimodal labeling, yet they contribute to improving the overall model performance by providing additional learning signals that complement the fully labeled multimodal multilingual data.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If diverse sample data is incorporated to increase training samples, then model performance is improved, but training complexity increases

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining process complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent designs a unified training framework that can handle three different data types (multimodal multilingual, unimodal multilingual, and multimodal monolingual) using a single model architecture. This multi-functional approach allows the model to learn from diverse data sources without requiring separate training pipelines, thereby improving performance while managing training complexity through a universal processing mechanism.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20260050754A1Translation method and apparatus, readable medium, and electronic device
Publication Date: 2026.02.19 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20260050754A1 patent drawing
  • US20260050754A1 patent drawing
  • US20260050754A1 patent drawing

AI summary

Embodiments of the present disclosure relate to a translation method and apparatus, a readable medium, and an electronic device. The method includes: determining a source text to be translated and a source associated image corresponding to the source text; and inputting the source text and the source associated image into a pre-generated target multimodal translation model to obtain a target translation text output by the target multimodal translation model. The target multimodal translation model is a model generated by training an undetermined multimodal translation model according to sample data, and the sample data includes at least two of multimodal multilingual data, unimodal multilingual data, and multimodal monolingual data. The multimodal multilingual data includes a first source-language text, a first target-language text, and a first image corresponding to the first source-language text, the unimodal multilingual data includes a second source-language text and a second target-language text, and the multimodal monolingual data includes a third target-language text and a second image.