Multi-lingual Model Training via Dual-task Semantic Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional multi-lingual models pre-trained using bilingual or monolingual corpora fail to accurately learn semantic alignment information between different languages, limiting their ability to perform effective information interaction across languages.

Innovation Solution

A method involving two training tasks: the first task uses bilingual corpora to predict masked semantic units in source and target languages, and the second task generates pseudo parallel corpora from monolingual corpora to further enhance semantic alignment, with convergence of loss functions indicating completion of training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional multi-lingual models are pre-trained using bilingual or monolingual corpora, then the model can be built with basic language capabilities, but the model cannot learn semantic alignment information between different languages

Engineering Contradiction:
Improvesemantic alignment accuracyVSAvoidcross-language information interaction capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the training process into two distinct tasks: the first training task uses bilingual corpora to learn semantic alignment between languages, while the second training task uses monolingual corpora to learn language-specific representations. This segmentation allows each task to focus on specific learning objectives without interference, resolving the contradiction between learning semantic alignment and maintaining language-specific capabilities

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension to the training approach by using pseudo-parallel corpora generated from monolingual corpora. This creates an additional training dimension that enables the model to learn semantic alignment information without relying solely on traditional bilingual corpora, thereby improving cross-language information interaction capability while maintaining basic language capabilities

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If the model is trained to predict masked semantic units in bilingual corpora, then semantic alignment can be learned, but the training process becomes complex with multiple training tasks

Engineering Contradiction:
Improvesemantic alignment learningVSAvoidtraining process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by first generating pseudo-parallel corpora from monolingual corpora before the main training process. This preparation step creates additional training data that facilitates the subsequent dual-task training process, making the complex training more effective by providing pre-processed semantic alignment information that the model can leverage during both training tasks

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11995405B2Multi-lingual model training method, apparatus, electronic device and readable storage medium
Publication Date: 2024.05.28 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11995405B2 patent drawing
  • US11995405B2 patent drawing
  • US11995405B2 patent drawing

AI summary

The present disclosure provides a multi-lingual model training method, apparatus, electronic device and readable storage medium and relates to the technical field of deep learning and natural language processing. A technical solution of the present disclosure when training the multi-lingual model is: obtaining training corpuses comprising a plurality of bilingual corpuses and a plurality of monolingual corpuses; training a multi-lingual model with a first training task by using the plurality of bilingual corpuses; training the multi-lingual model with a second training task by using the plurality of monolingual corpuses; and completing the training of the multi-lingual model in a case of determining that loss functions of the first training task and second training task converge. In the present disclosure, the multi-lingual model can be enabled to achieve semantic interaction between different languages and improve the accuracy of the multi-lingual model in learning the semantic representations of the multi-lingual model.