Language Model In-Domain Training for Translation Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine translation systems often generate translations that are grammatically correct but lack context relevance, as they are trained on general monolingual data, which may not cover specific topics commonly encountered in translation requests, leading to inconsistencies and reduced accuracy.

Innovation Solution

Supplementing the translation system's training data with in-domain data that includes materials relevant to the topics typically discussed in translation requests, allowing the language model to better understand context and preferences specific to source and destination languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the translation system is trained on general monolingual data, then the translation system can process a wide variety of topics, but the translation accuracy and context relevance deteriorate for specific domains

Engineering Contradiction:
Improvecoverage of topicsVSAvoidtranslation accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The training data is segmented into general monolingual data and in-domain supplemental data. The system uses both types of data in a combined training approach, where general data provides broad coverage and in-domain data provides specialized accuracy for specific topics and domains.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If the translation system uses in-domain supplemental training data, then the translation accuracy and context relevance improve, but the data processing complexity increases

Engineering Contradiction:
Improvetranslation accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system merges general monolingual training data with in-domain supplemental training data into a unified training process. The language model is trained on both data types simultaneously, allowing the system to handle both general and domain-specific translation tasks without requiring separate processing pipelines.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If the translation system is trained on in-domain data specific to source and destination languages, then the context relevance improves, but the training data requirements increase

Engineering Contradiction:
Improvecontext relevanceVSAvoidtraining data quantity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Instead of requiring uniform large amounts of training data for all topics, the system applies local quality by using in-domain supplemental data specifically for topics and domains where translation accuracy is most needed. The general monolingual data provides baseline coverage while in-domain data enhances specific local areas of the translation capability.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10460040B2Language model using reverse translations
Publication Date: 2019.10.29 META PLATFORMS INC
  • US10460040B2 patent drawing
  • US10460040B2 patent drawing
  • US10460040B2 patent drawing

AI summary

Exemplary embodiments relate to techniques for improving machine translation systems. The machine translation system may apply one or more models for translating material from a source language into a destination language. The models are initially trained using training data. According to exemplary embodiments, supplemental training data is used to train the models, where the supplemental training data uses in-domain material to improve the quality of output translations. In-domain data may include data that relates to the same or similar topics as those expected to be encountered in a translation of material from the source language into the destination language. In-domain data may include material previously translated from the source language into the destination language, material similar to previous translations, and destination language material that has previously been the subject of a request for translation into the source language.