Reinforcement Learning Data Filtering for Machine Translation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing methods for acquiring high-quality bilingual training data for machine translation models are costly due to the need for manual translation of source language data, resulting in poor quality source language data after filtering, which affects the translation performance of machine translation models.

Innovation Solution

A data processing method using a reinforcement learning algorithm to filter source language data and acquire corresponding markup language data, improving the quality of source language data and enhancing the translation performance of machine translation models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If source language data is filtered based on term frequency or model confidence, then the filtering process is simple and fast, but the quality of source language data remaining after filtering is poor

Engineering Contradiction:
Improvefiltering speedVSAvoiddata quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent changes the filtering parameters from traditional term frequency or model confidence to a reinforcement learning-based filtering model that evaluates data quality through multiple dimensions including terminology accuracy, grammatical correctness, and semantic coherence. This parameter transformation enables both high-speed filtering and high-quality data selection by using a trained RL model that can rapidly assess data quality without manual intervention.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces manual or simple rule-based filtering mechanisms with an automated reinforcement learning-based filtering system. The RL model learns optimal filtering strategies through training and can automatically identify high-quality source language data, substituting the need for complex manual curation or basic keyword filtering with an intelligent automated system that achieves both speed and quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If a large amount of source language data is filtered to acquire high-quality bilingual training data, then the translation performance can be improved, but the cost of obtaining markup language data increases

Engineering Contradiction:
Improvetranslation performanceVSAvoidcost of obtaining markup language data
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by filtering and selecting high-quality source language data before acquiring corresponding markup language data. The reinforcement learning model pre-evaluates source data quality and selects only the most promising candidates for translation, ensuring that markup language data is acquired only for high-quality source data. This preliminary selection action prevents waste of resources on low-quality data that would not contribute to improving translation performance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The reinforcement learning-based filtering system performs self-service by automatically identifying and selecting high-quality source language data without requiring manual intervention or expensive professional translation services for every data point. The RL model uses its trained knowledge to autonomously assess data quality and make selection decisions, reducing the need for costly human resources in the data acquisition process while maintaining high translation performance.

Inventive Principle:
Principle #25Self-service

3Ease of manufacture

If traditional filtering rules are used, then the filtering process is easy to implement, but the application scenarios are limited and data quality is poor

Engineering Contradiction:
Improvefiltering implementationVSAvoidapplication scenarios
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent achieves universality by developing a reinforcement learning-based filtering model that can adapt to multiple application scenarios and different language pairs. The RL model is trained to recognize patterns and qualities that are relevant across diverse domains, enabling it to effectively filter source language data for various translation tasks including technical, legal, medical, and general-purpose translation. This multi-functional approach replaces narrow, scenario-specific filtering rules with a general-purpose intelligent filtering system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12164879B2Data processing method, device, and storage medium
Publication Date: 2024.12.10 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12164879B2 patent drawing
  • US12164879B2 patent drawing
  • US12164879B2 patent drawing

AI summary

A data processing method is described. The method includes acquiring a to-be-filtered dataset, the to-be-filtered dataset including a plurality of pieces of to-be-filtered source language data; filtering all source language data in the to-be-filtered dataset based on a target data filtering model to obtain target source language data remaining after the filtering, the target data filtering model being obtained through training performed by using a reinforcement learning algorithm; and acquiring markup language data corresponding to the obtained target source language data, and acquiring a machine translation model based on the target source language data and the acquired markup language data. In such a data processing process, a filtering rule in the target data filtering model is automatically learned by a machine in a reinforcement learning process. Apparatus and non-transitory computer-readable storage medium counterpart embodiments are also provided.