Reinforcement Learning Data Filtering for Machine Translation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing methods for acquiring high-quality bilingual training data for machine translation models are costly due to the need for manual translation of source language data, resulting in poor quality source language data after filtering, which affects the translation performance of machine translation models.
Innovation Solution
A data processing method using a reinforcement learning algorithm to filter source language data and acquire corresponding markup language data, improving the quality of source language data and enhancing the translation performance of machine translation models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If source language data is filtered based on term frequency or model confidence, then the filtering process is simple and fast, but the quality of source language data remaining after filtering is poor
Solution Approach 1:
The patent changes the filtering parameters from traditional term frequency or model confidence to a reinforcement learning-based filtering model that evaluates data quality through multiple dimensions including terminology accuracy, grammatical correctness, and semantic coherence. This parameter transformation enables both high-speed filtering and high-quality data selection by using a trained RL model that can rapidly assess data quality without manual intervention.
Solution Approach 2:
The patent replaces manual or simple rule-based filtering mechanisms with an automated reinforcement learning-based filtering system. The RL model learns optimal filtering strategies through training and can automatically identify high-quality source language data, substituting the need for complex manual curation or basic keyword filtering with an intelligent automated system that achieves both speed and quality.
2Reliability
If a large amount of source language data is filtered to acquire high-quality bilingual training data, then the translation performance can be improved, but the cost of obtaining markup language data increases
Solution Approach 1:
The patent applies preliminary action by filtering and selecting high-quality source language data before acquiring corresponding markup language data. The reinforcement learning model pre-evaluates source data quality and selects only the most promising candidates for translation, ensuring that markup language data is acquired only for high-quality source data. This preliminary selection action prevents waste of resources on low-quality data that would not contribute to improving translation performance.
Solution Approach 2:
The reinforcement learning-based filtering system performs self-service by automatically identifying and selecting high-quality source language data without requiring manual intervention or expensive professional translation services for every data point. The RL model uses its trained knowledge to autonomously assess data quality and make selection decisions, reducing the need for costly human resources in the data acquisition process while maintaining high translation performance.
3Ease of manufacture
If traditional filtering rules are used, then the filtering process is easy to implement, but the application scenarios are limited and data quality is poor
Solution Approach 1:
The patent achieves universality by developing a reinforcement learning-based filtering model that can adapt to multiple application scenarios and different language pairs. The RL model is trained to recognize patterns and qualities that are relevant across diverse domains, enabling it to effectively filter source language data for various translation tasks including technical, legal, medical, and general-purpose translation. This multi-functional approach replaces narrow, scenario-specific filtering rules with a general-purpose intelligent filtering system.
Data Source
AI summary
A data processing method is described. The method includes acquiring a to-be-filtered dataset, the to-be-filtered dataset including a plurality of pieces of to-be-filtered source language data; filtering all source language data in the to-be-filtered dataset based on a target data filtering model to obtain target source language data remaining after the filtering, the target data filtering model being obtained through training performed by using a reinforcement learning algorithm; and acquiring markup language data corresponding to the obtained target source language data, and acquiring a machine translation model based on the target source language data and the acquired markup language data. In such a data processing process, a filtering rule in the target data filtering model is automatically learned by a machine in a reinforcement learning process. Apparatus and non-transitory computer-readable storage medium counterpart embodiments are also provided.


