Bilingual Corpora Screening via Multi-Model Quality Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current corpus cleaning methods in neural machine translation rely heavily on artificial rules or statistical methods, limiting data volume and efficiency in filtering and cleaning bilingual corpora.
Innovation Solution
A method involving acquiring multiple pairs of bilingual corpora, training machine translation and language models, obtaining feature vectors, and determining quality values to comprehensively screen and filter corpora based on these features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If artificial rules or statistical methods are used for corpus cleaning, then the cleaning process is simple to implement, but the data volume and efficiency of corpus cleaning are reduced
Solution Approach 1:
The patent transforms the corpus cleaning process from rule-based filtering to a quality-score-based selection process. It introduces multiple evaluation dimensions (translation quality, language model probability, feature vector matching) and calculates comprehensive quality scores to rank and select corpora, fundamentally changing the cleaning criterion from simple rule matching to multi-parameter optimization
Solution Approach 2:
The patent combines multiple evaluation methods (machine translation model, language model, feature vector analysis) into a composite evaluation system. Each method contributes different aspects of quality assessment, and their results are integrated through weighted scoring to form a comprehensive corpus quality evaluation framework
2Ease of manufacture
If artificial rules or statistical methods are used for corpus cleaning, then the implementation is straightforward, but the data volume of cleaned corpora is limited
Solution Approach 1:
The patent changes the selection criterion from binary rule matching to continuous quality scoring. By calculating quality scores and setting threshold values, the system can flexibly adjust the proportion of selected corpora, thereby increasing the data volume of cleaned corpora while maintaining quality standards
Solution Approach 2:
The patent creates a universal corpus cleaning framework that can handle multiple types of bilingual corpora (parallel corpora, sentence pairs, phrase pairs) through a unified quality evaluation mechanism. The system processes different corpus types using the same multi-dimensional assessment approach, expanding the applicable data volume
3Productivity
If targeted filtering through regular expressions is used, then the filtering is efficient for specific problems, but it cannot handle various pairs of bilingual corpora
Solution Approach 1:
The patent develops a universal corpus cleaning framework that processes different types of bilingual corpora (parallel corpora, sentence pairs, phrase pairs) through a unified quality evaluation mechanism. The system uses multiple evaluation dimensions (translation quality, language model probability, feature vector matching) that can assess various corpus types consistently, thereby achieving both efficiency and versatility
Solution Approach 2:
The patent transforms the filtering approach from problem-specific rule matching to a generalizable quality scoring system. By introducing multiple evaluation parameters and calculating comprehensive quality scores, the system can adapt to different corpus types and cleaning requirements while maintaining high processing efficiency
4Measurement precision
If corpus cleaning is performed for a specific situation, then the cleaning is focused, but it affects the data volume and reduces cleaning efficiency
Solution Approach 1:
The patent changes the cleaning approach from situation-specific filtering to comprehensive quality scoring. By evaluating corpora across multiple dimensions (translation quality, language model probability, feature vector matching) and calculating overall quality scores, the system achieves both precise quality control and high processing efficiency for large-scale corpora
Data Source
AI summary
A bilingual corpora screening method includes: acquiring multiple pairs of bilingual corpora, wherein each pair of the bilingual corpora comprises a source corpus and a target corpus; training a machine translation model based on the multiple pairs of bilingual corpora; obtaining a first feature of each pair of bilingual corpora based on the trained machine translation model; training a language model based on the multiple pairs of bilingual corpora; obtaining feature vectors of each pair of bilingual corpora and determining a second feature of each pair of bilingual corpora based on the trained language model; determining a quality value of each pair of bilingual corpora according to the first feature and the second feature of each pair of bilingual corpora; and screening each pair of bilingual corpora according to the quality value of each pair of bilingual corpora.


