Bilingual Corpus Update via Segmented Phrase Evaluation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for evaluating paraphrastic sentences in bilingual corpora are inefficient and imprecise, particularly when relying on databases with poor data quality or lacking sufficient illustrative sentences, which hampers the improvement of machine translation performance.
Innovation Solution
A method that involves inputting a third sentence created by replacing a phrase in an original sentence, judging its inclusion in databases for written and spoken text, calculating evaluation values, and adding the sentence to the bilingual corpus if it meets predetermined conditions, utilizing a combination of general-purpose and colloquial expression n-gram databases for precise evaluation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single database is used for evaluating paraphrastic sentences, then the evaluation process is simple, but the precision and comprehensiveness of evaluation deteriorates
Solution Approach 1:
The patent segments the evaluation process into two distinct stages: first evaluating whether the paraphrastic sentence contains phrases present in the original sentence (using the first database), and second evaluating whether the paraphrastic sentence contains phrases present in parallel sentences (using the second database). This segmentation allows comprehensive evaluation while maintaining clear procedural structure.
Solution Approach 2:
The patent introduces a two-dimensional evaluation framework by utilizing two different databases with different evaluation criteria. The first database evaluates phrase overlap with the original sentence, while the second database evaluates phrase overlap with parallel sentences. This multi-dimensional approach enhances evaluation precision without excessive complexity.
2Measurement precision
If multiple databases are used for evaluating paraphrastic sentences, then the evaluation precision is improved, but the device complexity increases
Solution Approach 1:
The patent divides the evaluation system into two independent modules, each using a different database. The first module evaluates phrase presence in the original sentence, and the second module evaluates phrase presence in parallel sentences. This segmentation reduces perceived complexity by making each module's function clear and distinct.
Solution Approach 2:
The patent implements a staged evaluation approach where the first database evaluation is performed partially (checking if any phrase is present), and only if this condition is met does the system proceed to the second database evaluation. This partial action approach reduces overall system complexity by avoiding unnecessary evaluations.
3Manufacturing precision
If strict evaluation criteria are applied to paraphrastic sentences, then the quality of bilingual corpus is improved, but the quantity of usable sentences decreases
Solution Approach 1:
The patent applies partial evaluation by checking whether at least one phrase is present in the original sentence and at least one phrase is present in parallel sentences, rather than requiring complete phrase overlap. This partial action approach maintains quality standards while increasing the quantity of acceptable paraphrastic sentences.
Solution Approach 2:
The patent changes the evaluation parameters from requiring complete sentence equivalence to requiring partial phrase presence. By evaluating whether phrases (rather than complete sentences) are present in the databases, the system maintains quality control while accepting a broader range of paraphrastic variations, thus increasing corpus quantity.
Data Source
AI summary
A third sentence obtained by replacing a first phrase of a first sentence with a second phrase is input, and it is judged whether a third phrase is included in a first database including at least a phrase used in written text. If the third phrase is not included, a first evaluation value in the first database is calculated for a seventh phrase obtained by replacing the second phrase of the third phrase with a sixth phrase. It is judged whether the third phrase is included in a second database including at least a phrase used in spoken text and whether a second evaluation value calculated from the first evaluation value satisfies a predetermined condition. If the third phrase is included, and the second evaluation value satisfies the predetermined condition, the third sentence and the second sentence as a pair are added to a bilingual corpus.


