Translation System Mistranslation Correction and Corpus Expansion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional machine translation devices face challenges in improving translation performance due to the inability to incorporate mistranslated text into parallel translation corpora, which limits the expansion of example text pairs and increases mistranslation generation.
Innovation Solution
A method in a translation system that receives source text, generates translated text, creates parallel translation data with mistranslated text when errors occur, transmits this data to terminal devices for revision, and registers revised text pairs into a database, allowing for the collection of large example text pairs and preventing mistranslations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If mistranslated text is excluded from parallel translation corpora, then translation quality is maintained, but the corpus size and diversity are limited
Solution Approach 1:
The patent segments the translation process into multiple stages: initial machine translation, quality assessment, human review/revision, and corpus integration. By separating the handling of mistranslated text from the main translation workflow, the system can process and correct errors systematically without compromising overall translation quality while still incorporating diverse examples into the corpus.
Solution Approach 2:
The patent introduces an intermediary quality assessment mechanism that evaluates machine translation outputs before they are integrated into the corpus. This intermediary layer identifies mistranslations and routes them for human review, allowing the system to maintain high quality standards while collecting a broader range of translation examples including initially incorrect ones that are subsequently corrected.
2Measurement precision
If only correctly translated text pairs are registered, then translation accuracy is improved, but the amount of example text for training is reduced
Solution Approach 1:
The patent converts mistranslations from harmful errors into beneficial training data. By systematically collecting, correcting, and integrating previously rejected mistranslated pairs, the system transforms errors into valuable examples that help the machine learning model learn from mistakes and improve future translation accuracy, thereby increasing corpus size without sacrificing quality.
Solution Approach 2:
The patent changes the status parameter of mistranslated text from 'rejected' to 'correctable'. By modifying how the system treats incorrect translations—viewing them as opportunities for improvement rather than failures to be discarded—it enables the integration of a larger volume of text pairs into the training corpus while maintaining accuracy through the correction and review process.
3Reliability
If a strict quality filter is applied to translation corpora, then translation reliability is maintained, but the corpus expansion speed is reduced
Solution Approach 1:
The patent applies preliminary quality assessment and error identification to translation outputs before they are fully integrated into the corpus. By performing quality checks and identifying mistranslations in advance, the system can efficiently process and correct errors at scale, maintaining high reliability standards while accelerating corpus expansion through automated preliminary filtering and routing of suspect translations for review.
Data Source
AI summary
A method executed in a translation system includes: (A) receiving source text in a first language; (B) translating the source text into a second language to generate first translated text; (C) acquiring a determination result as to whether or not the first translated text has been correctly translated; (D) in a case where the first translated text has not been correctly translated, generating first parallel translation data that includes the source text and the first translated text; (E) transmitting the first parallel translation data to a terminal device; (F) receiving, from the terminal device, second translated text obtained by the source text being correctly translated into the second language; and (G) registering, into a parallel translation database, second parallel translation data that includes the source text and the second translated text.


