Machine Translation Evaluation Ontology Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine translation evaluation systems, such as those using the BLEU metric, face limitations in providing detailed feedback on specific linguistic phenomena and ignore linguistic significance, leading to inadequate evaluation of machine translation quality.
Innovation Solution
A system and method for evaluating machine translation quality that includes a bilingual data generator, an example extraction component, and an evaluation component, which generate and update a bilingual corpus with semantic analysis, extract evaluation examples based on ontological categories, and score translation results to provide detailed feedback on various linguistic phenomena.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated evaluation systems use metrics like BLEU to evaluate machine translation, then evaluation speed and cost-effectiveness are improved, but measurement precision and ability to assess specific linguistic phenomena deteriorate
Solution Approach 1:
The evaluation system segments translation quality assessment into multiple independent dimensions including fluency, adequacy, and specific linguistic phenomena (ambiguity, idioms, sentence organization). Each dimension is evaluated separately using targeted metrics and rules, allowing precise measurement of different translation aspects while maintaining automated efficiency.
Solution Approach 2:
The system applies different evaluation criteria and weightings to different linguistic phenomena based on their specific characteristics. For example, ambiguity resolution is evaluated using context-aware rules, while idiom translation uses dedicated lexical databases. This localized quality approach enables precise assessment of each linguistic feature while preserving overall automation.
2Measurement precision
If manual evaluation methods are used to assess machine translation quality, then measurement precision and detailed feedback are improved, but productivity and cost-effectiveness deteriorate
Solution Approach 1:
The evaluation system performs self-assessment by automatically comparing machine translation output against reference translations and linguistic rules. The system generates its own evaluation metrics and feedback without requiring human intervention for each translation, achieving both high precision and automation. Evaluators can configure evaluation parameters once, and the system autonomously executes comprehensive assessments.
Solution Approach 2:
The system implements automated feedback loops where evaluation results are immediately generated and can be used to refine translation models. Detailed feedback on specific linguistic phenomena (e.g., which ambiguities were misresolved, which idioms translated incorrectly) is automatically produced, enabling continuous improvement without manual review while maintaining measurement precision.
3Ease of operation
If evaluation systems focus on overall translation scores, then ease of operation is improved, but measurement precision for specific linguistic phenomena deteriorates
Solution Approach 1:
The system segments evaluation into hierarchical levels: overall translation quality and detailed linguistic phenomenon dimensions. Users can access either the simplified overall score or drill down into specific linguistic aspects (ambiguity, idioms, sentence structure) as needed. This segmentation maintains ease of operation for quick assessments while providing precise measurement capability when detailed analysis is required.
Solution Approach 2:
The evaluation system is designed to serve multiple functions: it can provide quick overall scores for routine monitoring, detailed linguistic analysis for model improvement, and targeted assessment of specific phenomena. This multi-functionality allows the same system to accommodate both simple and complex evaluation needs without compromising precision in either mode.
Data Source
AI summary
A system for evaluating translation quality of a machine translator is discussed. The system includes a bilingual data generator configured to intermittently access a wide area network and generate a bilingual corpus from data received from the wide area network. The method also includes an example extraction component configured to receive an ontology input indicative of a plurality of ontological categories of evaluation and to extract evaluation examples from the bilingual corpus based on the ontology input. The system further includes an evaluation component configured to evaluate translation results from translation by a machine translator of the evaluation examples and to score the translation results according to the ontological categories.


