Text Region Extraction and Editing for Translation Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current translation technologies face challenges in accurately extracting and editing text regions from images, particularly in handling unreliable text recognition and merging or dividing text regions effectively for efficient translation processes.
Innovation Solution
A computer-readable medium storing a translation program that includes a document receiver, text region extractor, text recognizer, translator, display controller, text region editor, reliability degree calculator, and text region estimator, which work together to extract text regions from images, perform character recognition, edit text regions based on user operations, and calculate reliability degrees to improve translation accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If text regions are extracted from image information using character recognition, then text can be obtained for translation, but the reliability of text recognition may be insufficient leading to inaccurate text extraction
Solution Approach 1:
The system displays extracted text regions and recognized text to users for verification and correction. Users can provide feedback by selecting incorrect regions or text, which triggers re-extraction or correction processes, thereby improving both accuracy and reliability through iterative refinement
Solution Approach 2:
The system performs preliminary text extraction and recognition before final translation, allowing users to review and correct results in advance. This preliminary action enables error detection and correction before the translation process begins, improving overall reliability
2Manufacturing precision
If text regions are edited based on user operations, then translation accuracy can be improved, but the complexity of the system increases due to multiple processing steps
Solution Approach 1:
The display controller serves multiple functions: displaying image information, showing extracted text regions, presenting recognized text, and receiving user corrections. This multi-functionality reduces the need for separate dedicated components for each task, managing system complexity while maintaining high translation accuracy
Solution Approach 2:
The system divides the translation process into distinct segments: image processing, text extraction, character recognition, user verification, and translation. Each segment handles a specific task, making the overall complex process manageable and allowing for targeted improvements in each stage
3Reliability
If text regions are merged or divided based on reliability calculations, then translation quality can be enhanced, but the processing time increases
Solution Approach 1:
The system calculates reliability parameters for extracted text regions and uses these parameters to dynamically adjust processing decisions. High-reliability regions are processed directly while low-reliability regions trigger re-extraction or user verification, optimizing the balance between quality and processing time through parameter-based decision making
Data Source
AI summary
A non-transitory computer readable medium storing a translation program causes a computer to execute a process. The process includes: displaying image information, text regions, and original text in association with each other, the text regions being obtained by extracting regions including an image of text from the image information, the original text being obtained by performing character recognition on the text included in the text regions; and editing the text regions in accordance with the content of a received operation.


