Web Page Image Character Translation System
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine translation systems for web pages are inefficient in translating characters within images in real time, as they require long processing times, which makes them unsuitable for maintaining the visual appearance and layout of web pages during translation.
Innovation Solution
A machine translation system that connects to web data storage for HTML and image data, using dictionary data for text translation, and includes mechanisms to visualize un-visualized text and background images while un-visualizing visualized images, allowing for rapid translation without altering the web page's appearance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If characters in images are extracted and translated using conventional machine translation devices, then translation accuracy is improved, but processing time increases significantly
Solution Approach 1:
The patent applies preliminary action by pre-processing image data to extract character portions and store them in a specific format (PNG images with transparent backgrounds) before translation is needed. This allows the translation system to receive pre-prepared character images rather than processing raw images during translation, significantly reducing processing time while maintaining translation accuracy
Solution Approach 2:
The patent segments the web page content into different types: text portions (handled by conventional translation), image portions containing characters (extracted and processed separately), and other image portions. This segmentation allows each type to be processed through the most efficient method, reducing overall processing time while maintaining accuracy for character translation
2Adaptability or versatility
If characters in images are translated, then translation completeness is improved, but visual appearance and layout of the web page deteriorate
Solution Approach 1:
The patent uses copying by creating a duplicate of the original image containing characters, translating the character portion of this copy, and then replacing the original image with the translated copy. This ensures the visual appearance and layout are preserved while achieving translation completeness, as the translated character image maintains the same positioning, size, and styling as the original
Solution Approach 2:
The patent applies local quality by translating only the character portions of images while leaving other image portions unchanged. This selective translation approach maintains the overall visual appearance and layout of the web page while achieving translation completeness for the character elements
3Productivity
If simple machine translation of text data is performed, then processing speed is improved, but translation completeness deteriorates
Solution Approach 1:
The patent segments web page content into text portions and image portions, translating text data through simple machine translation for high speed, while extracting and translating character portions from images through a specialized process to ensure translation completeness. This segmentation allows the system to achieve both high processing speed and complete translation coverage
Data Source
AI summary
HTML data that contains at least a set of reference data (URL) of a visualized image containing characters, reference data (URL) of an un-visualized background image containing no characters whose display position is set to an area superimposed on the image, and un-visualized text data whose display position is set to an area superimposed on the background image is stored in a web DB, and the un-visualized background image data and the text data are visualized and the visualized image data is un-visualized in a translation process.


