Image Text Translation Preserving Display Attributes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional translation systems are unable to translate text embedded in images or videos as the text is encoded in pixel-based formats, making it difficult to extract and reinsert translated text in a manner consistent with the original display characteristics.
Innovation Solution
An image-based text translation system that automatically detects display attributes such as font size, color, and style from the source text within images, allowing translated text to be presented in a manner that matches the original, by using machine learning models and generative adversarial networks to analyze and replicate the visual characteristics of the source text.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional translation systems are used to translate text in images, then translation functionality is provided, but the translated text cannot be reinserted into the image maintaining original display characteristics
Solution Approach 1:
The system segments the image processing into distinct stages: text detection, text extraction, translation, and reinsertion. Each stage handles specific tasks independently, allowing the system to maintain original display attributes while providing translation functionality.
Solution Approach 2:
The system uses an intermediary process that analyzes the pixel-based format to extract display attributes (font size, color, style) before translation, then uses these attributes as a mediator to reinsert the translated text in a manner that matches the original appearance.
2Productivity
If text is extracted from pixel-based image formats, then translation can be performed, but the original display characteristics are lost
Solution Approach 1:
The system performs preliminary analysis of the image to detect and extract display attributes (font size, color, style) before the translation process begins. This preliminary action preserves the information needed to maintain original display characteristics throughout the translation and reinsertion process.
Solution Approach 2:
The system creates a copy of the display attributes from the original text region and applies this copied information to the translated text, ensuring the translated text visually matches the original without requiring the original pixel data to be preserved.
3Ease of operation
If translated text is reinserted into the image, then translation output is provided, but visual consistency with the original text may be compromised
Solution Approach 1:
The system applies local quality by analyzing the specific characteristics of each text region and applying appropriate display attributes (font size, color, style) tailored to that region. This ensures the translated text visually matches the original text in its specific context rather than applying uniform formatting.
Data Source
AI summary
Systems and methods are provided for translation of text in an image, and presentation of a version of the image in which the translated text is displayed a manner consistent with the original image. Text segments are automatically translated from their original source language to a target language. In order to provide presentation of the translated text in a manner that closely matches the source text, various display attributes of the source text are detected (e.g., font size, font color, font style, etc.).


