Image Text Translation Preserving Display Attributes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional translation systems are unable to translate text embedded in images or videos as the text is encoded in pixel-based formats, making it difficult to extract and reinsert translated text in a manner consistent with the original display characteristics.

Innovation Solution

An image-based text translation system that automatically detects display attributes such as font size, color, and style from the source text within images, allowing translated text to be presented in a manner that matches the original, by using machine learning models and generative adversarial networks to analyze and replicate the visual characteristics of the source text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional translation systems are used to translate text in images, then translation functionality is provided, but the translated text cannot be reinserted into the image maintaining original display characteristics

Engineering Contradiction:
Improvetranslation capabilityVSAvoiddisplay attribute consistency
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system segments the image processing into distinct stages: text detection, text extraction, translation, and reinsertion. Each stage handles specific tasks independently, allowing the system to maintain original display attributes while providing translation functionality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses an intermediary process that analyzes the pixel-based format to extract display attributes (font size, color, style) before translation, then uses these attributes as a mediator to reinsert the translated text in a manner that matches the original appearance.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If text is extracted from pixel-based image formats, then translation can be performed, but the original display characteristics are lost

Engineering Contradiction:
Improvetranslation processingVSAvoiddisplay attributes
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system performs preliminary analysis of the image to detect and extract display attributes (font size, color, style) before the translation process begins. This preliminary action preserves the information needed to maintain original display characteristics throughout the translation and reinsertion process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a copy of the display attributes from the original text region and applies this copied information to the translated text, ensuring the translated text visually matches the original without requiring the original pixel data to be preserved.

Inventive Principle:
Principle #26Copying

3Ease of operation

If translated text is reinserted into the image, then translation output is provided, but visual consistency with the original text may be compromised

Engineering Contradiction:
Improvetranslation output generationVSAvoidvisual consistency
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system applies local quality by analyzing the specific characteristics of each text region and applying appropriate display attributes (font size, color, style) tailored to that region. This ensures the translated text visually matches the original text in its specific context rather than applying uniform formatting.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240111968A1Image-based text translation and presentation
Publication Date: 2024.04.04 AMAZON TECH INC
  • US20240111968A1 patent drawing
  • US20240111968A1 patent drawing
  • US20240111968A1 patent drawing

AI summary

Systems and methods are provided for translation of text in an image, and presentation of a version of the image in which the translated text is displayed a manner consistent with the original image. Text segments are automatically translated from their original source language to a target language. In order to provide presentation of the translated text in a manner that closely matches the source text, various display attributes of the source text are detected (e.g., font size, font color, font style, etc.).