Dynamic Text Translation Overlay for Image Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing technologies struggle to effectively present additional information, such as language translations, in a manner that is clear and readable when multiple text blocks are present in an image, often leading to cluttered displays and user confusion.

Innovation Solution

The system identifies text blocks within an image and selects a presentation context based on their arrangement and visual characteristics, using prominence, collection, or map contexts to determine the most appropriate user interface for presenting translations, ensuring clarity and readability by overlaying translations over the relevant text blocks or providing them in separate screens as needed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If translations of multiple text blocks are presented simultaneously in an image, then comprehensive translation information is provided, but the display becomes cluttered and readability decreases

Engineering Contradiction:
Improvetranslation information completenessVSAvoidreadability
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent segments the translation presentation into different modes: overlay mode for prominent text blocks and separate screen mode for less prominent text blocks. This segmentation allows comprehensive translation information to be presented while maintaining readability by not overwhelming the user with all translations simultaneously in the same visual space.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different presentation qualities to different text blocks based on their prominence. Prominent text blocks receive overlay presentations for immediate visibility, while less prominent text blocks are presented separately. This local differentiation ensures that critical information is easily readable while comprehensive information remains accessible.

Inventive Principle:
Principle #3Local quality

2Device complexity

If a single user interface is used for all text arrangements, then the system is simple to implement, but it cannot adapt to different image contexts and arrangement patterns

Engineering Contradiction:
Improveuser interface simplicityVSAvoidcontext adaptation capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic user interface selection mechanism that adapts based on text block arrangement and prominence analysis. The system automatically determines whether to use overlay mode or separate screen mode based on the detected text characteristics, enabling context-aware adaptation without requiring manual user configuration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal translation system that can handle multiple text arrangement patterns (single prominent block, multiple prominent blocks, mixed arrangements) through a single integrated framework that selects appropriate presentation modes automatically, eliminating the need for separate specialized interfaces for each scenario.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of time

If translations are overlaid directly on the image, then the user can quickly locate translated text, but the overlay may reduce the clarity of the original image text

Engineering Contradiction:
Improvetranslation location timeVSAvoidtext clarity
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent applies overlay presentation selectively only to prominent text blocks where the text is naturally large and clear in the original image. For less prominent text blocks, the system uses separate screen presentation instead. This local quality differentiation ensures that overlays are only applied where they enhance usability without compromising the clarity of the original text.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9239833B2Presenting translations of text depicted in images
Publication Date: 2016.01.19 GOOGLE LLC
  • US9239833B2 patent drawing
  • US9239833B2 patent drawing
  • US9239833B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for presenting additional information for text depicted by an image. In one aspect, a method includes receiving an image. Text depicted in the image is identified. A presentation context is selected for the image based on an arrangement of the text depicted by the image. Each presentation context corresponds to a particular arrangement of text within images. Each presentation context has a corresponding user interface for presenting additional information related to the text. The user interface for each presentation context is different from the user interface for other presentation contexts. The user interface that corresponds to the selected presentation context is identified. Additional information is presented for at least a portion of the text depicted in the image using the identified user interface. The user interface can present the additional information in an overlay over the image.