Dynamic Text Translation Interface for Image Clutter Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing technologies fail to effectively present translations of text in images in a manner that is clear and user-friendly, especially when multiple text blocks are present, leading to cluttered displays and confusion for users.

Innovation Solution

The system identifies text in images, determines the arrangement and visual characteristics of the text, and selects a presentation context to dynamically choose a user interface that presents translations in an overlay or separate screens based on prominence, collection, or map contexts, ensuring readability and clarity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If translations of multiple text blocks are presented simultaneously, then complete information is provided, but display clutter increases and readability decreases

Engineering Contradiction:
Improvetranslation information completenessVSAvoidreadability
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent segments the presentation of translations by organizing text blocks into groups based on their spatial arrangement and visual characteristics in the image. Instead of presenting all translations at once, the system divides them into manageable groups that can be displayed together without creating excessive clutter, thereby maintaining both information completeness and readability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different presentation qualities to different text blocks based on their local characteristics. Text blocks are analyzed individually for their visual properties and spatial relationships, and translations are presented with varying levels of detail and formatting according to each text block's specific context, optimizing readability for each local region while maintaining overall information completeness.

Inventive Principle:
Principle #3Local quality

2Device complexity

If a single user interface is used for all image types, then device complexity is reduced, but adaptability to different text arrangements is insufficient

Engineering Contradiction:
Improveuser interface complexityVSAvoidcontext adaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic user interface that automatically adapts its presentation style based on the analyzed characteristics of the input image and its text blocks. The system evaluates spatial arrangement, visual properties, and contextual relationships to dynamically select appropriate presentation modes, thereby achieving high adaptability without requiring multiple static interfaces or complex manual configuration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes key parameters of the user interface presentation based on analyzed image characteristics. By adjusting parameters such as translation display format, text block grouping, overlay positioning, and presentation timing according to the detected spatial and visual properties, the system achieves context adaptability while maintaining a single unified interface framework.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If text arrangement analysis is performed to select presentation contexts, then translation presentation accuracy is improved, but processing time increases

Engineering Contradiction:
Improvepresentation context accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary analysis of text block arrangements and visual characteristics during the image processing stage, before final translation presentation is generated. By pre-analyzing spatial relationships, grouping text blocks into potential contexts, and preparing presentation strategies in advance, the system achieves high presentation context accuracy while minimizing additional processing time during the actual translation generation phase.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10726212B2Presenting translations of text depicted in images
Publication Date: 2020.07.28 GOOGLE LLC
  • US10726212B2 patent drawing
  • US10726212B2 patent drawing
  • US10726212B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for presenting additional information for text depicted by an image. In one aspect, a method includes receiving an image. Text depicted in the image is identified. The identified text can be in one or more text blocks. A prominence presentation context is selected for the image based on the relative prominence of the one or more text blocks. Each prominence presentation context corresponds to a relative prominence of each text block in which text is presented within images. Each prominence presentation context has a corresponding user interface for presenting additional information related to the identified text depicted in the image. A user interface is identified that corresponds to the selected prominence presentation context. Additional information is presented for at least a portion of the text depicted in the image using the identified user interface.