Image Text Relevancy Model for GUI Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face frustration when navigating through large amounts of text associated with images in graphical user interfaces (GUIs), as they struggle to find relevant information due to the manual annotation and linking of text to images, which is time-consuming and inefficient, especially when dealing with extensive user reviews and comments.

Innovation Solution

An image/text relevancy model is used to automatically determine correlations between image features and text, generating relevancy scores and storing them in a lookup table, allowing users to interact with images to surface relevant text without manual annotation, by using machine learning models like recurrent neural networks and convolutional neural networks to extract visual and textual features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual annotation and linking of text to images is used, then text can be associated with image components, but the process is time-consuming and inefficient

Engineering Contradiction:
Improvetext association accuracyVSAvoidannotation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables automatic text-image association through machine learning models that self-learn correlations between image features and text content, eliminating the need for manual annotation while maintaining reliable associations

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual annotation process with automated machine learning systems including recurrent neural networks and convolutional neural networks that automatically extract features and determine relevancy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If all text related to image components is displayed, then complete information is provided, but users struggle to find relevant information in large amounts of text

Engineering Contradiction:
Improveinformation completenessVSAvoidinformation retrieval ease
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system extracts only the most relevant text portions associated with selected image features by using machine learning models to compute relevancy scores, presenting a filtered subset rather than all available text

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different relevancy scoring and filtering criteria to different text portions based on their association strength with the selected image feature, highlighting the most relevant information while suppressing less relevant content

Inventive Principle:
Principle #3Local quality

3Ease of manufacture

If manual annotation of text to images is performed, then text can be linked to image components, but the process becomes inefficient with extensive user reviews and comments

Engineering Contradiction:
Improvetext linking feasibilityVSAvoidtext processing efficiency
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The machine learning system automatically processes extensive user reviews and comments by self-learning correlations between image features and text content, eliminating the need for manual annotation of large datasets

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the text processing task from manual annotation to automated feature extraction and relevancy scoring using machine learning models, fundamentally changing the approach to handling extensive text data

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11521018B1Relevant text identification based on image feature selection
Publication Date: 2022.12.06 AMAZON TECH INC
  • US11521018B1 patent drawing
  • US11521018B1 patent drawing
  • US11521018B1 patent drawing

AI summary

Techniques are generally described for predicting text relevant to image data. In various examples, the techniques may include receiving image data comprising a first portion. The first portion of the image data may correspond to a first plurality of pixels when rendered on the display. Text data comprising a first text related to the first portion of the image data may be received. A first vector representation of the first portion of the image data may be determined. In some examples, a correspondence between the first portion of the image data and the first text may be determined based at least in part on the first vector representation. A first identifier of the first portion of image data may be stored in a data structure in association with a second identifier of the first text.