Image Text Lookup Using Visual Context and Re-Ranking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional search engines fail to provide relevant search results when users input text from images due to a lack of contextual information, often returning unrelated results.

Innovation Solution

Utilizing machine learning models to identify and analyze contextual information within images, such as text, objects, and application metadata, to enhance search results by incorporating depth and location information, and locally or remotely re-ranking search results based on this data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional search engines are used for text lookup, then the search process is simple and fast, but the search results are irrelevant and inaccurate

Engineering Contradiction:
Improvesearch result relevanceVSAvoidsearch system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by extracting contextual information from the image before the search query is submitted. The machine learning model analyzes the image to identify objects, text, and their spatial relationships, preparing enriched search queries that include both the text content and contextual metadata about the image elements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary layer between the user's simple text query and the search engine. This intermediary is the contextual enrichment system that translates image-based queries into enhanced search parameters, including extracted text, object identifiers, spatial coordinates, and visual embeddings, which then guide the search results.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If contextual information is extracted from images using machine learning, then search result accuracy improves, but processing time and computational resources increase

Engineering Contradiction:
Improvesearch result accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies partial action by selectively extracting only the most relevant contextual information from the image rather than analyzing every possible element. The machine learning model identifies and extracts key elements such as prominent text, main objects, and their spatial relationships, while filtering out less relevant details to reduce processing overhead.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent utilizes parameter changes by transforming the search query from simple text to enriched parameters that include contextual metadata. The system changes the representation of the search query to include multiple dimensions such as text content, object type, spatial position, and visual features, enabling more accurate search results while managing processing time through efficient parameter transformation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250371073A1Contextual text lookup for images
Publication Date: 2025.12.04 APPLE INC
  • US20250371073A1 patent drawing
  • US20250371073A1 patent drawing
  • US20250371073A1 patent drawing

AI summary

The subject technology provides for contextual text lookup for images. When a request is received by an electronic device to perform a lookup or search for text in an image that is displayed at the electronic device, the electronic device may obtain one or more search results, based on the text itself and based on contextual information derived, by the electronic device, from the image. In one or more implementations, application information associated with an application that displays the image may also be used as contextual metadata for enhancing the results of the search for the text from the image.