Image Text Lookup Using Visual Context and Re-Ranking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional search engines fail to provide relevant search results when users input text from images due to a lack of contextual information, often returning unrelated results.
Innovation Solution
Utilizing machine learning models to identify and analyze contextual information within images, such as text, objects, and application metadata, to enhance search results by incorporating depth and location information, and locally or remotely re-ranking search results based on this data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional search engines are used for text lookup, then the search process is simple and fast, but the search results are irrelevant and inaccurate
Solution Approach 1:
The system performs preliminary actions by extracting contextual information from the image before the search query is submitted. The machine learning model analyzes the image to identify objects, text, and their spatial relationships, preparing enriched search queries that include both the text content and contextual metadata about the image elements.
Solution Approach 2:
The patent introduces an intermediary layer between the user's simple text query and the search engine. This intermediary is the contextual enrichment system that translates image-based queries into enhanced search parameters, including extracted text, object identifiers, spatial coordinates, and visual embeddings, which then guide the search results.
2Measurement precision
If contextual information is extracted from images using machine learning, then search result accuracy improves, but processing time and computational resources increase
Solution Approach 1:
The system applies partial action by selectively extracting only the most relevant contextual information from the image rather than analyzing every possible element. The machine learning model identifies and extracts key elements such as prominent text, main objects, and their spatial relationships, while filtering out less relevant details to reduce processing overhead.
Solution Approach 2:
The patent utilizes parameter changes by transforming the search query from simple text to enriched parameters that include contextual metadata. The system changes the representation of the search query to include multiple dimensions such as text content, object type, spatial position, and visual features, enabling more accurate search results while managing processing time through efficient parameter transformation.
Data Source
AI summary
The subject technology provides for contextual text lookup for images. When a request is received by an electronic device to perform a lookup or search for text in an image that is displayed at the electronic device, the electronic device may obtain one or more search results, based on the text itself and based on contextual information derived, by the electronic device, from the image. In one or more implementations, application information associated with an application that displays the image may also be used as contextual metadata for enhancing the results of the search for the text from the image.


