Search Enriched Metadata for Document Image Text Highlighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional document editor/viewer applications fail to highlight search terms within document images, even if the image content matches the search term, limiting the ability to locate images of interest within electronic documents.
Innovation Solution
The technique involves creating search enriched metadata for documents that includes images, which associates selected image portions with text search terms and location coordinates, allowing the system to visually identify and highlight matching terms within images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional document search algorithms are used, then text search functionality is provided, but image content with matching search terms cannot be located
Solution Approach 1:
The patent merges text search and image search functionalities into a unified search system. The search algorithm now processes both document text and image content (via OCR) to locate matching terms, combining previously separate search domains into one integrated capability that addresses the limitation of conventional text-only search.
Solution Approach 2:
The patent introduces an intermediary layer (OCR processing and metadata generation) that converts image content into searchable text representations. This intermediary transformation enables the search system to process images without requiring direct image pattern matching, bridging the gap between conventional text search and image content search.
2Measurement precision
If image metadata search is performed, then images of potential interest can be located, but search terms within image content cannot be highlighted
Solution Approach 1:
The patent segments the search results into different types: text results and image results. For image results, it further segments by identifying specific regions of interest within images where matching terms are located. This segmentation enables precise location of images while also providing highlighting of specific term locations within those images.
Solution Approach 2:
The patent adds a new dimension to search results by incorporating spatial location information within images. Instead of only providing image-level results, the system now annotates the specific positions (coordinates) of matching terms within image boundaries, enabling both location accuracy and visual highlighting functionality.
3Device complexity
If document images are ignored in search, then search algorithm complexity is reduced, but relevant image content is missed
Solution Approach 1:
The patent performs preliminary OCR processing and metadata generation for images before the actual search operation. By pre-processing images to extract text content and associate it with spatial locations, the system prepares search-ready data structures that simplify the main search algorithm while enabling comprehensive search capability including image content.
Data Source
AI summary
A technique for facilitating identification of a matching search term in one or more images includes selecting at least a portion of an image and creating search enriched metadata for a document that includes the image. The search enriched metadata includes a text portion that provides one or more search terms that are associated with the selected portion of the image and a location portion that provides a location of the selected portion of the image.


