Structured Image Knowledge Extraction for Search Precision

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image search techniques are prone to errors and inaccuracies due to the lack of structured relationships between tags and images, leading to inefficient searches for complex queries, as they rely on unstructured tagging and fail to define relationships between image elements.

Innovation Solution

A digital medium environment is configured to automatically extract and model structured knowledge from images using machine learning, generating a descriptive summarization that associates text features with image features, enabling precise image searches and automatic caption generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional unstructured tagging techniques are used for image search, then the implementation is simple and fast, but the search accuracy and precision are poor

Engineering Contradiction:
Improvesearch accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the image into multiple regions and extracts features from each region separately, then combines them with text features to create a structured representation. This segmentation approach improves search accuracy by capturing spatial relationships and local details while maintaining manageable system complexity through modular feature extraction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from unstructured text tags to a multi-dimensional structured representation that includes spatial coordinates, region boundaries, and hierarchical relationships. This dimensional expansion enables precise spatial queries and relationship-based searches while the systematic framework keeps implementation complexity controlled.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If conventional tag-based search is used, then the process is quick and easy to implement, but it requires navigating through many images to find the desired one

Engineering Contradiction:
Improvesearch efficiencyVSAvoidtime to locate image
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces manual browsing and keyword matching with automated machine learning models that compute structured representations and perform semantic matching. This substitution dramatically improves search efficiency by directly identifying relevant images based on complex queries without requiring users to manually navigate through large datasets.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If structured knowledge modeling is implemented, then the search precision and semantic understanding are improved, but the processing complexity and computational requirements increase

Engineering Contradiction:
Improvesemantic search precisionVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary extraction of structured features from images during an offline preprocessing stage, creating reusable structured representations that can be quickly queried online. This preliminary action shifts computational complexity from the search phase to the indexing phase, improving online search precision while managing processing complexity through batch processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically extracts structured features, generates region proposals, and builds knowledge representations without requiring manual annotation or intervention. This self-service approach improves semantic search precision through automated structured modeling while reducing operational complexity by eliminating manual processing steps.

Inventive Principle:
Principle #25Self-service

4Reliability

If manual tagging is used for image organization, then the tagging process allows human judgment and context, but it is time-consuming and labor-intensive

Engineering Contradiction:
Improvetagging accuracyVSAvoidtagging speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces manual human tagging with automated machine learning models that extract structured features and generate captions. This substitution maintains high reliability through learned semantic understanding while dramatically improving productivity by processing images at machine speed without human intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-annotation by automatically extracting structured representations and generating text descriptions from image content. This self-service capability achieves tagging accuracy comparable to or exceeding manual methods while eliminating the time-consuming and labor-intensive nature of manual tagging through automated feature extraction and caption generation.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10460033B2Structured knowledge modeling, extraction and localization from images
Publication Date: 2019.10.29 ADOBE INC
  • US10460033B2 patent drawing
  • US10460033B2 patent drawing
  • US10460033B2 patent drawing

AI summary

Techniques and systems are described to model and extract knowledge from images. A digital medium environment is configured to learn and use a model to compute a descriptive summarization of an input image automatically and without user intervention. Training data is obtained to train a model using machine learning in order to generate a structured image representation that serves as the descriptive summarization of an input image. The images and associated text are processed to extract structured semantic knowledge from the text, which is then associated with the images. The structured semantic knowledge is processed along with corresponding images to train a model using machine learning such that the model describes a relationship between text features within the structured semantic knowledge. Once the model is learned, the model is usable to process input images to generate a structured image representation of the image.