Visual Semantics Learning for Conceptual Image Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional image search methods rely on low-level visual features, failing to effectively retrieve similar images without explicit labeling or a sense of visual semantics, making it difficult to find conceptually similar images.

Innovation Solution

A system and method for learning scene embeddings via visual semantics, where machine learning is used to create representations of images based on annotations of co-existing concepts, enabling the identification of conceptually similar images and inferring image context.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional low-level visual feature methods are used for image search, then the system is simpler and faster, but the ability to retrieve conceptually similar images deteriorates

Engineering Contradiction:
Improveability to retrieve conceptually similar imagesVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces visual semantics as an intermediary layer between low-level visual features and high-level conceptual understanding. This intermediary captures semantic relationships among collocated concepts in images, enabling the system to retrieve conceptually similar images without requiring explicit labels or complex high-level processing alone

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions from low-level visual feature space to a higher-dimensional semantic embedding space. By learning scene embeddings that capture relationships among collocated concepts, the system operates in a new dimensional space where conceptually similar images are naturally clustered, improving retrieval accuracy without simply adding more features to the traditional approach

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If explicit labeling is required for image search, then the search accuracy improves, but the ease of operation deteriorates due to manual labeling requirements

Engineering Contradiction:
Improvesearch accuracyVSAvoidease of image search
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs self-service by automatically learning visual semantics from image data without requiring manual labeling. The machine learning model autonomously extracts and learns relationships among collocated concepts, enabling accurate image search while maintaining ease of operation for users who simply need to query without providing labels

Inventive Principle:
Principle #25Self-service

3Productivity

If low-level visual features are used for searching, then the processing speed is faster, but the ability to understand visual semantics deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidvisual semantics understanding
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing visual semantics embeddings for images during a training phase. This preliminary processing captures semantic relationships among collocated concepts and stores them as scene embeddings, enabling fast retrieval during actual search operations without requiring slow real-time semantic analysis

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11886488B2System and method for learning scene embeddings via visual semantics and application thereof
Publication Date: 2024.01.30 YAHOO ASSETS LLC
  • US11886488B2 patent drawing
  • US11886488B2 patent drawing
  • US11886488B2 patent drawing

AI summary

The present teaching relates to method, system, and programming for responding to an image related query. Information related to each of a plurality of images is received, wherein the information represents concepts co-existing in the image. Visual semantics for each of the plurality of images are created based on the information related thereto. Representations of scenes of the plurality of images are obtained via machine learning, based on the visual semantics of the plurality of images, wherein the representations capture concepts associated with the scenes.