Probabilistic Semantic Model for Multi-Modal Image Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multimedia retrieval systems face challenges in bridging the semantic gap between low-level image features and high-level semantic concepts, particularly in multi-modal scenarios where imagery is often accompanied by collateral information, leading to insufficient retrieval accuracy and limited querying modalities.

Innovation Solution

A probabilistic semantic modeling approach that exploits synergy between image and text modalities through a hidden layer, using an Expectation-Maximization based iterative learning procedure to determine conditional probabilities and perform image-to-text and text-to-image retrieval in a Bayesian framework, allowing for automated annotation and improved retrieval performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If image features are used solely to find similar images, then the retrieval process is simple, but retrieval accuracy is insufficient due to the semantic gap between low-level features and high-level semantic concepts

Engineering Contradiction:
Improveretrieval accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a hidden layer of semantic concepts as an intermediary between low-level image features and high-level semantic descriptions. This hidden layer acts as a mediator that bridges the semantic gap, allowing the system to learn and exploit correlations between visual features and textual annotations without directly mapping between them, thereby improving retrieval accuracy while managing system complexity through probabilistic modeling.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multi-modal approaches are used to exploit collateral information, then retrieval accuracy improves, but the system complexity increases

Engineering Contradiction:
Improveretrieval accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple information modalities (image features, textual annotations, and hidden semantic concepts) into a unified probabilistic framework. By combining these diverse data sources and learning their joint distribution through the EM algorithm, the system exploits the redundancy and complementary information across modalities to achieve improved retrieval accuracy while integrating them into a single coherent model.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The hidden layer of semantic concepts serves multiple functions simultaneously: it acts as a bridge between image features and textual annotations, enables automatic image annotation, supports multi-modal image retrieval, and captures underlying semantic structures. This multi-functionality allows the system to address multiple retrieval challenges with a single unified mechanism, improving accuracy without proportionally increasing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If automated annotation is performed using probabilistic semantic models, then retrieval performance improves, but computational requirements increase

Engineering Contradiction:
Improveretrieval performanceVSAvoidcomputational energy
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary learning of the probabilistic semantic model and hidden semantic concepts during an offline training phase using the EM algorithm. By pre-computing and storing the learned correlations between image features, semantic concepts, and textual annotations, the system prepares the model in advance so that online retrieval and annotation operations can be performed more efficiently with reduced computational energy requirements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The EM algorithm incorporates feedback mechanisms where the model parameters are iteratively refined based on the observed correlations between image features and textual annotations. The algorithm uses the current parameter estimates to compute expected values, then updates parameters to maximize likelihood, creating a feedback loop that converges to optimal values. This feedback-driven learning improves retrieval performance by capturing accurate semantic relationships while efficiently utilizing computational resources through iterative optimization.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10614366B1System and method for multimedia ranking and multi-modal image retrieval using probabilistic semantic models and expectation-maximization (EM) learning
Publication Date: 2020.04.07 THE RES FOUNDATION FOR THE STATE UNIV OF NEW YORK
  • US10614366B1 patent drawing
  • US10614366B1 patent drawing
  • US10614366B1 patent drawing

AI summary

Systems and Methods for multi-modal or multimedia image retrieval are provided. Automatic image annotation is achieved based on a probabilistic semantic model in which visual features and textual words are connected via a hidden layer comprising the semantic concepts to be discovered, to explicitly exploit the synergy between the two modalities. The association of visual features and textual words is determined in a Bayesian framework to provide confidence of the association. A hidden concept layer which connects the visual feature(s) and the words is discovered by fitting a generative model to the training image and annotation words. An Expectation-Maximization (EM) based iterative learning procedure determines the conditional probabilities of the visual features and the textual words given a hidden concept class. Based on the discovered hidden concept layer and the corresponding conditional probabilities, the image annotation and the text-to-image retrieval are performed using the Bayesian framework.