Statistical Image Indexing via Hidden Markov Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing content-based image retrieval systems face challenges in automatically assigning comprehensive textual descriptions to images due to the difficulty in recognizing a large number of objects, relying on semantically meaningful segmentation, which is often unavailable in image databases, limiting their scalability and ability to handle a large number of categories and images.

Innovation Solution

A system and method that uses statistical models, specifically two-dimensional multi-resolution hidden Markov models (2-D MHMMs), to automatically index images by profiling categories, extracting feature vectors, and comparing them to determine statistically significant index terms for linguistic descriptions, allowing for large-scale training and indexing without relying on segmentation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If semantically meaningful segmentation is used for image indexing, then indexing accuracy is improved, but device complexity and scalability deteriorate due to unavailability of segmentation in most image databases

Engineering Contradiction:
Improveindexing accuracyVSAvoidsegmentation requirement
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the segmentation requirement from the image indexing process. Instead of requiring segmented images as input, the system directly processes raw images by extracting visual features and comparing them against statistical models, thereby eliminating the complexity barrier while maintaining indexing capability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The statistical modeling approach serves multiple functions: it enables image categorization, linguistic indexing, and semantic understanding without requiring segmentation. This universal method works across diverse image types and domains, making the system scalable and adaptable to different applications

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If a large number of categories are trained simultaneously, then system versatility is improved, but processing time and computational resources worsen

Engineering Contradiction:
Improvenumber of categoriesVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-training statistical models for multiple categories beforehand. These models are built in advance and stored for later use, allowing the system to quickly process new images by comparing them against the pre-existing model library without performing extensive training at query time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates simplified copies of image categories in the form of statistical models. Instead of storing and processing large numbers of actual training images during query operations, the system uses compact statistical model representations that capture the essential characteristics of each category, enabling fast comparison and classification

Inventive Principle:
Principle #26Copying

3Loss of information

If comprehensive textual descriptions are assigned to images, then information completeness is improved, but processing complexity worsens due to difficulty in recognizing large numbers of objects

Engineering Contradiction:
Improvetextual description completenessVSAvoidobject recognition complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces statistical models as intermediary representations between raw images and linguistic descriptions. These models serve as mediators that capture visual patterns and characteristics, enabling the system to generate comprehensive textual descriptions without directly solving the complex problem of recognizing and categorizing every object in the image

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces the mechanical approach of explicit object recognition and categorization with a statistical pattern matching approach. Instead of identifying and labeling individual objects, the system compares overall image patterns against statistical models, substituting complex object recognition mechanics with probabilistic pattern matching that achieves comprehensive description with less complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS7394947B2System and method for automatic linguistic indexing of images by a statistical modeling approach
Publication Date: 2008.07.01 THE PENN STATE RES FOUND INC
  • US7394947B2 patent drawing
  • US7394947B2 patent drawing
  • US7394947B2 patent drawing

AI summary

The present invention provides a statistical modeling approach to automatic linguistic indexing of photographic images. The invention uses categorized images to train a dictionary of hundreds of statistical models each representing a concept. Images of any given concept are regarded as instances of a stochastic process that characterizes the concept. To measure the extent of association between an image and a textual description associated with a predefined concept, the likelihood of the occurrence of the image based on the characterizing stochastic process is computed. A high likelihood indicates a strong association between the textual description and the image. The invention utilizes two-dimensional multi-resolution hidden Markov models that demonstrate accuracy and high potential in linguistic indexing of photographic images.