Statistical Image Indexing via Hidden Markov Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing content-based image retrieval systems face challenges in automatically assigning comprehensive textual descriptions to images due to the difficulty in recognizing a large number of objects, relying on semantically meaningful segmentation, which is often unavailable in image databases, limiting their scalability and ability to handle a large number of categories and images.
Innovation Solution
A system and method that uses statistical models, specifically two-dimensional multi-resolution hidden Markov models (2-D MHMMs), to automatically index images by profiling categories, extracting feature vectors, and comparing them to determine statistically significant index terms for linguistic descriptions, allowing for large-scale training and indexing without relying on segmentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If semantically meaningful segmentation is used for image indexing, then indexing accuracy is improved, but device complexity and scalability deteriorate due to unavailability of segmentation in most image databases
Solution Approach 1:
The patent extracts and removes the segmentation requirement from the image indexing process. Instead of requiring segmented images as input, the system directly processes raw images by extracting visual features and comparing them against statistical models, thereby eliminating the complexity barrier while maintaining indexing capability
Solution Approach 2:
The statistical modeling approach serves multiple functions: it enables image categorization, linguistic indexing, and semantic understanding without requiring segmentation. This universal method works across diverse image types and domains, making the system scalable and adaptable to different applications
2Adaptability or versatility
If a large number of categories are trained simultaneously, then system versatility is improved, but processing time and computational resources worsen
Solution Approach 1:
The patent performs preliminary actions by pre-training statistical models for multiple categories beforehand. These models are built in advance and stored for later use, allowing the system to quickly process new images by comparing them against the pre-existing model library without performing extensive training at query time
Solution Approach 2:
The system creates simplified copies of image categories in the form of statistical models. Instead of storing and processing large numbers of actual training images during query operations, the system uses compact statistical model representations that capture the essential characteristics of each category, enabling fast comparison and classification
3Loss of information
If comprehensive textual descriptions are assigned to images, then information completeness is improved, but processing complexity worsens due to difficulty in recognizing large numbers of objects
Solution Approach 1:
The patent introduces statistical models as intermediary representations between raw images and linguistic descriptions. These models serve as mediators that capture visual patterns and characteristics, enabling the system to generate comprehensive textual descriptions without directly solving the complex problem of recognizing and categorizing every object in the image
Solution Approach 2:
The system replaces the mechanical approach of explicit object recognition and categorization with a statistical pattern matching approach. Instead of identifying and labeling individual objects, the system compares overall image patterns against statistical models, substituting complex object recognition mechanics with probabilistic pattern matching that achieves comprehensive description with less complexity
Data Source
AI summary
The present invention provides a statistical modeling approach to automatic linguistic indexing of photographic images. The invention uses categorized images to train a dictionary of hundreds of statistical models each representing a concept. Images of any given concept are regarded as instances of a stochastic process that characterizes the concept. To measure the extent of association between an image and a textual description associated with a predefined concept, the likelihood of the occurrence of the image based on the characterizing stochastic process is computed. A high likelihood indicates a strong association between the textual description and the image. The invention utilizes two-dimensional multi-resolution hidden Markov models that demonstrate accuracy and high potential in linguistic indexing of photographic images.


