Image Search Integrating Textual and Visual Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image search technologies face challenges in bridging the semantic gap between low-level visual features and high-level content, and suffer from computational overload, especially when dealing with large datasets, leading to inefficient and irrelevant search results.

Innovation Solution

The integration of textual features from surrounding web pages with visual features from image content, using a system that captures the meaning of text terms in a visual feature space and reweights visual features according to their significance, to improve image search relevance and recall.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If content-based image retrieval (CBIR) is used to detect low-level visual features, then automated image tagging becomes possible, but the system cannot detect high-level image content due to the semantic gap

Engineering Contradiction:
Improveautomated image taggingVSAvoidhigh-level image content
Core Design Contradiction:
Extent of automationVSLoss of information

Solution Approach 1:

The patent combines CBIR technology with folksonomic tagging systems to merge automated low-level feature detection with human-generated high-level semantic labels, thereby bridging the semantic gap between visual features and meaningful content description

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary classification system that translates low-level visual features into high-level semantic concepts by learning from folksonomic tags, acting as a bridge between automated detection and human understanding

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated image tagging using classification and supervised learning is implemented, then text descriptions can be generated automatically, but the system is limited by low-level visual features only

Engineering Contradiction:
Improvetext description generationVSAvoidhigh-level content understanding
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent performs preliminary action by pre-collecting and organizing folksonomic tags from multiple users before implementing the classification system, creating a rich semantic database that enhances the automated tagging capability beyond low-level features

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameter space by incorporating diverse folksonomic tags as additional training data, transforming the classification system from relying solely on low-level visual parameters to including high-level semantic parameters derived from user-generated content

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If conventional CBIR methods are used for image searching, then low-level visual features can be detected, but computational overload occurs when dealing with large datasets

Engineering Contradiction:
Improvevisual feature detectionVSAvoidsearch processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the image database into clusters based on folksonomic tags and visual features, allowing the search system to process only relevant segments rather than the entire dataset, thereby reducing computational overload while maintaining detection precision

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9075825B2System and methods of integrating visual features with textual features for image searching
Publication Date: 2015.07.07 UNIVERSITY OF KANSAS
  • US9075825B2 patent drawing
  • US9075825B2 patent drawing
  • US9075825B2 patent drawing

AI summary

The present invention includes a system and methods for image searching. Embodiments of the system integrate textual features and visual features for improved search performance. The system represents text terms in the visual feature space and develops a text-guided weighting scheme for visual features. The weighting scheme infers intention from query terms and enhances the visual features that are significant in light of such intention.