Cross-Modality Visual Search via Learned Appearance Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multi-camera video monitoring systems face challenges in searching for individuals or objects across different camera modalities and lighting conditions, as appearance changes significantly between RGB and infrared modes, and varying lighting conditions affect object recognition, leading to difficulties in mapping and false positives.

Innovation Solution

A system that builds appearance models based on video analysis, learns from user input to identify beneficial mappings, and prunes likely false positives, using confidence factors to generalize searches across modalities and lighting conditions, enabling efficient visual search across multiple camera views and conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If appearance-based video search is performed across multiple camera modalities (RGB and infrared), then the search coverage and detection capability are improved, but the mapping complexity and false positive rate increase due to appearance changes between modalities

Engineering Contradiction:
Improvesearch coverage across modalitiesVSAvoidmapping complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary mapping system that learns correspondence relationships between RGB and infrared modalities. This mapping acts as a mediator that translates appearance features from one modality to another, enabling cross-modality search without direct complex comparisons. The mapping is learned through training data and can be pruned to remove incorrect mappings, thus reducing false positives while maintaining search coverage.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If appearance-based video search is performed across different lighting conditions, then the search robustness is improved, but the mapping accuracy deteriorates due to drastic appearance changes under varying lighting

Engineering Contradiction:
Improvesearch robustnessVSAvoidmapping accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent addresses lighting condition variations by learning mappings that are robust to parameter changes in appearance. The system trains on diverse lighting conditions and learns to map appearances across different lighting scenarios. The mapping can be pruned based on performance metrics to remove mappings that produce false positives under specific lighting conditions, thus maintaining accuracy while preserving robustness.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If all possible mappings between modalities and lighting conditions are built, then the search completeness is improved, but the computational complexity and processing time increase

Engineering Contradiction:
Improvesearch completenessVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts and retains only the beneficial mappings from the set of all possible mappings. Through training and evaluation, the system identifies which mappings produce accurate results and which produce false positives. Incorrect or redundant mappings are pruned and removed, keeping only the essential mappings needed for effective cross-modality and cross-lighting-condition search. This reduces computational complexity while maintaining search completeness.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11762906B2Modality mapping for visual search
Publication Date: 2023.09.19 OBJECTVIDEO LABS LLC
  • US11762906B2 patent drawing
  • US11762906B2 patent drawing
  • US11762906B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for integrating a monitoring system with one or more air quality sensors. The method includes obtaining a request that includes an indication of an appearance of an object, selecting a first set of images in an initial modality based on the indication of the appearance of the object, determining an additional set of images in a different modality based on mappings between the first set of images and the additional set of images; and providing the first set of images and the additional set of images in response to the request.