Multi-view Interactive Digital Media Representations for Visual Feature Tagging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image tagging technologies are limited by the information available in a single image, failing to provide comprehensive information about an object's features when viewed from different perspectives.

Innovation Solution

The use of multi-view interactive digital media representations (MIDMRs) that incorporate spatial information, scale information, and multiple viewpoint images to identify and tag visual features across different images, enabling feature tagging that links corresponding locations in various viewpoints.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If single image tagging is used, then the tagging process is simple, but the information completeness about object features is limited

Engineering Contradiction:
Improveinformation completenessVSAvoidtagging system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent transitions from single-image tagging to multi-view tagging by adding the dimension of multiple viewing angles. Instead of tagging features in one static image, the system captures and tags the same object from multiple perspectives (front, back, left, right views), thereby recovering information that would be invisible or ambiguous in a single view without proportionally increasing system complexity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system creates multiple copies of the object representation from different viewpoints and performs tagging on each copy. The tagged results from multiple viewpoint copies are then integrated to produce comprehensive feature information, allowing the system to overcome single-image limitations while maintaining manageable processing complexity through parallel independent tagging operations

Inventive Principle:
Principle #26Copying

2Measurement precision

If multiple viewpoint images are used for feature tagging, then the accuracy and comprehensiveness of feature identification improves, but the processing complexity increases

Engineering Contradiction:
Improvefeature identification accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the multi-view tagging problem into independent segments - each viewpoint image is processed separately through its own tagging pipeline. This segmentation allows parallel processing of multiple views without requiring complex inter-view coordination during the tagging phase, maintaining processing efficiency while improving feature identification accuracy through comprehensive multi-angle analysis

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces correspondence information as an intermediary element that links features across different viewpoint images. This intermediary structure enables the integration of tagging results from multiple views by establishing relationships between corresponding features in different perspectives, thereby improving accuracy without requiring direct complex interactions between all viewpoint processing components

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If visual feature correspondence information is created across multiple viewpoints, then the navigation and identification capability improves, but the data processing requirements increase

Engineering Contradiction:
Improvenavigation capabilityVSAvoiddata volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system extracts only the essential correspondence information needed for navigation and identification from multiple viewpoint images, rather than processing and storing all possible feature relationships. By selecting and extracting only the critical correspondence data that enables cross-view navigation, the system improves adaptability while controlling data volume through selective information extraction

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20220012495A1Visual feature tagging in multi-view interactive digital media representations
Publication Date: 2022.01.13 FUSION INC
  • US20220012495A1 patent drawing
  • US20220012495A1 patent drawing
  • US20220012495A1 patent drawing

AI summary

Provided are mechanisms and processes for visual feature tagging in multi-view interactive digital media representations (MIDMRs). In one example, a process includes receiving a visual feature tagging request that includes an MIDMR of an object to be searched, where the MIDMR includes spatial information, scale information, and different viewpoint images of the object. A visual feature in the MIDMR is identified, and visual feature correspondence information is created that links information identifying the visual feature with locations in the viewpoint images. At least one image associated with the MIDMR is transmitted in response to the feature tagging request.