Multi-view Interactive Digital Media Representations for Visual Feature Tagging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image tagging technologies are limited by the information available in a single image, failing to provide comprehensive information about an object's features when viewed from different perspectives.
Innovation Solution
The use of multi-view interactive digital media representations (MIDMRs) that incorporate spatial information, scale information, and multiple viewpoint images to identify and tag visual features across different images, enabling feature tagging that links corresponding locations in various viewpoints.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If single image tagging is used, then the tagging process is simple, but the information completeness about object features is limited
Solution Approach 1:
The patent transitions from single-image tagging to multi-view tagging by adding the dimension of multiple viewing angles. Instead of tagging features in one static image, the system captures and tags the same object from multiple perspectives (front, back, left, right views), thereby recovering information that would be invisible or ambiguous in a single view without proportionally increasing system complexity
Solution Approach 2:
The system creates multiple copies of the object representation from different viewpoints and performs tagging on each copy. The tagged results from multiple viewpoint copies are then integrated to produce comprehensive feature information, allowing the system to overcome single-image limitations while maintaining manageable processing complexity through parallel independent tagging operations
2Measurement precision
If multiple viewpoint images are used for feature tagging, then the accuracy and comprehensiveness of feature identification improves, but the processing complexity increases
Solution Approach 1:
The patent divides the multi-view tagging problem into independent segments - each viewpoint image is processed separately through its own tagging pipeline. This segmentation allows parallel processing of multiple views without requiring complex inter-view coordination during the tagging phase, maintaining processing efficiency while improving feature identification accuracy through comprehensive multi-angle analysis
Solution Approach 2:
The system introduces correspondence information as an intermediary element that links features across different viewpoint images. This intermediary structure enables the integration of tagging results from multiple views by establishing relationships between corresponding features in different perspectives, thereby improving accuracy without requiring direct complex interactions between all viewpoint processing components
3Adaptability or versatility
If visual feature correspondence information is created across multiple viewpoints, then the navigation and identification capability improves, but the data processing requirements increase
Solution Approach 1:
The system extracts only the essential correspondence information needed for navigation and identification from multiple viewpoint images, rather than processing and storing all possible feature relationships. By selecting and extracting only the critical correspondence data that enables cross-view navigation, the system improves adaptability while controlling data volume through selective information extraction
Data Source
AI summary
Provided are mechanisms and processes for visual feature tagging in multi-view interactive digital media representations (MIDMRs). In one example, a process includes receiving a visual feature tagging request that includes an MIDMR of an object to be searched, where the MIDMR includes spatial information, scale information, and different viewpoint images of the object. A visual feature in the MIDMR is identified, and visual feature correspondence information is created that links information identifying the visual feature with locations in the viewpoint images. At least one image associated with the MIDMR is transmitted in response to the feature tagging request.


