Multi-view Interactive Digital Media Representation for Object Tagging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional digital media formats, such as 2D images and videos, are inadequate for capturing and indexing visual data effectively, limiting user interaction and search capabilities, especially with the increasing quantity of visual data being captured.
Innovation Solution
The development of multi-view interactive digital media representations (MIDMRs) that analyze spatial relationships between images and location information to create immersive, interactive 3D-like experiences without requiring actual 3D modeling, allowing for efficient indexing and user-controlled viewing angles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional 2D flat images and videos are used to capture visual data, then the implementation is simple and memory requirements are low, but the user interaction capability and search effectiveness are limited
Solution Approach 1:
The patent transitions from traditional 2D flat images to multi-view interactive digital media representations that incorporate spatial relationships and multiple viewing angles. This dimensional enhancement allows users to interact with visual data from different perspectives while maintaining computational efficiency through selective processing of view data.
Solution Approach 2:
The patent segments visual data into multiple discrete views that can be independently processed and selectively rendered. Each view represents a specific angle or perspective, allowing the system to load and process only the necessary portions based on user interaction, thereby improving ease of operation without requiring complete 3D models.
2Adaptability or versatility
If actual 3D modeling approaches are used to create immersive experiences, then user interaction and viewing flexibility are improved, but memory and CPU requirements increase significantly
Solution Approach 1:
The patent creates lightweight digital representations that capture essential spatial and geometric information without requiring complete 3D models. These simplified copies retain sufficient detail for interactive viewing from multiple angles while consuming significantly less computational resources than full 3D modeling approaches.
Solution Approach 2:
The patent processes and stores only the necessary portions of spatial information required for multi-view interaction, rather than creating complete 3D models. This partial processing approach provides adequate viewing flexibility while avoiding the excessive computational burden of full 3D reconstruction and rendering.
3Measurement precision
If comprehensive search and indexing mechanisms are developed for increasing visual data, then search effectiveness is improved, but the system complexity and processing requirements increase
Solution Approach 1:
The patent performs preliminary organization of visual data into structured multi-view representations with embedded spatial relationships during the capture and storage phase. This pre-structuring enables more effective search and indexing operations later, as the data is already arranged in a format that facilitates efficient querying and retrieval without requiring complex processing during search operations.
Data Source
AI summary
Various embodiments of the present disclosure relate generally to systems and methods for automatic tagging of objects on a multi-view interactive digital media representation of a dynamic entity. According to particular embodiments, the spatial relationship between multiple images and video is analyzed together with location information data, for purposes of creating a representation referred to herein as a multi-view interactive digital media representation for presentation on a device. Multi-view interactive digital media representations correspond to multi-view interactive digital media representations of the dynamic objects in backgrounds. A first multi-view interactive digital media representation of a dynamic object is obtained. Next, the dynamic object is tagged. Then, a second multi-view interactive digital media representation of the dynamic object is generated. Finally, the dynamic object in the second multi-view interactive digital media representation is automatically identified and tagged.


