Multi-view Interactive Digital Media Representation for 3D Tagging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional digital media formats, such as 2D flat images and videos, are limited in providing an immersive and interactive experience, especially when trying to capture and present three-dimensional data, as they require significant additional data for interpolation or extrapolation, leading to high processing demands and inefficient data transfer.
Innovation Solution
The development of multi-view interactive digital media representations (MVIDMR) that analyze spatial relationships between images and video with location information to create an immersive, interactive, and efficient format for capturing and presenting 3D data, allowing for the generation and tagging of media content associated with specific locations within the media.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional 2D flat images and videos are used to capture and present 3D data, then the format is simple and widely compatible, but significant additional data is required for interpolation or extrapolation, leading to high processing demands and inefficient data transfer
Solution Approach 1:
The patent transitions from 2D flat images to multi-view 3D representations by capturing images from multiple angles and depths. This dimensional expansion enables true 3D data representation while using efficient encoding methods that avoid the need for extensive interpolation or extrapolation, thereby reducing processing demands despite the increased adaptability for 3D content
Solution Approach 2:
The patent divides the 3D scene into multiple discrete views captured from different positions and angles. Each view is processed and encoded independently, allowing for efficient data management and transfer. This segmentation approach reduces the overall processing complexity compared to attempting to represent the entire 3D scene as a single interpolated model
2Manufacturing precision
If dense depth maps or optical flow maps are used to describe scene structure, then 3D reconstruction accuracy is improved, but the amount of additional data required increases significantly
Solution Approach 1:
Instead of creating dense depth maps or optical flow maps that require storing data for every pixel, the patent captures multiple discrete 2D images from different viewpoints. These images serve as copies of the scene from various angles, allowing 3D reconstruction through geometric relationships between the copies rather than through dense per-pixel depth information, thereby maintaining accuracy while reducing data volume
3Manufacturing precision
If computer generation of polygons or texture mapping is used to create 3D models, then 3D model quality is improved, but processing times and resource requirements increase
Solution Approach 1:
The patent performs view synthesis and 3D reconstruction operations in advance during the content creation phase, generating the multi-view representation before playback. This preliminary processing allows the system to store pre-computed 3D data that can be efficiently transmitted and displayed without requiring real-time polygon generation or texture mapping during playback, thereby improving processing speed while maintaining model quality
Data Source
AI summary
Various embodiments of the present invention relate generally to systems and methods for analyzing and manipulating images and video. In particular, a multi-view interactive digital media representation (MVIDMR) of an object can be generated from live images of an object captured from a camera. After the MVIDMR of the object is generated, a tag can be placed at a location on the object in the MVIDMR. The locations of the tag in the frames of the MVIDMR can vary from frame to frame as the view of the object changes. When the tag is selected, media content can be output which shows details of the object at location where the tag is placed. In one embodiment, the object can be car and tags can be used to link to media content showing details of the car at the locations where the tags are placed.


