3D Auto Tagging for Multi-View Object Models Without Dense Depth Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for creating three-dimensional (3D) image models require significant data interpolation or extrapolation, dense depth maps, and high processing times, limiting efficiency and transfer rates, especially in augmented and virtual reality systems.
Innovation Solution
The method involves analyzing spatial relationships between multiple images and video with location information to generate a multi-view interactive digital media representation (MVIDMR), using machine learning to tag and stabilize images, and allowing user interaction with the viewing angle, reducing the need for dense depth maps and 3D modeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional methods of creating 3D image models using dense depth maps and optical flow maps are used, then measurement precision and reliability are improved, but device complexity and processing time increase significantly
Solution Approach 1:
The patent extracts only the essential 3D structure information needed for rendering, rather than computing complete dense depth maps. By taking out only the necessary geometric data points and using sparse sampling combined with neural network inference, the system achieves adequate depth precision without the computational burden of dense optical flow maps.
Solution Approach 2:
The patent uses neural networks to learn and copy 3D scene structure from multiple 2D images, creating a simplified 3D representation that captures essential geometry. This copied structural information is then used for rendering, replacing the need for computationally intensive dense depth map generation while maintaining sufficient measurement precision.
2Manufacturing precision
If computer generation of polygons and texture mapping over 3D meshes is used, then manufacturing precision of 3D models is improved, but productivity and processing speed decrease
Solution Approach 1:
The patent replaces the mechanical process of manual or algorithmic polygon generation and texture mapping with a neural network-based system. The neural network directly predicts 3D scene structure and generates renderable representations, substituting the traditional multi-step geometric processing pipeline with a learned end-to-end approach that maintains accuracy while dramatically improving processing speed.
Solution Approach 2:
The patent changes the fundamental parameters of 3D model representation from detailed polygon meshes with textures to simplified geometric primitives and learned scene embeddings. This parameter change allows the system to achieve sufficient manufacturing precision for AR/VR applications while reducing computational complexity and improving productivity.
3Reliability
If dense depth maps and optical flow maps are computed, then reliability of 3D reconstruction is improved, but loss of time and computational resources increase
Solution Approach 1:
The patent applies partial action by computing only the necessary portions of depth information needed for reliable 3D reconstruction. Instead of generating complete dense depth maps across all pixels, the system uses sparse sampling combined with neural network inference to predict depth values only where needed for accurate scene understanding, reducing processing time while maintaining reconstruction reliability.
Solution Approach 2:
The patent performs preliminary action by pre-training neural networks on large datasets of images and depth maps. This preliminary learning phase allows the network to encode reliable 3D reconstruction knowledge, which can then be applied rapidly to new images without requiring computationally intensive real-time depth map computation, thus reducing loss of time during actual operation.
4Ease of operation
If traditional 2D flat image formats are used, then ease of operation and device simplicity are maintained, but adaptability to immersive experiences and viewing angle flexibility are limited
Solution Approach 1:
The patent applies dimensionality change by transforming traditional 2D flat image representations into 3D scene understanding. The neural network infers three-dimensional scene structure from two-dimensional images, enabling the system to generate multiple viewing angles and immersive experiences while maintaining ease of operation with standard camera inputs. This adds the dimension of depth and spatial awareness without complicating the input process.
Data Source
AI summary
A multi-view interactive digital media representation (MVIDMR) of an object can be generated from live images of an object captured from a camera. Selectable tags can be placed at locations on the object in the MVIDMR. When the selectable tags are selected, media content can be output which shows details of the object at location where the selectable tag is placed. A machine learning algorithm can be used to automatically recognize landmarks on the object in the frames of the MVIDMR and a structure from motion calculation can be used to determine 3-D positions associated with the landmarks. A 3-D skeleton associated with the object can be assembled from the 3-D positions and projected into the frames associated with the MVIDMR. The 3-D skeleton can be used to determine the selectable tag locations in the frames of the MVIDMR of the object.


