Multi-View Interactive Digital Media Representation Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating multi-view interactive digital media representations (MIDMRs) are inefficient due to the need for dense data descriptions, high processing times, and resource requirements, particularly in producing 3D models from 2D images, which limits their applicability in augmented and virtual reality systems.
Innovation Solution
A method for automatically generating MIDMRs using convex or concave motion capture, combining general object and specific feature MIDMRs, and embedding specific views within general views, allowing for interactive viewing with selectable tags, and utilizing user templates and neural networks for image processing and enhancement algorithms to reduce data redundancy and enhance user experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If dense data descriptions (depth maps, optical flow maps) are used to generate 3D models, then manufacturing precision is improved, but device complexity and processing resources increase significantly
Solution Approach 1:
The patent segments the 3D model generation process into two distinct phases: an offline training phase where a neural network learns from dense data (depth maps, optical flow), and an online inference phase where the trained network generates 3D models from sparse 2D images. This segmentation allows the system to benefit from dense data training without requiring dense data during actual operation, thus reducing device complexity while maintaining manufacturing precision.
Solution Approach 2:
The patent performs preliminary action by training the neural network offline using dense data descriptions before deployment. The pre-trained network encapsulates the complex processing requirements, allowing the runtime system to operate with reduced complexity while still achieving high 3D model accuracy through the pre-learned patterns and features.
2Manufacturing precision
If traditional 3D model generation methods are used, then manufacturing precision is improved, but productivity decreases due to high processing times
Solution Approach 1:
The patent replaces traditional mechanical 3D reconstruction algorithms (which involve complex geometric computations, mesh generation, and texture mapping) with a neural network-based system. This substitution leverages parallel processing capabilities of neural networks, significantly improving processing speed while maintaining 3D model quality through the network's learned understanding of scene geometry and appearance.
Solution Approach 2:
The patent changes the fundamental parameters of the 3D generation process by transitioning from deterministic algorithmic parameters to probabilistic neural network parameters. The network learns optimal parameters for 3D reconstruction from training data, enabling faster processing while maintaining or improving model quality through data-driven parameter optimization rather than fixed algorithmic rules.
3Adaptability or versatility
If complete 360-degree capture is performed, then adaptability is improved, but loss of time increases due to extended capture duration
Solution Approach 1:
The patent applies partial action by capturing images from a limited set of viewpoints rather than completing a full 360-degree capture. The neural network compensates for the missing views by inferring unseen portions from the captured images, allowing the system to achieve comprehensive adaptability with reduced capture time by performing only partial capture actions.
Solution Approach 2:
The patent uses the neural network to create virtual copies of viewpoints that were not physically captured. By learning the scene structure from limited captured images, the network generates synthetic images representing uncaptured angles, effectively copying the appearance and geometry of unseen portions without requiring physical capture, thus maintaining adaptability while reducing capture time.
Data Source
AI summary
Various embodiments describe systems and processes for capturing and generating multi-view interactive digital media representations (MIDMRs). In one aspect, a method for automatically generating a MIDMR comprises obtaining a first MIDMR and a second MIDMR. The first MIDMR includes a convex or concave motion capture using a recording device and is a general object MIDMR. The second MIDMR is a specific feature MIDMR. The first and second MIDMRs may be obtained using different capture motions. A third MIDMR is generated from the first and second MIDMRs, and is a combined embedded MIDMR. The combined embedded MIDMR may comprise the second MIDMR being embedded in the first MIDMR, forming an embedded second MIDMR. The third MIDMR may include a general view in which the first MIDMR is displayed for interactive viewing by a user on a user device. The embedded second MIDMR may not be viewable in the general view.


