3D Auto Tagging for Multi-View Object Models on Mobile Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for creating three-dimensional (3D) image models require significant additional data and high processing resources, limiting efficiency and speed, especially in mobile and wearable devices.
Innovation Solution
The method involves analyzing spatial relationships between multiple images and video with location information to generate a multi-view interactive digital media representation (MVIDMR), using machine learning to tag and stabilize images, and allowing user interaction for immersive viewing experiences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional methods of creating 3D image models using dense depth maps or optical flow maps are used, then measurement precision and scene structure description are improved, but device complexity and processing resources increase significantly
Solution Approach 1:
The patent extracts only the essential depth information needed for 3D model construction from multiple 2D images, rather than using dense depth maps for every pixel. The system identifies and processes only relevant features and landmarks to derive 3D spatial relationships, significantly reducing data volume while maintaining measurement precision for critical scene structures.
Solution Approach 2:
The system performs preliminary processing of images to identify key features, landmarks, and spatial relationships before constructing the 3D model. By pre-processing images to extract essential geometric information and establish correspondences across views, the system reduces the computational burden during the actual 3D reconstruction phase, lowering overall processing resources required.
2Manufacturing precision
If computer generation of polygons or texture mapping over 3D mesh is used, then manufacturing precision of 3D models is improved, but productivity and processing speed decrease
Solution Approach 1:
The patent applies partial action by generating 3D models only for relevant objects and regions of interest rather than complete scene reconstruction. The system selectively processes images to create 3D representations of specific targets, maintaining manufacturing precision for those elements while significantly improving overall processing speed by avoiding unnecessary computation for entire scenes.
Solution Approach 2:
The system segments the 3D modeling process into distinct stages: feature detection, correspondence matching, spatial relationship extraction, and model construction. This segmentation allows parallel processing of different image pairs and features, improving productivity while maintaining accuracy through systematic verification at each stage.
3Measurement precision
If dense depth maps or optical flow maps are used for 3D reconstruction, then measurement precision is improved, but loss of time and processing efficiency worsen
Solution Approach 1:
The system extracts only the necessary depth information from multiple 2D images by identifying corresponding features and landmarks across views. Instead of computing dense depth maps for all pixels, the system derives depth relationships from key points and propagates this information selectively, maintaining measurement precision while dramatically reducing processing time.
Solution Approach 2:
The patent changes the parameter representation from dense per-pixel depth values to sparse feature-based depth descriptors. By representing scene geometry through landmark positions and spatial relationships rather than complete depth maps, the system maintains sufficient measurement precision for 3D reconstruction while reducing data volume and processing time proportionally.
Data Source
AI summary
A multi-view interactive digital media representation (MVIDMR) of an object can be generated from live images of an object captured from a camera. Selectable tags can be placed at locations on the object in the MVIDMR. When the selectable tags are selected, media content can be output which shows details of the object at location where the selectable tag is placed. A machine learning algorithm can be used to automatically recognize landmarks on the object in the frames of the MVIDMR and a structure from motion calculation can be used to determine 3-D positions associated with the landmarks. A 3-D skeleton associated with the object can be assembled from the 3-D positions and projected into the frames associated with the MVIDMR. The 3-D skeleton can be used to determine the selectable tag locations in the frames of the MVIDMR of the object.


