3D Auto Tagging With Skeleton-Based Landmark Projection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for creating three-dimensional (3D) image models require significant data interpolation or extrapolation, dense depth maps, and high processing times, limiting efficiency and transfer rates, especially in augmented and virtual reality systems.
Innovation Solution
The method involves analyzing spatial relationships between multiple images and video with location information to generate a multi-view interactive digital media representation (MVIDMR), using machine learning algorithms to tag and render 2D or 3D skeletons, allowing users to interactively control the viewing experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional methods are used to produce 3D models using dense depth maps or optical flow maps, then 3D model accuracy is improved, but data requirements and processing complexity increase significantly
Solution Approach 1:
The patent extracts only the essential information needed for 3D reconstruction from the image sequences, specifically identifying and using only the necessary image frames and their corresponding camera poses, rather than processing all available data. This selective extraction reduces the data quantity required while maintaining sufficient accuracy for the application.
Solution Approach 2:
The system uses a partial set of images and camera poses that is sufficient for generating the 3D model, rather than requiring complete dense coverage. By using only the minimal necessary data (partial action), the system achieves acceptable 3D reconstruction accuracy without the computational burden of processing all available image data.
2Ease of manufacture
If computer generation of polygons or texture mapping is used to produce 3D models, then 3D model creation is achieved, but processing time and computational resources increase
Solution Approach 1:
Instead of generating 3D models through complex computer synthesis (polygons and texture mapping), the patent directly copies and utilizes existing 2D image data from the image sequences. By working with actual captured images rather than synthesized representations, the system avoids the time-consuming computational processes of 3D modeling while still achieving the desired visual representation.
Solution Approach 2:
The patent replaces the mechanical/computational process of generating 3D models through polygon construction and texture mapping with a more direct approach using image processing and camera pose data. This substitution of the computational mechanism reduces processing time and resource requirements while maintaining the ability to create interactive 3D representations.
3Measurement precision
If dense depth maps or optical flow maps are used for 3D reconstruction, then spatial accuracy is improved, but device complexity and processing requirements increase
Solution Approach 1:
The patent segments the 3D reconstruction task into discrete, manageable components: identifying key frames, determining camera poses for each frame, and using only the essential image data. This segmentation of the processing workflow reduces overall system complexity while maintaining spatial accuracy through careful selection and processing of individual image components.
Solution Approach 2:
The system changes the parameters of data processing by working with a reduced set of essential parameters (key frame indices, camera poses, and selected images) rather than processing complete dense depth maps or optical flow fields. This parameter transformation simplifies the processing requirements while preserving the necessary spatial information for accurate 3D reconstruction.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A multi-view interactive digital media representation (MVIDMR) of an object can be generated from live images of an object captured from a camera. Selectable tags can be placed at locations on the object in the MVIDMR. When the selectable tags are selected, media content can be output which shows details of the object at location where the selectable tag is placed. A machine learning algorithm can be used to automatically recognize landmarks on the object in the frames of the MVIDMR and a structure from motion calculation can be used to determine 3-D positions associated with the landmarks. A 3-D skeleton associated with the object can be assembled from the 3-D positions and projected into the frames associated with the MVIDMR. The 3-D skeleton can be used to determine the selectable tag locations in the frames of the MVIDMR of the object.