3D Auto Tagging for Multi-View Object Models Without Dense Depth Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for creating three-dimensional (3D) image models require significant data interpolation or extrapolation, dense depth maps, and high processing times, limiting efficiency and transfer rates, especially in augmented and virtual reality systems.

Innovation Solution

The method involves analyzing spatial relationships between multiple images and video with location information to generate a multi-view interactive digital media representation (MVIDMR), using machine learning to tag and stabilize images, and allowing user interaction with the viewing angle, reducing the need for dense depth maps and 3D modeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional methods of creating 3D image models using dense depth maps and optical flow maps are used, then measurement precision and reliability are improved, but device complexity and processing time increase significantly

Engineering Contradiction:
Improvedepth map precisionVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential 3D structure information needed for rendering, rather than computing complete dense depth maps. By taking out only the necessary geometric data points and using sparse sampling combined with neural network inference, the system achieves adequate depth precision without the computational burden of dense optical flow maps.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses neural networks to learn and copy 3D scene structure from multiple 2D images, creating a simplified 3D representation that captures essential geometry. This copied structural information is then used for rendering, replacing the need for computationally intensive dense depth map generation while maintaining sufficient measurement precision.

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If computer generation of polygons and texture mapping over 3D meshes is used, then manufacturing precision of 3D models is improved, but productivity and processing speed decrease

Engineering Contradiction:
Improve3D model accuracyVSAvoidprocessing speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent replaces the mechanical process of manual or algorithmic polygon generation and texture mapping with a neural network-based system. The neural network directly predicts 3D scene structure and generates renderable representations, substituting the traditional multi-step geometric processing pipeline with a learned end-to-end approach that maintains accuracy while dramatically improving processing speed.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of 3D model representation from detailed polygon meshes with textures to simplified geometric primitives and learned scene embeddings. This parameter change allows the system to achieve sufficient manufacturing precision for AR/VR applications while reducing computational complexity and improving productivity.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If dense depth maps and optical flow maps are computed, then reliability of 3D reconstruction is improved, but loss of time and computational resources increase

Engineering Contradiction:
Improve3D reconstruction reliabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by computing only the necessary portions of depth information needed for reliable 3D reconstruction. Instead of generating complete dense depth maps across all pixels, the system uses sparse sampling combined with neural network inference to predict depth values only where needed for accurate scene understanding, reducing processing time while maintaining reconstruction reliability.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent performs preliminary action by pre-training neural networks on large datasets of images and depth maps. This preliminary learning phase allows the network to encode reliable 3D reconstruction knowledge, which can then be applied rapidly to new images without requiring computationally intensive real-time depth map computation, thus reducing loss of time during actual operation.

Inventive Principle:
Principle #10Preliminary action

4Ease of operation

If traditional 2D flat image formats are used, then ease of operation and device simplicity are maintained, but adaptability to immersive experiences and viewing angle flexibility are limited

Engineering Contradiction:
Improvesystem simplicityVSAvoidviewing angle adaptability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent applies dimensionality change by transforming traditional 2D flat image representations into 3D scene understanding. The neural network infers three-dimensional scene structure from two-dimensional images, enabling the system to generate multiple viewing angles and immersive experiences while maintaining ease of operation with standard camera inputs. This adds the dimension of depth and spatial awareness without complicating the input process.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12525045B2Method and apparatus for 3-D auto tagging
Publication Date: 2026.01.13 FUSION INC
  • US12525045B2 patent drawing
  • US12525045B2 patent drawing
  • US12525045B2 patent drawing

AI summary

A multi-view interactive digital media representation (MVIDMR) of an object can be generated from live images of an object captured from a camera. Selectable tags can be placed at locations on the object in the MVIDMR. When the selectable tags are selected, media content can be output which shows details of the object at location where the selectable tag is placed. A machine learning algorithm can be used to automatically recognize landmarks on the object in the frames of the MVIDMR and a structure from motion calculation can be used to determine 3-D positions associated with the landmarks. A 3-D skeleton associated with the object can be assembled from the 3-D positions and projected into the frames associated with the MVIDMR. The 3-D skeleton can be used to determine the selectable tag locations in the frames of the MVIDMR of the object.