Triangular 3D Model Extraction from Images via Joint Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current 3D modeling techniques, such as photogrammetry, are complex and require significant manual adjustments, leading to inefficiencies and errors in creating high-quality 3D models, especially in automating the process of extracting triangular 3D models, materials, and lighting from images for use in traditional graphics engines.

Innovation Solution

A method for joint optimization of topology, materials, and lighting from multi-view image observations, which constructs triangular 3D models that can be deployed unmodified in traditional graphics engines, using a system that includes a topology construction unit, rendering unit, and image space loss unit to iteratively refine the 3D model based on image space losses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional photogrammetry pipeline is used to create 3D models, then detailed 3D models can be generated, but the process becomes complex with multiple stages and manual adjustments required

Engineering Contradiction:
Improve3D model qualityVSAvoidphotogrammetry pipeline complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple photogrammetry stages (multi-view stereo, geometric simplification, texture parameterization, material baking, and relighting) into a unified neural network framework. This integration allows joint optimization of all components simultaneously, eliminating the need for sequential manual processing while maintaining high 3D model quality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent replaces the traditional mechanical/multi-stage photogrammetry pipeline with a neural network-based system. The neural network automatically performs geometry extraction, material decomposition, and lighting estimation from input images, substituting manual artistic modeling skills and technical knowledge with automated intelligent processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If conventional photogrammetry is used, then 3D models can be created, but significant manual adjustments and multiple software tools are required

Engineering Contradiction:
Improve3D model qualityVSAvoidoperation simplicity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The neural network system performs self-service by automatically extracting geometry, decomposing materials, and estimating lighting from input images without requiring manual artistic modeling skills or technical knowledge. The system self-optimizes all parameters jointly, eliminating the need for artists to use multiple software tools and make manual adjustments.

Inventive Principle:
Principle #25Self-service

3Extent of automation

If neural network representations are used for 3D modeling, then automation is improved, but errors propagate between stages and optimization goals conflict

Engineering Contradiction:
Improve3D modeling automationVSAvoiderror propagation
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent merges all photogrammetry stages into a single neural network that jointly optimizes geometry, materials, and lighting. This unified approach eliminates error propagation between sequential stages because all components are optimized simultaneously with a single loss function, resolving conflicting optimization goals through unified gradient descent.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11967024B2Extracting triangular 3-D models, materials, and lighting from images
Publication Date: 2024.04.23 NVIDIA CORP
  • US11967024B2 patent drawing
  • US11967024B2 patent drawing
  • US11967024B2 patent drawing

AI summary

A technique is described for extracting or constructing a three-dimensional (3D) model from multiple two-dimensional (2D) images. In an embodiment, a foreground segmentation mask or depth field may be provided as an additional supervision input with each 2D image. In an embodiment, the foreground segmentation mask or depth field is automatically generated for each 2D image. The constructed 3D model comprises a triangular mesh topology, materials, and environment lighting. The constructed 3D model is represented in a format that can be directly edited and/or rendered by conventional application programs, such as digital content creation (DCC) tools. For example, the constructed 3D model may be represented as a triangular surface mesh (with arbitrary topology), a set of 2D textures representing spatially-varying material parameters, and an environment map. Furthermore, the constructed 3D model may be included in 3D scenes and interacts realistically with other objects.