3D Object Generation via Joint 2D and 3D Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI-based methods for generating 3D content from 2D images face challenges in quality due to limited generalization and poor geometry, particularly when dealing with out-of-distribution test images.

Innovation Solution

A method involving joint 2D and 3D training of a machine learning model using datasets with single-view images labeled with texture information and multi-view images labeled with geometry information, enabling the generation of 3D content from a single 2D image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multi-view supervision is used to construct learnable 3D priors, then 3D information can be generated from 2D images, but the model suffers from poor generalization when test images are out-of-distribution due to limited 3D training data

Engineering Contradiction:
Improvegeneralization capabilityVSAvoidamount of 3D training data
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent combines 2D and 3D training datasets into a unified joint training framework. The model simultaneously processes both 2D images (with texture labels) and 3D images (with geometry labels) during training, allowing the limited 3D data to be augmented and enhanced by the larger 2D dataset while still learning accurate 3D priors from the 3D portion.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The trained model achieves multi-functionality by being capable of generating 3D content from 2D inputs (the core function) while also maintaining robust generalization to out-of-distribution images. The joint training approach makes the model universal across both 2D and 3D data distributions, enabling it to handle diverse input scenarios effectively.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of manufacture

If 2D priors only are used to construct 3D content, then text-to-3D synthesis can be performed, but the generated content suffers from poor geometry due to ignorance of 3D information

Engineering Contradiction:
Improvesimplicity of training processVSAvoidgeometry accuracy
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent merges 2D prior knowledge (from pre-trained 2D text-to-image diffusion models) with 3D geometric information (from labeled 3D datasets) in a joint training framework. This combination allows the model to benefit from the simplicity of 2D priors while correcting their geometric deficiencies through supervised learning on 3D data.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The model undergoes parameter updates during joint training that transition it from relying solely on 2D priors to incorporating 3D geometric constraints. The training process modifies the model parameters to balance 2D texture fidelity with 3D geometric accuracy, achieving improved geometry while maintaining the ease of 2D-based synthesis.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If per-instance optimization is performed for each 3D generation task, then the model can adapt to specific inputs, but the process becomes computationally expensive and time-consuming

Engineering Contradiction:
Improveadaptability to specific inputsVSAvoidoptimization time per instance
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary joint training on comprehensive 2D and 3D datasets before deployment. This pre-training establishes robust 3D priors and geometric understanding in the model, eliminating the need for extensive per-instance optimization. The model is prepared in advance to handle specific inputs effectively without requiring time-consuming fine-tuning for each new generation task.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250111592A1Single image to realistic 3D object generation via semi-supervised 2d and 3D joint training
Publication Date: 2025.04.03 NVIDIA CORP
  • US20250111592A1 patent drawing
  • US20250111592A1 patent drawing
  • US20250111592A1 patent drawing

AI summary

Virtual reality and augmented reality bring increasing demand for 3D content creation. In an effort to automate the generation of 3D content, artificial intelligence-based processes have been developed. However, these processes are limited in terms of the quality of their output because they typically involve a model trained on limited 3D data thereby resulting in a model that does not generalize well to unseen objects, or a model trained on 2D data thereby resulting in a model that suffers from poor geometry due to ignorance of 3D information. The present disclosure jointly uses both 2D and 3D data to train a machine learning model to be able to generate 3D content from a single 2D image.