Text-to-3D Model Generation With Multi-View Diffusion Debiasing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Neural networks generating 3D models from images often produce inaccuracies due to limited or biased training data, such as occlusions and limited angles, leading to artifacts like extra features appearing in different viewpoints.

Innovation Solution

Training neural networks using multiple viewpoints of an object, creating a dataset of image-text pairs from 3D CAD models, and employing a debiased text-to-image diffusion model to generate 3D models that account for consistent appearance across various angles, using a fine-tuned diffusion model as a critic to ensure the generated 3D models match expected viewpoints.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If training data is collected from limited angles or with occlusions, then data collection is easier and faster, but the generated 3D models contain artifacts and inaccuracies

Engineering Contradiction:
Improve3D model accuracyVSAvoiddata collection complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by generating multiple synthetic viewpoints from 3D CAD models before training the neural network. This pre-processing step creates a comprehensive training dataset that eliminates the need for complex multi-angle physical photography, resolving the contradiction between model accuracy and data collection complexity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates copies of the object from different virtual viewpoints by rendering 2D images from 3D CAD models. These synthetic copies serve as training data, eliminating the need for physical multi-angle photography while ensuring complete and accurate representation of the object from all perspectives

Inventive Principle:
Principle #26Copying

2Stability of the object's composition

If training data contains biases toward certain viewpoints, then training is simpler, but the generated models show inconsistent appearance across different angles

Engineering Contradiction:
Improveviewpoint consistencyVSAvoidtraining process complexity
Core Design Contradiction:
Stability of the object's compositionVSEase of manufacture

Solution Approach 1:

The system deliberately creates an asymmetric training approach by generating viewpoints at multiple specific angles (front, back, left, right, and intermediate angles) rather than relying on biased single-viewpoint data. This asymmetric multi-angle training ensures the neural network learns consistent object representation across all perspectives

Inventive Principle:
Principle #4Asymmetry

Solution Approach 2:

The system changes the parameter of viewpoint angle by generating training images at multiple discrete angles around the object. This parameter variation ensures the neural network learns viewpoint-invariant features, maintaining object consistency across different viewing directions while using systematically varied training data

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If a critic model is used to evaluate generated 3D models, then model accuracy improves, but computational resources and training time increase

Engineering Contradiction:
Improvemodel evaluation accuracyVSAvoidcomputational energy
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system implements feedback by using a critic neural network that evaluates generated 3D models and provides guidance for improvement. The critic compares generated models against ground truth multi-view images and returns loss signals that guide the generator to produce more accurate and consistent 3D models, resolving the contradiction between evaluation accuracy and computational cost through iterative refinement

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250259388A1Using one or more neural networks to generate three-dimensional (3D) models
Publication Date: 2025.08.14 NVIDIA CORP
  • US20250259388A1 patent drawing
  • US20250259388A1 patent drawing
  • US20250259388A1 patent drawing

AI summary

Apparatuses, systems, and techniques to use one or more neural networks to perform one or more tasks. In at least one embodiment, said one or more neural networks use one or more textual descriptions to generate one or more three-dimensional (3D) models of one or more first objects based, at least in part, on two or more images of one or more second objects from two or more viewpoints.