Multimodal 3D Model Generation With Text and 2D Priors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional generative models struggle to generate diverse and high-resolution 3D objects due to limited training datasets, lacking the ability to effectively utilize text inputs and 2D priors for enhanced diversity and realism.

Innovation Solution

A scalable 3D generative model utilizing a Variational Auto-Encoder (VAE) trained on large datasets, incorporating text encoders and 2D/3D encoders, with Score Distillation Sampling (SDS) loss to improve generalization and diversity, and capable of generating material and physical properties of 3D objects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional generative models are trained on limited 3D object datasets, then training time and computational resources are reduced, but the diversity and quality of generated 3D objects deteriorate

Engineering Contradiction:
Improvetraining efficiencyVSAvoiddiversity of generated 3D objects
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent introduces 2D images as an intermediary medium to bridge the gap between limited 3D data and diverse generation needs. By training on abundant 2D image data and using it as a mediator to guide 3D generation through SDS loss, the system achieves high diversity without requiring extensive 3D training datasets.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions from traditional 3D-only training to a multi-dimensional approach by incorporating 2D image data and text descriptions. This dimensionality expansion allows the model to leverage diverse 2D priors and semantic information to generate varied 3D objects despite limited 3D training examples.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If conventional generative models use basic training approaches, then model complexity is reduced, but the realism and resolution of generated 3D objects deteriorate

Engineering Contradiction:
Improvemodel complexityVSAvoidresolution and realism of 3D objects
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent combines multiple data types (2D images, text descriptions) and multiple training objectives (reconstruction loss, SDS loss) to create a composite training framework. This composite approach integrates diverse information sources to achieve high-resolution and realistic 3D generation without requiring a single overly complex model architecture.

Inventive Principle:
Principle #40Composite materials

Solution Approach 2:

The patent performs preliminary encoding of text descriptions and 2D images into latent representations before the main generation process. This preliminary action prepares structured priors that guide the 3D generation, improving realism and resolution without adding complexity to the core generative model.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If conventional models cannot effectively utilize text inputs, then model simplicity is maintained, but text-to-shape generation quality and composability deteriorate

Engineering Contradiction:
Improvemodel simplicityVSAvoidtext-to-shape generation quality
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent implements a universal text encoding mechanism that serves multiple functions: guiding 3D shape generation, enabling compositional editing, and providing semantic priors for the SDS loss. This multi-functional text processing approach achieves high text-to-shape quality without requiring separate specialized modules for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12555343B23D model generation using multimodal generative AI
Publication Date: 2026.02.17 NVIDIA CORP
  • US12555343B2 patent drawing
  • US12555343B2 patent drawing
  • US12555343B2 patent drawing

AI summary

In various examples, systems and methods are disclosed relating to generating an output 3D latent representation by encoding, using a text encoder, a text prompt and encoding, using a 2D/3D encoder, a 2D image of an object or a 3D representation of the object. A 3D output is generated by applying the output 3D latent representation to a decoder. A reconstruction loss and a SDS loss are determined for the 3D output. At least one of the text encoder, the 2D/3D encoder, and the decoder is updated using the reconstruction loss and the SDS loss.