Text-Driven 3D Object Stylization for Rapid Style Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating realistic and adaptable 3D object models across different content types is challenging due to the need for significant skill and time, and existing methods often require major modifications to change styles.

Innovation Solution

A system utilizing a generative neural network, such as a camera-conditional generative adversarial network (ccGAN), combined with a 3D stylization module, processes a 3D shape with texture and text input to generate stylized 3D objects, incorporating features from models like 3DStyleNet and Text2Mesh, guided by a language-vision model like CLIP for embedding and loss tuning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional 3D model creation methods are used, then model realism and quality can be achieved, but significant skill and time are required

Engineering Contradiction:
Improvemodel qualityVSAvoidcreation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent replaces traditional manual 3D modeling mechanical processes with a neural network-based automated system. The neural network learns from training data to generate 3D models directly from text inputs, substituting the manual skill-based creation process with an automated intelligent system that maintains high model quality while dramatically reducing creation time

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary action by pre-training the neural network on large datasets of 3D models and their corresponding text descriptions. This pre-training enables the network to already possess the knowledge and patterns needed for high-quality model generation, so when actual 3D models need to be created, the network can rapidly generate them without requiring time-consuming manual refinement

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If existing 3D models are used across different content types, then model reusability is maintained, but major modifications are required to change styles

Engineering Contradiction:
Improvestyle adaptabilityVSAvoidmodification complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamics by making the 3D model generation process adaptive and flexible through the neural network. The system can dynamically adjust model characteristics based on text inputs, allowing the same underlying technology to generate models in different styles without requiring static, hard-coded modification procedures. The neural network continuously adapts to generate appropriate styles based on the input description

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The neural network system serves multiple functions: it can generate 3D models from scratch, modify existing models, adapt to different styles, and respond to various text inputs. This universal system replaces the need for separate tools and processes for different model creation and modification tasks, allowing a single system to handle diverse content type requirements without major modifications

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12417602B2Text-driven 3D object stylization using neural networks
Publication Date: 2025.09.16 NVIDIA CORP
  • US12417602B2 patent drawing
  • US12417602B2 patent drawing
  • US12417602B2 patent drawing

AI summary

Generation of three-dimensional (3D) object models may be challenging for users without a sufficient skill set for content creation and may also be resource intensive. One or more style transfer networks may be combined with a generative network to generate objects based on parameters associated with a textual input. An input including a 3D mesh and texture may be provided to a trained system along with a textual input that includes parameters for object generation. Features of the input object may be identified and then tuned in accordance with the textual input to generate a modified 3D object that includes a new texture along with one or more geometric adjustments.