Text-Guided 3D Object Stylization With Identity-Preserving Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating realistic and versatile 3D object models is time-consuming and requires significant expertise, and existing models often need substantial modifications to adapt to different styles across various media applications.

Innovation Solution

A system utilizing a generative neural network, such as a camera-conditional generative adversarial network (ccGAN), combined with a 3D stylization module, processes a 3D shape with texture and a text input to generate stylized 3D objects by disentangling identity from camera view and applying text-guided geometric and textural variations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional manual methods are used to create 3D object models, then the models can be realistic and detailed, but the process requires significant time and expertise

Engineering Contradiction:
Improvemodel qualityVSAvoidcreation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical modeling processes with an automated neural network system. The neural network automatically generates 3D object models from 2D input images, eliminating the need for manual 3D modeling operations and significantly reducing both time and expertise requirements while maintaining model quality

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service 3D model generation where the neural network autonomously performs the entire modeling process without human intervention. Users simply provide 2D images and desired style parameters, and the system automatically generates styled 3D models, making the process accessible to non-experts

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If existing 3D models are modified to adapt to different media styles, then the models can be used across various applications, but significant refinement and modifications are required

Engineering Contradiction:
Improvestyle adaptabilityVSAvoidmodification complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses parameter-based style control where different media styles are represented as adjustable parameters in the neural network. By changing these style parameters, the same base model can be adapted to different media styles (e.g., cartoon, realistic, stylized) without requiring manual geometric modifications, greatly simplifying the adaptation process

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a universal 3D modeling system that can generate models suitable for multiple media applications simultaneously. The neural network is trained on diverse style data and can produce models in various styles from a single system, eliminating the need for separate refinement processes for different media types

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If manual refinement is performed to change model styles, then the models can conform to different media requirements, but the process requires significant expertise and time

Engineering Contradiction:
Improvestyle versatilityVSAvoidoperation ease
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent replaces manual refinement operations with automated neural network processing. The system automatically adjusts model styles by processing input images through trained neural network layers that encode different style characteristics, eliminating the need for manual refinement operations and making style changes accessible to non-experts

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The neural network is pre-trained on diverse style data during an offline training phase. This preliminary action embeds various media style knowledge into the network weights, so that during actual use, style changes can be achieved simply by adjusting parameters without requiring users to perform complex refinement operations

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260112137A1Text-driven 3D object stylization using neural networks
Publication Date: 2026.04.23 NVIDIA CORP
  • US20260112137A1 patent drawing
  • US20260112137A1 patent drawing
  • US20260112137A1 patent drawing

AI summary

Generation of three-dimensional (3D) object models may be challenging for users without a sufficient skill set for content creation and may also be resource intensive. One or more style transfer networks may be combined with a generative network to generate objects based on parameters associated with a textual input. An input including a 3D mesh and texture may be provided to a trained system along with a textual input that includes parameters for object generation. Features of the input object may be identified and then tuned in accordance with the textual input to generate a modified 3D object that includes a new texture along with one or more geometric adjustments.