Text-Driven 3D Mask Generation With Differentiable Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Producing 3D masks, such as 3D face masks, is difficult and requires sophisticated software, and generating them from a textual description is computationally prohibitive due to the large number of possible masks and high computational requirements.

Innovation Solution

A generative neural network (NN) is trained to generate 2D images from input text, and an initial 3D mask is updated based on the output of the NN compared with a 2D image rendered with the initial 3D mask, using differentiable rendering and backpropagation to iteratively improve the 3D mask.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If sophisticated 3D modeling software is used to produce 3D masks, then manufacturing precision is improved, but device complexity increases and ease of manufacture deteriorates

Engineering Contradiction:
Improve3D mask qualityVSAvoidsoftware accessibility
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent replaces traditional mechanical 3D modeling software with a neural network-based generative system. The neural network learns to generate 3D masks by training on 2D image data and automatically produces 3D outputs, eliminating the need for users to operate complex 3D modeling tools while maintaining high manufacturing precision

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system creates 3D masks by learning from and copying patterns in 2D image data. The neural network extracts features from 2D images and generates corresponding 3D representations, allowing high-quality 3D mask generation without requiring users to manually create models in sophisticated software

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If the number of possible 3D masks is increased to provide flexibility, then adaptability is improved, but computational requirements increase prohibitively

Engineering Contradiction:
Improvemask varietyVSAvoidcomputational resources
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary action by pre-training the neural network on a diverse dataset of 2D images and their corresponding 3D representations. This pre-training enables the network to generate a wide variety of 3D masks adaptively without requiring computational resources to evaluate all possible mask configurations in real-time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The neural network uses parameter changes in the latent space to generate diverse 3D masks efficiently. By manipulating a small set of latent variables rather than exploring the full space of possible 3D masks, the system achieves high adaptability with computationally feasible resource requirements

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If a generative neural network is trained to generate 2D images from text, then ease of operation is improved, but device complexity increases

Engineering Contradiction:
Improveuser interface simplicityVSAvoidneural network system
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent uses text as an intermediary between the user and the complex neural network system. Users provide simple text descriptions of desired masks, and the neural network translates these into 2D images and subsequently 3D masks, shielding users from the underlying system complexity while maintaining ease of operation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The neural network system performs multiple functions: it generates 2D images from text, converts 2D images to 3D masks, and iteratively refines outputs. This multi-functionality is encapsulated in a single system that users interact with through a simple text interface, improving ease of operation despite the inherent complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Manufacturing precision

If iterative improvement of 3D mask is implemented using backpropagation, then manufacturing precision is improved, but loss of time increases

Engineering Contradiction:
Improve3D mask accuracyVSAvoidgeneration time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system uses periodic action by implementing iterative refinement cycles where the 3D mask is generated, rendered to 2D, compared with the target image, and updated through backpropagation. This periodic cycle of generation and refinement improves manufacturing precision while keeping time loss manageable through efficient convergence

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The iterative improvement process uses feedback by comparing the rendered 2D representation of the 3D mask with the target 2D image and using the difference to guide updates via backpropagation. This feedback mechanism efficiently directs the refinement process toward the optimal solution, improving precision without excessive time loss

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250259405A13D mask generation
Publication Date: 2025.08.14 SNAP INC
  • US20250259405A1 patent drawing
  • US20250259405A1 patent drawing
  • US20250259405A1 patent drawing

AI summary

A three-dimensional (3D) trainable mask is geometrically adjusted based on two-dimensional (2D) images of a target image. The 3D trainable mask is applied to a face of a user within an image. Example methods include rendering a trainable three-dimensional (3D) mask to generate a two-dimensional (2D) image, adding noise to the 2D image to generate a 2D image with added noise, inputting the 2D image with added noise and input text into a trained neural network to generate a predicted noise of the 2D image with added noise, determining a loss between the 2D image with added noise and the predicted noise, and updating the trainable 3D mask based on the loss.