Neural Texture Synthesis for Realistic Face Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing facial capture systems and machine learning models face challenges in generating realistic digital faces due to the need for extensive and varied training data, which is time- and resource-intensive to collect.

Innovation Solution

The technique involves generating segmentation masks, texture maps, and neural textures for input geometries using neural networks, allowing for the rendering of realistic images without the need for extensive real-world training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a deep neural network is trained to perform 3D reconstruction or animation of a face using images captured under uncontrolled conditions, then the model can generalize to new data, but a large amount and variety of training data is required which is time- and resource-intensive to collect

Engineering Contradiction:
Improvegeneralization abilityVSAvoidtraining data collection time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent uses a learned identity representation from a trained identity network as a compact copy or encoding of the identity information contained in images. This learned representation serves as a surrogate for using大量actual training images, allowing the system to generalize to new identities without requiring extensive training data collection for each identity. The identity embedding captures essential identity features in a compressed form that can be reused across different imaging conditions.

Inventive Principle:
Principle #26Copying

2Productivity

If a deep neural network is trained to perform 3D reconstruction using a relatively small number of training samples, then the training process is faster and less resource-intensive, but the model's ability to generalize to new data deteriorates

Engineering Contradiction:
Improvetraining efficiencyVSAvoidgeneralization ability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary training of an identity network on a diverse set of faces to learn robust identity representations before using these representations for the actual 3D reconstruction task. This preliminary action of pre-learning identity features allows the subsequent 3D reconstruction model to focus on geometric and appearance variations without needing to relearn identity information, thereby improving generalization with fewer training samples.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an identity embedding or identity representation as an intermediary element that mediates between the input images and the 3D reconstruction process. This intermediary captures identity-specific information separately, allowing the reconstruction network to handle universal facial geometry and appearance while the identity embedding provides subject-specific customization, improving generalization without requiring extensive identity-specific training data.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If a facial capture system employs a specialized light stage and hundreds of lights to capture numerous images under multiple illumination conditions, then photorealistic faces can be captured, but the system becomes complex and resource-intensive

Engineering Contradiction:
Improveface capture realismVSAvoidcapture system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces the complex mechanical illumination system (light stage with hundreds of lights) with a learning-based approach. Instead of using physical lights to capture images under multiple illumination conditions, the system uses a neural network trained to synthesize appearance under different lighting conditions from a single or limited number of images. This substitution of mechanical illumination with computational rendering dramatically reduces device complexity while maintaining capture quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the approach from physically varying illumination parameters (using multiple lights at different positions and intensities) to computationally varying illumination parameters in the neural network. The network learns to adjust appearance parameters such as albedo, normal maps, and lighting conditions through training, allowing realistic face rendering without requiring physical changes to the capture setup.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12243140B2Synthesizing sequences of images for movement-based performance
Publication Date: 2025.03.04 DISNEY ENTERPRISES INC
  • US12243140B2 patent drawing
  • US12243140B2 patent drawing
  • US12243140B2 patent drawing

AI summary

A technique for rendering an input geometry includes generating a first segmentation mask for a first input geometry and a first set of texture maps associated with one or more portions of the first input geometry. The technique also includes generating, via one or more neural networks, a first set of neural textures for the one or more portions of the first input geometry. The technique further includes rendering a first image corresponding to the first input geometry based on the first segmentation mask, the first set of texture maps, and the first set of neural textures.