Neural Texture Rendering for Realistic Human Avatars

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional methods for capturing realistic three-dimensional human body appearances require complex and expensive setups, making it difficult to digitize and transfer models, and struggle with view-dependent effects and pose-dependent deformations.

Innovation Solution

A system using a simple statistical human body model fitted to a training video to capture body shape statistics and pose information, employing a convolutional rendering network with neural latent texture optimized for keyframes to produce realistic, view-dependent effects, allowing for the rendering of multiple identities with a single set of rendering network parameters and enabling virtual try-on and animation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional complex capture setups are used to capture realistic three-dimensional human body appearances, then the realism and quality of the captured appearance is improved, but the device complexity and cost increase significantly

Engineering Contradiction:
Improveappearance realismVSAvoidcapture setup complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical capture setups with a neural network-based rendering system. Instead of using multiple cameras and complex hardware to capture appearance from different views, the system uses a single video input and processes it through neural networks (rendering network and texture synthesis network) to generate realistic three-dimensional avatars with view-dependent effects, thereby substituting mechanical complexity with computational processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the parameter representation from direct multi-view capture to latent texture encoding. The neural texture encodes appearance information in a compressed latent space with view-dependent parameters, allowing the system to generate realistic appearances by adjusting rendering parameters rather than capturing all views physically, thus reducing device complexity while maintaining appearance quality

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If complex capture setups are used to obtain realistic human body models, then the appearance quality is improved, but the ease of digitization and model transfer deteriorates

Engineering Contradiction:
Improveappearance qualityVSAvoidmodel digitization ease
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent replaces complex physical capture infrastructure with a software-based neural rendering pipeline that processes standard video input. This substitution enables easy digitization using common video devices and facilitates model transfer through the compact neural texture representation, which can be efficiently stored and transferred without requiring complex calibration data or proprietary formats

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The rendering network is trained on multiple identities simultaneously, leading to a strong decoupling of the neural texture and the rendering network. As a result, a system may capture and render multiple identities with only one set of rendering network parameters in addition to an identity specific neural texture map.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Manufacturing precision

If neural texture is optimized with all video frames, then the texture quality is improved, but the training time and computational resources increase

Engineering Contradiction:
Improvetexture qualityVSAvoidtraining time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the training process into two distinct phases: first, the neural texture is optimized using only keyframes (selected important frames) to establish the base appearance; second, the rendering network is trained on all frames to learn pose-dependent effects. This segmentation allows the texture optimization to focus on quality without being overwhelmed by the full video duration, reducing training time while maintaining texture fidelity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses a partial action approach by selecting only keyframes for neural texture optimization rather than processing every frame. The keyframes are sufficient to capture the essential appearance characteristics, providing adequate texture quality without the excessive computational cost of processing the entire video sequence, thus achieving a balance between quality and efficiency

Inventive Principle:
Principle #16Partial or excessive action

4Adaptability or versatility

If a single rendering network is used for multiple identities, then the system versatility is improved, but the texture rendering quality for each identity may deteriorate

Engineering Contradiction:
Improvemulti-identity renderingVSAvoidtexture rendering quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the appearance representation into two independent components: a shared rendering network that handles pose-dependent effects and lighting, and identity-specific neural textures that encode appearance characteristics. This segmentation allows the rendering network to be universal across identities while each identity maintains its own optimized texture, preventing quality deterioration despite multi-identity versatility

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by allowing different parts of the system to have different optimization levels. The neural texture for each identity is independently optimized to capture that specific identity's appearance characteristics with high fidelity, while the rendering network uses standardized parameters for all identities. This localized optimization ensures each identity receives appropriate attention for quality while the system maintains overall versatility

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11651540B2Learning a realistic and animatable full body human avatar from monocular video
Publication Date: 2023.05.16 META PLATFORMS TECHNOLOGIES LLC
  • US11651540B2 patent drawing
  • US11651540B2 patent drawing
  • US11651540B2 patent drawing

AI summary

In one embodiment, a method includes adjusting parameters of a three-dimensional geometry corresponding to a first person to make the three-dimensional geometry represent a desired pose for the first person, accessing a neural texture encoding an appearance of the first person, generating a first rendered neural texture based on a mapping between (1) a portion of the three-dimensional geometry that is visible from a viewing direction and (2) the neural texture, generating a second rendered neural texture by processing the first rendered neural texture using a first neural network, determining normal information associated with the portion of the three-dimensional geometry that is visible from the viewing direction, and generating a rendered image for the first person in the desired pose by processing the second rendered neural texture and the normal information using a second neural network.