Deep Relightable Appearance Model for Real-Time Face Avatars
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating relightable three-dimensional computer models of human faces are limited to single lighting conditions, making them unsuitable for dynamic renderings under novel expressions and lighting conditions, which hinders their adoption in applications like game and film production where consistency between character and environment is desirable.
Innovation Solution
A learning-based method using a Deep Relightable Appearance Model (DRAM) that leverages neural networks to generate expression-dependent and view-dependent textures, allowing for real-time rendering under novel viewpoints and lighting conditions, including challenging natural illumination and near-field lighting, through a teacher-student network framework and conditional variational auto-encoders.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Illumination intensity
If traditional 3D rendering models are used, then rendering speed is fast, but lighting realism is poor and limited to single lighting conditions
Solution Approach 1:
The patent pre-computes and stores lighting responses in the form of texture maps (view-dependent texture maps and expression-dependent texture maps) during an offline training phase. These pre-computed textures encode how the face reflects light from various directions and under various expressions, allowing real-time relighting without intensive processing during runtime.
Solution Approach 2:
The patent creates simplified representations (texture maps) that copy and encode the complex lighting interactions. Instead of performing full physical lighting simulations during rendering, the system uses pre-computed texture maps that replicate lighting effects, enabling fast real-time relighting while maintaining realism.
2Illumination intensity
If intensive processing is used to achieve realism, then lighting quality is improved, but real-time application capability is lost
Solution Approach 1:
The system performs intensive processing offline to train neural networks and generate texture maps. Once trained, the model can rapidly generate relit images during real-time applications. The heavy computational work is shifted from runtime to training time, enabling both high quality and real-time performance.
Solution Approach 2:
The patent replaces traditional mechanical/lighting simulation systems with a learning-based neural network system. The neural network learns complex lighting patterns during training and can rapidly infer relit images during runtime, substituting intensive physical simulations with efficient learned predictions.
3Productivity
If learning-based relighting approaches are applied on 2D images or static scenes, then processing complexity is reduced, but applicability to dynamic renderings under novel expressions and lighting conditions is limited
Solution Approach 1:
The patent creates a universal model that handles multiple functions: it works with novel lighting conditions, novel expressions, and novel viewpoints simultaneously. The view-dependent and expression-dependent texture maps are designed to generalize across different conditions, making the system adaptable to diverse scenarios beyond the training data.
Solution Approach 2:
The patent transitions from 2D image processing to 3D-aware processing by incorporating view-dependent texture maps and expression-dependent texture maps. This multi-dimensional approach (combining view, expression, and lighting dimensions) enables the system to handle dynamic renderings under novel conditions while maintaining processing efficiency.
Data Source
AI summary
A method for providing a relightable avatar of a subject to a virtual reality application is provided. The method includes retrieving multiple images including multiple views of a subject and generating an expression-dependent texture map and a view-dependent texture map for the subject, based on the images. The method also includes generating, based on the expression-dependent texture map and the view-dependent texture map, a view of the subject illuminated by a light source selected from an environment in an immersive reality application, and providing the view of the subject to an immersive reality application running in a client device. A non-transitory, computer-readable medium storing instructions and a system that executes the instructions to perform the above method are also provided.


