NeRF Facial Modeling Decoupling Latent Codes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing facial reconstruction technologies face challenges in accurately simulating three-dimensional (3D) facial expressions and movements, particularly when driven by audio or text inputs, leading to poor simulation effects.
Innovation Solution
A training method for a facial modeling model that uses a neural radiance field (NeRF) to perform 3D facial reconstruction based on facial region latent codes and action latent codes, obtained through feature encoding of sample facial images, to achieve decoupling control of facial attributes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If audio-driven facial reconstruction is used, then facial animation can be generated, but different voices cause inaccurate driving and poor simulation effect
Solution Approach 1:
The patent introduces phoneme sequences as an intermediary between audio input and facial animation generation. Instead of directly using audio signals which vary across different voices, the system converts audio to text and then to phoneme sequences, which serve as a voice-independent intermediate representation that accurately drives facial expressions.
Solution Approach 2:
The patent replaces the audio-driven mechanism with a text-driven mechanism. By substituting the direct audio-to-animation pathway with a text-based phoneme sequence pathway, the system eliminates voice-dependent variations and achieves more consistent and accurate facial animation across different speakers.
2Ease of manufacture
If text-driven facial reconstruction is used, then 2D facial images can be generated, but 3D face driving simulation effect is poor
Solution Approach 1:
The patent transitions from 2D facial image generation to 3D facial model animation by introducing spatial dimensionality. The system uses 3D facial landmarks, mesh vertices, and NeRF (Neural Radiance Fields) to create three-dimensional facial representations that can be animated from text phoneme sequences, significantly improving 3D simulation effects.
3Productivity
If traditional facial reconstruction methods are used, then facial images can be generated, but long-sequence 3D facial modeling with synchronized actions is difficult to achieve
Solution Approach 1:
The patent performs preliminary extraction and encoding of phoneme sequences from text input before generating facial animations. By pre-processing the text to obtain phoneme sequences and their corresponding temporal information, the system establishes a structured foundation that enables accurate synchronized 3D facial modeling without compromising generation speed.
Solution Approach 2:
The patent transforms the input representation from raw text to phoneme sequences with associated temporal and acoustic features. This parameter transformation includes extracting phoneme duration, pitch, and energy information, which are then used to control 3D facial landmark positions and mesh deformations, achieving reliable synchronization.
Data Source
AI summary
A training method performed by an electronic device includes obtaining a sample facial image including facial images of a same object from different perspectives at a same moment, performing feature encoding on the sample facial image through an encoder of a facial modeling model to obtain a facial action latent code corresponding to a facial action and one or more facial region latent codes each corresponding to one of one or more facial regions, performing, through a neural radiance field (NeRF) of the facial modeling model, three-dimensional (3D) facial reconstruction on the object based on the one or more facial region latent codes and the facial action latent code to obtain a 3D facial image of the object; obtaining a 3D reconstruction loss between the 3D facial image and the sample facial image, and training the facial modeling model based on the 3D reconstruction loss.


