NeRF Facial Modeling Decoupling Latent Codes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing facial reconstruction technologies face challenges in accurately simulating three-dimensional (3D) facial expressions and movements, particularly when driven by audio or text inputs, leading to poor simulation effects.

Innovation Solution

A training method for a facial modeling model that uses a neural radiance field (NeRF) to perform 3D facial reconstruction based on facial region latent codes and action latent codes, obtained through feature encoding of sample facial images, to achieve decoupling control of facial attributes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If audio-driven facial reconstruction is used, then facial animation can be generated, but different voices cause inaccurate driving and poor simulation effect

Engineering Contradiction:
Improvefacial animation generationVSAvoiddriving accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces phoneme sequences as an intermediary between audio input and facial animation generation. Instead of directly using audio signals which vary across different voices, the system converts audio to text and then to phoneme sequences, which serve as a voice-independent intermediate representation that accurately drives facial expressions.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the audio-driven mechanism with a text-driven mechanism. By substituting the direct audio-to-animation pathway with a text-based phoneme sequence pathway, the system eliminates voice-dependent variations and achieves more consistent and accurate facial animation across different speakers.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of manufacture

If text-driven facial reconstruction is used, then 2D facial images can be generated, but 3D face driving simulation effect is poor

Engineering Contradiction:
Improvefacial image generationVSAvoid3D simulation effect
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent transitions from 2D facial image generation to 3D facial model animation by introducing spatial dimensionality. The system uses 3D facial landmarks, mesh vertices, and NeRF (Neural Radiance Fields) to create three-dimensional facial representations that can be animated from text phoneme sequences, significantly improving 3D simulation effects.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If traditional facial reconstruction methods are used, then facial images can be generated, but long-sequence 3D facial modeling with synchronized actions is difficult to achieve

Engineering Contradiction:
Improvefacial image generation speedVSAvoid3D facial synchronization
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary extraction and encoding of phoneme sequences from text input before generating facial animations. By pre-processing the text to obtain phoneme sequences and their corresponding temporal information, the system establishes a structured foundation that enables accurate synchronized 3D facial modeling without compromising generation speed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the input representation from raw text to phoneme sequences with associated temporal and acoustic features. This parameter transformation includes extracting phoneme duration, pitch, and energy information, which are then used to control 3D facial landmark positions and mesh deformations, achieving reliable synchronization.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250200878A1Training method and apparatus for facial modeling model, modeling method and apparatus, electronic device, storage medium, and program product
Publication Date: 2025.06.19 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US20250200878A1 patent drawing
  • US20250200878A1 patent drawing
  • US20250200878A1 patent drawing

AI summary

A training method performed by an electronic device includes obtaining a sample facial image including facial images of a same object from different perspectives at a same moment, performing feature encoding on the sample facial image through an encoder of a facial modeling model to obtain a facial action latent code corresponding to a facial action and one or more facial region latent codes each corresponding to one of one or more facial regions, performing, through a neural radiance field (NeRF) of the facial modeling model, three-dimensional (3D) facial reconstruction on the object based on the one or more facial region latent codes and the facial action latent code to obtain a 3D facial image of the object; obtaining a 3D reconstruction loss between the 3D facial image and the sample facial image, and training the facial modeling model based on the 3D reconstruction loss.