3D Face Animation from Speech Using Segmented Mesh Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio-driven facial animation technologies struggle to produce highly realistic and scalable three-dimensional facial animations that accurately capture co-articulation and are not limited to person-specific models.

Innovation Solution

The development of a machine learning-based approach that utilizes a categorical latent space for facial expressions, disentangling audio-correlated and uncorrelated information, and employing a cross-modality loss function to ensure accurate upper face reconstruction and lip motion synchronization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing audio-driven facial animation approaches are used, then implementation is simpler, but the animation quality becomes uncanny or static and fails to produce accurate co-articulation

Engineering Contradiction:
Improvefacial animation accuracyVSAvoidmodel complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the facial animation generation into distinct components: audio-driven lower face mesh generation for lip-sync, and expression-driven upper face mesh generation for facial expressions. This segmentation allows each component to be optimized independently, achieving high accuracy without requiring a single complex unified model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary expression embedding that bridges audio input and facial animation output. This embedding captures expression information from video frames and mediates between audio-driven lip motion and the final synthesized facial animation, enabling accurate co-articulation without direct complex audio-to-animation mapping.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If person-specific models are used, then animation accuracy for specific individuals improves, but scalability to new identities is limited

Engineering Contradiction:
Improvescalability to new identitiesVSAvoidanimation accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent employs a universal expression embedding that captures general facial expression patterns applicable across different identities. This embedding is trained on diverse video data and can be applied to any target identity, enabling the system to generate accurate facial animations for new identities without requiring identity-specific training data.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent copies expression patterns from video frames into the expression embedding, which then serves as a reusable template for generating facial animations. This copying mechanism allows expression information to be transferred across different identities, achieving both accuracy and scalability.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If audio-driven approaches are used, then lip motion synchronization improves, but upper face animation becomes uncanny or static

Engineering Contradiction:
Improvelip motion accuracyVSAvoidupper face animation naturalness
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent segments facial animation into audio-driven lower face (for accurate lip-sync) and expression-driven upper face (for natural expressions). This segmentation allows each region to be controlled by the most appropriate input modality, ensuring both lip motion accuracy and upper face naturalness simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The expression embedding acts as an intermediary that provides natural upper face animation independent of audio input. This intermediary ensures the upper face remains expressive and natural while the lower face maintains accurate lip-sync to audio, resolving the trade-off between the two regions.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250131631A1Three-dimensional face animation from speech
Publication Date: 2025.04.24 META PLATFORMS TECHNOLOGIES LLC
  • US20250131631A1 patent drawing
  • US20250131631A1 patent drawing
  • US20250131631A1 patent drawing

AI summary

A method for training a three-dimensional model face animation model from speech, is provided. The method includes determining a first correlation value for a facial feature based on an audio waveform from a first subject, generating a first mesh for a lower portion of a human face, based on the facial feature and the first correlation value, updating the first correlation value when a difference between the first mesh and a ground truth image of the first subject is greater than a pre-selected threshold, and providing a three-dimensional model of the human face animated by speech to an immersive reality application accessed by a client device based on the difference between the first mesh and the ground truth image of the first subject. A non-transitory, computer-readable medium storing instructions to cause a system to perform the above method, and the system, are also provided.