3D Face Animation from Speech Using Segmented Mesh Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio-driven facial animation technologies struggle to produce highly realistic and scalable three-dimensional facial animations that accurately capture co-articulation and are not limited to person-specific models.
Innovation Solution
The development of a machine learning-based approach that utilizes a categorical latent space for facial expressions, disentangling audio-correlated and uncorrelated information, and employing a cross-modality loss function to ensure accurate upper face reconstruction and lip motion synchronization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If existing audio-driven facial animation approaches are used, then implementation is simpler, but the animation quality becomes uncanny or static and fails to produce accurate co-articulation
Solution Approach 1:
The patent segments the facial animation generation into distinct components: audio-driven lower face mesh generation for lip-sync, and expression-driven upper face mesh generation for facial expressions. This segmentation allows each component to be optimized independently, achieving high accuracy without requiring a single complex unified model.
Solution Approach 2:
The patent introduces an intermediary expression embedding that bridges audio input and facial animation output. This embedding captures expression information from video frames and mediates between audio-driven lip motion and the final synthesized facial animation, enabling accurate co-articulation without direct complex audio-to-animation mapping.
2Adaptability or versatility
If person-specific models are used, then animation accuracy for specific individuals improves, but scalability to new identities is limited
Solution Approach 1:
The patent employs a universal expression embedding that captures general facial expression patterns applicable across different identities. This embedding is trained on diverse video data and can be applied to any target identity, enabling the system to generate accurate facial animations for new identities without requiring identity-specific training data.
Solution Approach 2:
The patent copies expression patterns from video frames into the expression embedding, which then serves as a reusable template for generating facial animations. This copying mechanism allows expression information to be transferred across different identities, achieving both accuracy and scalability.
3Manufacturing precision
If audio-driven approaches are used, then lip motion synchronization improves, but upper face animation becomes uncanny or static
Solution Approach 1:
The patent segments facial animation into audio-driven lower face (for accurate lip-sync) and expression-driven upper face (for natural expressions). This segmentation allows each region to be controlled by the most appropriate input modality, ensuring both lip motion accuracy and upper face naturalness simultaneously.
Solution Approach 2:
The expression embedding acts as an intermediary that provides natural upper face animation independent of audio input. This intermediary ensures the upper face remains expressive and natural while the lower face maintains accurate lip-sync to audio, resolving the trade-off between the two regions.
Data Source
AI summary
A method for training a three-dimensional model face animation model from speech, is provided. The method includes determining a first correlation value for a facial feature based on an audio waveform from a first subject, generating a first mesh for a lower portion of a human face, based on the facial feature and the first correlation value, updating the first correlation value when a difference between the first mesh and a ground truth image of the first subject is greater than a pre-selected threshold, and providing a three-dimensional model of the human face animated by speech to an immersive reality application accessed by a client device based on the difference between the first mesh and the ground truth image of the first subject. A non-transitory, computer-readable medium storing instructions to cause a system to perform the above method, and the system, are also provided.


