Facial Animation Generation Using Blendshape Control for Semantic Lip Sync
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing animation generation methods for lip sync animations are vertex-driven, leading to animations that do not accurately represent the semantics of speech data and lack diversity in facial expressions.
Innovation Solution
An animation generation method that utilizes blendshape parameters, where control conditions are determined based on speech feature subsequences and facial expression style features, using a preset decoder to generate target blendshape parameters, and deforming an object model to create a facial animation that aligns with speech data while allowing for diverse facial expressions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If vertex-driven animation generation is used, then the animation generation process is simple, but the animation accuracy in representing speech semantics deteriorates
Solution Approach 1:
The patent changes the driving parameters from vertex coordinates to blendshape parameters. The blendshape parameters directly control facial muscle movements and are semantically correlated with speech content, thereby improving animation accuracy while maintaining a manageable generation process through parameter transformation.
Solution Approach 2:
The patent introduces a semantic correlation module as an intermediary between speech data and animation generation. This module extracts semantic features from speech and correlates them with blendshape parameters, serving as a bridge that improves semantic representation accuracy without requiring direct vertex manipulation.
2Device complexity
If vertex-driven animation generation is used, then the computational complexity is low, but the facial expression diversity deteriorates
Solution Approach 1:
The patent transforms the animation control from vertex positions to blendshape parameters, which inherently provide greater facial expression diversity. Blendshape parameters can independently control different facial muscle groups, enabling more varied and natural expressions while maintaining reasonable computational complexity through efficient parameter management.
Solution Approach 2:
The patent introduces dynamic semantic correlation that adapts blendshape parameter selection based on speech content. The system dynamically adjusts which blendshape parameters are activated and to what extent, enabling diverse facial expressions that match the semantic meaning of different speech segments without requiring excessive computational resources.
3Manufacturing precision
If blendshape parameters are used with semantic correlation, then the animation accuracy improves, but the model complexity increases
Solution Approach 1:
The patent extracts only the semantically relevant features from speech data and correlates them with corresponding blendshape parameters. By selectively extracting and using only the necessary semantic correlations rather than processing all speech features, the model achieves high animation accuracy while controlling complexity through feature selection.
Solution Approach 2:
The patent segments the facial animation control into multiple independent blendshape parameter groups, each correlated with specific semantic features of speech. This segmentation allows the model to process and control different facial expressions independently, improving overall animation accuracy while managing complexity through modular parameter organization.
Data Source
AI summary
An animation generation method and apparatus, an electronic device, and a storage medium. The method includes: obtaining at least one control condition determined based on at least one speech feature subsequence and a facial expression style feature; generating, by using a preset decoder, a target blendshape parameter corresponding to the at least one speech feature subsequence based on the at least one control condition and at least one preset variable, where the preset decoder is included in a preset generation model constructed based on at least one sample control condition and at least one second blendshape parameter sequence, and the at least one second blendshape parameter sequence is related to semantics of sample speech data corresponding to the at least one sample control condition; and deforming an object model based on at least one target blendshape parameter in sequence to generate a facial animation corresponding to speech data.


