Facial Animation Generation Using Decoupled Blendshape Sequences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating lip sync animations are vertex-driven, limiting the diversity and consistency of facial expressions in animations, especially when dealing with speech data.
Innovation Solution
A method that utilizes blendshape parameters unrelated to speech semantics, combined with a preset generation model and decoder, to drive an object model for generating facial animations that match speech data while maintaining consistent facial expressions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If vertex-driven animation generation is used, then the animation can be generated based on speech data, but the diversity and consistency of facial expressions are limited
Solution Approach 1:
The patent segments the facial animation generation into two independent parts: speech-driven lip movement generation and expression-driven facial animation generation. The first blendshape parameter sequence captures speech-related lip movements, while the second blendshape parameter sequence captures expression-related facial movements. This segmentation allows each sequence to be optimized independently, improving both diversity and consistency of facial expressions while maintaining speech accuracy.
Solution Approach 2:
The patent introduces an intermediary mechanism using a generation model that takes both the speech feature and the first blendshape parameter sequence as inputs, and outputs the second blendshape parameter sequence. This intermediary allows the system to combine speech information with expression information in a controlled manner, enabling diverse and consistent facial expressions that accurately represent speech data.
2Stability of the object's composition
If vertex-driven animation generation is used, then the animation can be generated based on speech data, but the consistency of facial expressions is limited
Solution Approach 1:
The patent segments the facial animation generation into two independent parts: speech-driven lip movement generation and expression-driven facial animation generation. The first blendshape parameter sequence captures speech-related lip movements, while the second blendshape parameter sequence captures expression-related facial movements. This segmentation allows each sequence to be optimized independently, improving both diversity and consistency of facial expressions while maintaining speech accuracy.
Solution Approach 2:
The patent introduces an intermediary mechanism using a generation model that takes both the speech feature and the first blendshape parameter sequence as inputs, and outputs the second blendshape parameter sequence. This intermediary allows the system to combine speech information with expression information in a controlled manner, enabling diverse and consistent facial expressions that accurately represent speech data.
3Measurement precision
If speech data is directly used to drive facial animation, then the lip sync accuracy is improved, but the facial expression diversity is reduced
Solution Approach 1:
The patent segments the facial animation generation into two independent parts: speech-driven lip movement generation and expression-driven facial animation generation. The first blendshape parameter sequence captures speech-related lip movements, while the second blendshape parameter sequence captures expression-related facial movements. This segmentation allows each sequence to be optimized independently, improving both diversity and consistency of facial expressions while maintaining speech accuracy.
Solution Approach 2:
The patent introduces an intermediary mechanism using a generation model that takes both the speech feature and the first blendshape parameter sequence as inputs, and outputs the second blendshape parameter sequence. This intermediary allows the system to combine speech information with expression information in a controlled manner, enabling diverse and consistent facial expressions that accurately represent speech data.
Data Source
AI summary
A method of animation generation, an electronic device, and a storage medium are provided. The method includes: obtaining a speech feature of speech data and a first blendshape parameter sequence, the first blendshape parameter sequence being unrelated to semantics of the speech data; generating an encoded sequence based on the speech feature and the first blendshape parameter sequence by using a preset generation model; decoding the encoded sequence into a second blendshape parameter sequence by using a preset decoder; and driving an object model based on the second blendshape parameter sequence to generate a facial animation corresponding to the speech data.


