Facial Animation Generation Using Decoupled Blendshape Sequences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating lip sync animations are vertex-driven, limiting the diversity and consistency of facial expressions in animations, especially when dealing with speech data.

Innovation Solution

A method that utilizes blendshape parameters unrelated to speech semantics, combined with a preset generation model and decoder, to drive an object model for generating facial animations that match speech data while maintaining consistent facial expressions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If vertex-driven animation generation is used, then the animation can be generated based on speech data, but the diversity and consistency of facial expressions are limited

Engineering Contradiction:
Improvediversity of facial expressionsVSAvoidcomplexity of animation generation system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the facial animation generation into two independent parts: speech-driven lip movement generation and expression-driven facial animation generation. The first blendshape parameter sequence captures speech-related lip movements, while the second blendshape parameter sequence captures expression-related facial movements. This segmentation allows each sequence to be optimized independently, improving both diversity and consistency of facial expressions while maintaining speech accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism using a generation model that takes both the speech feature and the first blendshape parameter sequence as inputs, and outputs the second blendshape parameter sequence. This intermediary allows the system to combine speech information with expression information in a controlled manner, enabling diverse and consistent facial expressions that accurately represent speech data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Stability of the object's composition

If vertex-driven animation generation is used, then the animation can be generated based on speech data, but the consistency of facial expressions is limited

Engineering Contradiction:
Improveconsistency of facial expressionsVSAvoidcomplexity of animation generation system
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The patent segments the facial animation generation into two independent parts: speech-driven lip movement generation and expression-driven facial animation generation. The first blendshape parameter sequence captures speech-related lip movements, while the second blendshape parameter sequence captures expression-related facial movements. This segmentation allows each sequence to be optimized independently, improving both diversity and consistency of facial expressions while maintaining speech accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism using a generation model that takes both the speech feature and the first blendshape parameter sequence as inputs, and outputs the second blendshape parameter sequence. This intermediary allows the system to combine speech information with expression information in a controlled manner, enabling diverse and consistent facial expressions that accurately represent speech data.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If speech data is directly used to drive facial animation, then the lip sync accuracy is improved, but the facial expression diversity is reduced

Engineering Contradiction:
Improvelip sync accuracyVSAvoidfacial expression diversity
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the facial animation generation into two independent parts: speech-driven lip movement generation and expression-driven facial animation generation. The first blendshape parameter sequence captures speech-related lip movements, while the second blendshape parameter sequence captures expression-related facial movements. This segmentation allows each sequence to be optimized independently, improving both diversity and consistency of facial expressions while maintaining speech accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism using a generation model that takes both the speech feature and the first blendshape parameter sequence as inputs, and outputs the second blendshape parameter sequence. This intermediary allows the system to combine speech information with expression information in a controlled manner, enabling diverse and consistent facial expressions that accurately represent speech data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250371776A1Method of animation generation, electronic device, and storage medium
Publication Date: 2025.12.04 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250371776A1 patent drawing
  • US20250371776A1 patent drawing
  • US20250371776A1 patent drawing

AI summary

A method of animation generation, an electronic device, and a storage medium are provided. The method includes: obtaining a speech feature of speech data and a first blendshape parameter sequence, the first blendshape parameter sequence being unrelated to semantics of the speech data; generating an encoded sequence based on the speech feature and the first blendshape parameter sequence by using a preset generation model; decoding the encoded sequence into a second blendshape parameter sequence by using a preset decoder; and driving an object model based on the second blendshape parameter sequence to generate a facial animation corresponding to the speech data.