Audio-Driven 3D Facial Animation Model for Lip-Sync Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio-driven three-dimensional facial animation technologies face challenges with mismatched lip-sync actions and speech, and poor naturalness of expression actions, leading to reduced accuracy in audio-driven three-dimensional facial animation.

Innovation Solution

An audio-driven three-dimensional facial animation model generation method that involves acquiring sample data including audio and speaking style data, performing feature extraction and encoding to improve blend shape value accuracy, and updating model parameters based on loss function values to enhance animation accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If speech-driven expression technologies are used to generate facial animation, then facial expressions can be generated from speech, but the lip-sync actions do not match the speech and the expression actions lack naturalness

Engineering Contradiction:
Improvefacial animation accuracyVSAvoidlip-sync synchronization
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent segments the speech signal into multiple feature dimensions (spectral features, temporal features, prosodic features) and processes each dimension separately through dedicated neural network branches. This segmentation allows precise control over different aspects of facial animation (lip-sync, expression, pitch) independently, resolving the contradiction between overall animation accuracy and specific lip-sync synchronization reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the speech signal from traditional time-domain representation to multiple parameter spaces including spectral parameters (Mel-frequency cepstral coefficients), temporal parameters (energy contours, zero-crossing rates), and prosodic parameters (pitch contours, stress patterns). These parameter transformations enable more accurate mapping to facial animation parameters, improving both manufacturing precision and reliability of lip-sync synchronization.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If traditional speech preprocessing is used to generate blend shape values, then processing is simple, but the output has poor naturalness of expression actions

Engineering Contradiction:
Improveprocessing simplicityVSAvoidexpression naturalness
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent introduces intermediate representation layers that mediate between raw speech signals and final blend shape values. Speech signals first pass through feature extraction modules that create intermediate representations (spectrograms, mel-frequency features), which then undergo further processing through multiple neural network layers before generating facial animation parameters. These intermediaries enable complex expression naturalness while maintaining processing efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent adds temporal dimension processing by analyzing speech signals across multiple time frames and using recurrent neural network layers to capture temporal dynamics. This dimensional expansion from static spectral analysis to dynamic temporal-spectral analysis enables more natural expression actions that reflect the temporal evolution of speech, while the parallel processing architecture maintains computational efficiency.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12254552B1Audio-driven three-dimensional facial animation model generation method and apparatus, and electronic device
Publication Date: 2025.03.18 NANJING SILICON INTELLIGENCE TECH CO LTD
  • US12254552B1 patent drawing
  • US12254552B1 patent drawing
  • US12254552B1 patent drawing

AI summary

This application provides a audio-driven three-dimensional facial animation model generation method and apparatus, and an electronic device. The method includes: acquiring sample data including sample audio data, sample speaking style data, and a sample blend shape value; performing feature extraction on the sample audio data to obtain a sample audio feature; performing convolution on the sample audio feature based on a to-be-trained audio-driven three-dimensional facial animation model to obtain an initial audio feature, and performing encoding on the sample speaking style data based on the to-be-trained audio-driven three-dimensional facial animation model to obtain a sample speaking style feature; performing encoding on the initial audio feature and the sample speaking style feature based on the to-be-trained audio-driven three-dimensional facial animation model, to obtain an output blend shape value; and performing calculation on the sample blend shape value and the output blend shape value to obtain a loss function value.