Audio-Driven Facial Motion Alignment for Realistic Video Conferencing Avatars

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing 2D digital human products lack natural head posture changes, leading to unnatural facial movements and poor realism in generated videos due to fixed head positions during voice-driven synthesis.

Innovation Solution

A method for data processing that adjusts motion feature information of a human face based on audio information to generate a target human face image sequence, incorporating motion feature information and target human face images to enhance consistency between facial movements and audio.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If motion feature information is adjusted according to audio information to improve facial movement consistency, then the realism of generated videos is improved, but the computational complexity increases

Engineering Contradiction:
Improvefacial movement consistencyVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the facial motion adjustment process into distinct modules: audio feature extraction, motion feature extraction, feature alignment, and image generation. This modular segmentation allows each component to be optimized independently, improving facial movement consistency while managing computational complexity through distributed processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary extraction and alignment of audio and motion features before the actual video generation process. By pre-processing and storing aligned feature representations, the system reduces real-time computational requirements during video synthesis while maintaining high facial movement consistency.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If head posture information is kept fixed to simplify processing, then the computational tasks are reduced, but the naturalness of digital human face deteriorates

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidnaturalness of digital human face
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent implements dynamic head posture adjustment by integrating audio-driven motion features that automatically modulate head position and orientation based on speech content. This dynamic approach maintains processing efficiency through learned motion patterns while significantly improving the naturalness of digital human faces compared to fixed posture methods.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes key parameters including head pose angles, facial expression intensities, and motion feature weights based on audio characteristics. By dynamically adjusting these parameters according to speech emotion and content, the system achieves high naturalness without proportionally increasing processing complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250342640A1Data processing method and device, video conferencing system, storage medium
Publication Date: 2025.11.06 ZTE CORP
  • US20250342640A1 patent drawing
  • US20250342640A1 patent drawing
  • US20250342640A1 patent drawing

AI summary

A method for data processing and, a device, a video conferencing system, and a computer readable storage medium. The method may include, acquiring audio information, motion feature information of a human face, and a target human face image; adjusting the motion feature information according to the audio information to acquire target motion feature information of the human face; and generating a target human face image sequence corresponding to the audio information according to the target motion feature information and the target human face image.