Audio-Driven Facial Motion Alignment for Realistic Video Conferencing Avatars
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing 2D digital human products lack natural head posture changes, leading to unnatural facial movements and poor realism in generated videos due to fixed head positions during voice-driven synthesis.
Innovation Solution
A method for data processing that adjusts motion feature information of a human face based on audio information to generate a target human face image sequence, incorporating motion feature information and target human face images to enhance consistency between facial movements and audio.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If motion feature information is adjusted according to audio information to improve facial movement consistency, then the realism of generated videos is improved, but the computational complexity increases
Solution Approach 1:
The patent segments the facial motion adjustment process into distinct modules: audio feature extraction, motion feature extraction, feature alignment, and image generation. This modular segmentation allows each component to be optimized independently, improving facial movement consistency while managing computational complexity through distributed processing.
Solution Approach 2:
The patent performs preliminary extraction and alignment of audio and motion features before the actual video generation process. By pre-processing and storing aligned feature representations, the system reduces real-time computational requirements during video synthesis while maintaining high facial movement consistency.
2Productivity
If head posture information is kept fixed to simplify processing, then the computational tasks are reduced, but the naturalness of digital human face deteriorates
Solution Approach 1:
The patent implements dynamic head posture adjustment by integrating audio-driven motion features that automatically modulate head position and orientation based on speech content. This dynamic approach maintains processing efficiency through learned motion patterns while significantly improving the naturalness of digital human faces compared to fixed posture methods.
Solution Approach 2:
The patent changes key parameters including head pose angles, facial expression intensities, and motion feature weights based on audio characteristics. By dynamically adjusting these parameters according to speech emotion and content, the system achieves high naturalness without proportionally increasing processing complexity.
Data Source
AI summary
A method for data processing and, a device, a video conferencing system, and a computer readable storage medium. The method may include, acquiring audio information, motion feature information of a human face, and a target human face image; adjusting the motion feature information according to the audio information to acquire target motion feature information of the human face; and generating a target human face image sequence corresponding to the audio information according to the target motion feature information and the target human face image.


