Identity-Preserving Talking Face Generation via Attention Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating realistic talking faces from speech input fail to preserve the identity of the target individual and often produce implausible mouth shapes or lack natural eye blink movements, leading to unrealistic and visually displeasing animations.
Innovation Solution
The system employs a processor-implemented method that extracts DeepSpeech features from audio speech, generates speech-induced motion on a neutral mean face using a speech-to-landmark generation network, and combines this with attention-based texture generation using an attention map and color map learned from identity images, ensuring accurate audio-visual synchronization and natural eye blink movements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing methods generate realistic talking faces from speech input, then audio-visual synchronization is achieved, but identity preservation deteriorates
Solution Approach 1:
The method segments the facial animation generation into independent components: extracting speech-driven motion from audio input, generating eye blink movements separately, and combining them with identity-preserving texture mapping. This segmentation allows each component to be optimized independently while maintaining overall coherence.
Solution Approach 2:
The patent introduces an intermediary attention map that selectively preserves identity-related regions while allowing speech-driven motion in other areas. This attention map acts as a mediator between the speech input and the final facial animation, ensuring identity preservation while maintaining audio-visual synchronization.
2Productivity
If existing methods generate talking face animations, then speech-driven motion is achieved, but natural eye blink movements are lost
Solution Approach 1:
The method performs preliminary generation of eye blink movements independent of the speech input, then integrates them with the speech-driven facial animations. This preliminary action ensures that natural eye blink patterns are established before being combined with speech-driven motions, preventing their loss in the final animation.
Solution Approach 2:
The patent merges speech-driven facial motion with independently generated eye blink movements by combining their respective landmark point sequences. This merging process integrates two separate motion generation processes into a unified facial animation that preserves both speech synchronization and natural eye behavior.
3Adaptability or versatility
If existing methods generate talking faces from speech, then audio input is utilized, but implausible mouth shapes are produced
Solution Approach 1:
The method applies local quality control by generating mouth shapes that are specifically adapted to the speech content while maintaining anatomical plausibility. The speech-driven motion is applied locally to relevant facial regions, ensuring that mouth shapes are both responsive to audio input and visually plausible.
4Productivity
If existing methods generate facial animations, then speech input is processed, but identity of target individual is not preserved
Solution Approach 1:
The patent utilizes color map manipulation to preserve identity characteristics in the generated facial animations. By carefully controlling the color and texture mapping between the identity image and the animated face, the method maintains identity preservation while enabling dynamic facial expressions and motions.
Data Source
AI summary
Speech-driven facial animation is useful for a variety of applications such as telepresence, chatbots, etc. The necessary attributes of having a realistic face animation are: 1) audiovisual synchronization, (2) identity preservation of the target individual, (3) plausible mouth movements, and (4) presence of natural eye blinks. Existing methods mostly address audio-visual lip synchronization, and synthesis of natural facial gestures for overall video realism. However, existing approaches are not accurate. Present disclosure provides system and method that learn motion of facial landmarks as an intermediate step before generating texture. Person-independent facial landmarks are generated from audio for invariance to different voices, accents, etc. Eye blinks are imposed on facial landmarks and the person-independent landmarks are retargeted to person-specific landmarks to preserve identity related facial structure. Facial texture is then generated from person-specific facial landmarks that helps to preserve identity-related texture.


