Identity-Preserving Talking Face Generation via Attention Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating realistic talking faces from speech input fail to preserve the identity of the target individual and often produce implausible mouth shapes or lack natural eye blink movements, leading to unrealistic and visually displeasing animations.

Innovation Solution

The system employs a processor-implemented method that extracts DeepSpeech features from audio speech, generates speech-induced motion on a neutral mean face using a speech-to-landmark generation network, and combines this with attention-based texture generation using an attention map and color map learned from identity images, ensuring accurate audio-visual synchronization and natural eye blink movements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing methods generate realistic talking faces from speech input, then audio-visual synchronization is achieved, but identity preservation deteriorates

Engineering Contradiction:
Improveaudio-visual synchronization accuracyVSAvoididentity preservation
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The method segments the facial animation generation into independent components: extracting speech-driven motion from audio input, generating eye blink movements separately, and combining them with identity-preserving texture mapping. This segmentation allows each component to be optimized independently while maintaining overall coherence.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary attention map that selectively preserves identity-related regions while allowing speech-driven motion in other areas. This attention map acts as a mediator between the speech input and the final facial animation, ensuring identity preservation while maintaining audio-visual synchronization.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If existing methods generate talking face animations, then speech-driven motion is achieved, but natural eye blink movements are lost

Engineering Contradiction:
Improvespeech-driven motion generationVSAvoidnaturalness of eye blink movements
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The method performs preliminary generation of eye blink movements independent of the speech input, then integrates them with the speech-driven facial animations. This preliminary action ensures that natural eye blink patterns are established before being combined with speech-driven motions, preventing their loss in the final animation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent merges speech-driven facial motion with independently generated eye blink movements by combining their respective landmark point sequences. This merging process integrates two separate motion generation processes into a unified facial animation that preserves both speech synchronization and natural eye behavior.

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If existing methods generate talking faces from speech, then audio input is utilized, but implausible mouth shapes are produced

Engineering Contradiction:
Improveaudio input utilizationVSAvoidmouth shape plausibility
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The method applies local quality control by generating mouth shapes that are specifically adapted to the speech content while maintaining anatomical plausibility. The speech-driven motion is applied locally to relevant facial regions, ensuring that mouth shapes are both responsive to audio input and visually plausible.

Inventive Principle:
Principle #3Local quality

4Productivity

If existing methods generate facial animations, then speech input is processed, but identity of target individual is not preserved

Engineering Contradiction:
Improvefacial animation generationVSAvoididentity preservation
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent utilizes color map manipulation to preserve identity characteristics in the generated facial animations. By carefully controlling the color and texture mapping between the identity image and the animated face, the method maintains identity preservation while enabling dynamic facial expressions and motions.

Inventive Principle:
Principle #32Color changes

Data Source

PatentUS20210366173A1Identity preserving realistic talking face generation using audio speech of a user
Publication Date: 2021.11.25 TATA CONSULTANCY SERVICES LTD
  • US20210366173A1 patent drawing
  • US20210366173A1 patent drawing
  • US20210366173A1 patent drawing

AI summary

Speech-driven facial animation is useful for a variety of applications such as telepresence, chatbots, etc. The necessary attributes of having a realistic face animation are: 1) audiovisual synchronization, (2) identity preservation of the target individual, (3) plausible mouth movements, and (4) presence of natural eye blinks. Existing methods mostly address audio-visual lip synchronization, and synthesis of natural facial gestures for overall video realism. However, existing approaches are not accurate. Present disclosure provides system and method that learn motion of facial landmarks as an intermediate step before generating texture. Person-independent facial landmarks are generated from audio for invariance to different voices, accents, etc. Eye blinks are imposed on facial landmarks and the person-independent landmarks are retargeted to person-specific landmarks to preserve identity related facial structure. Facial texture is then generated from person-specific facial landmarks that helps to preserve identity-related texture.