Emotion-Controlled Talking Face Generation via Geometry-Aware Landmarks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques for realistic talking face generation struggle with efficiently controlling emotions over the face, often resulting in inconsistent emotions and limited generalization to arbitrary unknown target faces.

Innovation Solution

A processor-implemented method and system for emotion-controllable generalized talking face generation, which involves training a geometry-aware landmark generation network and a flow-guided texture generation network using speech audio, emotion data, and neutral face images to predict emotion-driven facial landmarks and texture maps for arbitrary target faces.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional techniques use external emotion control using one-hot emotion vector, then emotion control is achieved, but the lower part of the face is animated from audio independently resulting in inconsistent emotions over the face

Engineering Contradiction:
Improveemotion controlVSAvoidemotional consistency
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent merges emotion control and audio-driven animation into a unified framework. The emotion encoder processes both emotion labels and audio signals, and the landmark generation network integrates these inputs to produce consistent facial expressions across the entire face, eliminating the independence problem between upper and lower face animation

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary landmark generation network that acts as a mediator between emotion control signals and facial animation. This network processes emotion embeddings and audio features to generate geometry-aware facial landmarks, which then guide the texture generation network to produce consistent emotional expressions throughout the face

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If conventional techniques focus on generating consistent emotions over the full face using disentangled emotion latent feature, then emotional consistency is improved, but the technique relies on intermediate global landmarks or edge maps to generate texture directly which limits generalization to arbitrary unknown target face

Engineering Contradiction:
Improveemotional consistencyVSAvoidgeneralization to arbitrary face
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments the facial animation process into two independent but coordinated networks: a geometry-aware landmark generation network that handles facial structure and landmark detection, and a flow-guided texture generation network that handles texture synthesis. This segmentation allows each network to specialize in specific tasks, improving both emotional consistency and generalization capability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from working directly with pixel-level textures and global landmarks to operating in the dimension of geometry-aware facial landmarks. By representing facial geometry through a graph of landmarks with explicit spatial relationships, the system achieves better generalization to arbitrary faces while maintaining emotional consistency through the structured representation

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If conventional techniques learn facial emotions implicitly from audio, then emotion generation is achieved, but the technique is not efficient to control the emotion over the face

Engineering Contradiction:
Improveemotion generation efficiencyVSAvoidemotion control
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent introduces an emotion encoder as an intermediary that explicitly processes emotion labels and audio signals to generate emotion embeddings. This explicit emotion representation serves as a mediator that can be independently controlled and adjusted, allowing efficient emotion generation while maintaining precise control over facial expressions

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary processing of audio signals to extract emotion-relevant features before they are used for facial animation. The audio encoder processes audio input to generate emotion-invariant speech embedding features, which are then combined with emotion labels to create comprehensive emotion representations that guide the landmark generation network

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12307567B2Methods and systems for emotion-controllable generalized talking face generation
Publication Date: 2025.05.20 TATA CONSULTANCY SERVICES LTD
  • US12307567B2 patent drawing
  • US12307567B2 patent drawing
  • US12307567B2 patent drawing

AI summary

This disclosure relates generally to methods and systems for emotion-controllable generalized talking face generation of an arbitrary face image. Most of the conventional techniques for the realistic talking face generation may not be efficient to control the emotion over the face and have limited scope of generalization to an arbitrary unknown target face. The present disclosure proposes a graph convolutional network that uses speech content feature along with an independent emotion input to generate emotion and speech-induced motion on facial geometry-aware landmark representation. The facial geometry-aware landmark representation is further used in by an optical flow-guided texture generation network for producing the texture. A two-branch optical flow-guided texture generation network with motion and texture branches is designed to consider the motion and texture content independently. The optical flow-guided texture generation network then renders emotional talking face animation from a single image of any arbitrary target face.