Emotion-Controlled Talking Face Generation via Geometry-Aware Landmarks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional techniques for realistic talking face generation struggle with efficiently controlling emotions over the face, often resulting in inconsistent emotions and limited generalization to arbitrary unknown target faces.
Innovation Solution
A processor-implemented method and system for emotion-controllable generalized talking face generation, which involves training a geometry-aware landmark generation network and a flow-guided texture generation network using speech audio, emotion data, and neutral face images to predict emotion-driven facial landmarks and texture maps for arbitrary target faces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional techniques use external emotion control using one-hot emotion vector, then emotion control is achieved, but the lower part of the face is animated from audio independently resulting in inconsistent emotions over the face
Solution Approach 1:
The patent merges emotion control and audio-driven animation into a unified framework. The emotion encoder processes both emotion labels and audio signals, and the landmark generation network integrates these inputs to produce consistent facial expressions across the entire face, eliminating the independence problem between upper and lower face animation
Solution Approach 2:
The patent introduces an intermediary landmark generation network that acts as a mediator between emotion control signals and facial animation. This network processes emotion embeddings and audio features to generate geometry-aware facial landmarks, which then guide the texture generation network to produce consistent emotional expressions throughout the face
2Reliability
If conventional techniques focus on generating consistent emotions over the full face using disentangled emotion latent feature, then emotional consistency is improved, but the technique relies on intermediate global landmarks or edge maps to generate texture directly which limits generalization to arbitrary unknown target face
Solution Approach 1:
The patent segments the facial animation process into two independent but coordinated networks: a geometry-aware landmark generation network that handles facial structure and landmark detection, and a flow-guided texture generation network that handles texture synthesis. This segmentation allows each network to specialize in specific tasks, improving both emotional consistency and generalization capability
Solution Approach 2:
The patent transitions from working directly with pixel-level textures and global landmarks to operating in the dimension of geometry-aware facial landmarks. By representing facial geometry through a graph of landmarks with explicit spatial relationships, the system achieves better generalization to arbitrary faces while maintaining emotional consistency through the structured representation
3Productivity
If conventional techniques learn facial emotions implicitly from audio, then emotion generation is achieved, but the technique is not efficient to control the emotion over the face
Solution Approach 1:
The patent introduces an emotion encoder as an intermediary that explicitly processes emotion labels and audio signals to generate emotion embeddings. This explicit emotion representation serves as a mediator that can be independently controlled and adjusted, allowing efficient emotion generation while maintaining precise control over facial expressions
Solution Approach 2:
The patent performs preliminary processing of audio signals to extract emotion-relevant features before they are used for facial animation. The audio encoder processes audio input to generate emotion-invariant speech embedding features, which are then combined with emotion labels to create comprehensive emotion representations that guide the landmark generation network
Data Source
AI summary
This disclosure relates generally to methods and systems for emotion-controllable generalized talking face generation of an arbitrary face image. Most of the conventional techniques for the realistic talking face generation may not be efficient to control the emotion over the face and have limited scope of generalization to an arbitrary unknown target face. The present disclosure proposes a graph convolutional network that uses speech content feature along with an independent emotion input to generate emotion and speech-induced motion on facial geometry-aware landmark representation. The facial geometry-aware landmark representation is further used in by an optical flow-guided texture generation network for producing the texture. A two-branch optical flow-guided texture generation network with motion and texture branches is designed to consider the motion and texture content independently. The optical flow-guided texture generation network then renders emotional talking face animation from a single image of any arbitrary target face.


