Virtual Avatar Expression Transfer Using Facial Keypoint Coefficients
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating realistic and expressive portrait videos from single face images face challenges in computational efficiency, diversity, and stability, particularly in deep learning-based approaches like GANs and diffusion models, which are resource-intensive and prone to mode collapse.
Innovation Solution
A data processing method that involves facial keypoint extraction, adjustment of expression coefficients based on preset bases, and generation of transformed images using these coefficients, optimizing computational resources and improving generation quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If deep learning-based approaches (GANs and diffusion models) are used for generating portrait videos, then the realism and expressiveness of generated videos are improved, but computational resource consumption increases and stability deteriorates due to mode collapse
Solution Approach 1:
The patent extracts and processes only the essential facial keypoint information from input images, separating the critical expressive features from the complete image data. By focusing computation on extracted facial keypoints rather than entire images, the method reduces computational resource consumption while maintaining generation quality.
Solution Approach 2:
The patent transforms facial expression information into a parameterized representation using expression coefficients that control morphable face models. This parameterization converts complex image processing into manageable coefficient adjustments, improving both computational efficiency and stability in video generation.
2Manufacturing precision
If deep learning-based approaches (GANs and diffusion models) are used for generating portrait videos, then the realism and expressiveness of generated videos are improved, but stability deteriorates due to mode collapse
Solution Approach 1:
The patent replaces unstable deep learning generation processes with a more stable pipeline that uses extracted facial keypoints and expression coefficients to control morphable face models. This substitution of the generation mechanism eliminates mode collapse issues while maintaining realistic and expressive video output.
Solution Approach 2:
By representing facial expressions through controlled parameter adjustments (expression coefficients) rather than relying on probabilistic deep learning sampling, the method achieves stable and reproducible generation results, preventing mode collapse while preserving generation quality.
3Productivity
If facial keypoint extraction and expression coefficient adjustment are used, then computational efficiency is improved, but the complexity of the processing pipeline increases
Solution Approach 1:
The patent segments the face processing into distinct components: keypoint extraction, expression coefficient calculation, and model parameter adjustment. This segmentation allows each component to be optimized independently, improving overall computational efficiency while making the pipeline more manageable despite increased steps.
Data Source
AI summary
A method includes: obtaining a target image including a face of a target object; performing facial keypoint extraction on the target image to obtain a first facial keypoint image; obtaining a first set of expression coefficients based on the first facial keypoint image and a preset set of expression bases; adjusting a corresponding expression coefficient in the first set of expression coefficients to obtain a second set of expression coefficients; obtaining a second facial keypoint image based on the second set of expression coefficients and the set of expression bases; and obtaining, based on the second facial keypoint image and the target image, a first image corresponding to the target image that has undergone a facial expression transformation.


