Virtual Avatar Expression Transfer Using Facial Keypoint Coefficients

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating realistic and expressive portrait videos from single face images face challenges in computational efficiency, diversity, and stability, particularly in deep learning-based approaches like GANs and diffusion models, which are resource-intensive and prone to mode collapse.

Innovation Solution

A data processing method that involves facial keypoint extraction, adjustment of expression coefficients based on preset bases, and generation of transformed images using these coefficients, optimizing computational resources and improving generation quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If deep learning-based approaches (GANs and diffusion models) are used for generating portrait videos, then the realism and expressiveness of generated videos are improved, but computational resource consumption increases and stability deteriorates due to mode collapse

Engineering Contradiction:
Improvegeneration qualityVSAvoidcomputational resource consumption
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent extracts and processes only the essential facial keypoint information from input images, separating the critical expressive features from the complete image data. By focusing computation on extracted facial keypoints rather than entire images, the method reduces computational resource consumption while maintaining generation quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms facial expression information into a parameterized representation using expression coefficients that control morphable face models. This parameterization converts complex image processing into manageable coefficient adjustments, improving both computational efficiency and stability in video generation.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If deep learning-based approaches (GANs and diffusion models) are used for generating portrait videos, then the realism and expressiveness of generated videos are improved, but stability deteriorates due to mode collapse

Engineering Contradiction:
Improvegeneration qualityVSAvoidstability
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent replaces unstable deep learning generation processes with a more stable pipeline that uses extracted facial keypoints and expression coefficients to control morphable face models. This substitution of the generation mechanism eliminates mode collapse issues while maintaining realistic and expressive video output.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

By representing facial expressions through controlled parameter adjustments (expression coefficients) rather than relying on probabilistic deep learning sampling, the method achieves stable and reproducible generation results, preventing mode collapse while preserving generation quality.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If facial keypoint extraction and expression coefficient adjustment are used, then computational efficiency is improved, but the complexity of the processing pipeline increases

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidprocessing pipeline complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the face processing into distinct components: keypoint extraction, expression coefficient calculation, and model parameter adjustment. This segmentation allows each component to be optimized independently, improving overall computational efficiency while making the pipeline more manageable despite increased steps.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20260073729A1Data processing method and apparatus for virtual avatar, electronic device, and medium
Publication Date: 2026.03.12 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20260073729A1 patent drawing
  • US20260073729A1 patent drawing
  • US20260073729A1 patent drawing

AI summary

A method includes: obtaining a target image including a face of a target object; performing facial keypoint extraction on the target image to obtain a first facial keypoint image; obtaining a first set of expression coefficients based on the first facial keypoint image and a preset set of expression bases; adjusting a corresponding expression coefficient in the first set of expression coefficients to obtain a second set of expression coefficients; obtaining a second facial keypoint image based on the second set of expression coefficients and the set of expression bases; and obtaining, based on the second facial keypoint image and the target image, a first image corresponding to the target image that has undergone a facial expression transformation.