Expression Transfer Model Training for Realistic Facial Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing facial expression transfer methods require complex image processing, leading to increased training difficulty and cost, and often result in unrealistic facial images.

Innovation Solution

A training method for an expression transfer model that utilizes a source and target domain facial images, along with facial feature images, to update the model parameters based on discriminative and classification results, reducing the need for complex image processing and enhancing the model's ability to output realistic facial images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If complex image processing is performed on facial images to achieve expression transfer, then the expression transfer accuracy is improved, but the training difficulty and cost increase

Engineering Contradiction:
Improveexpression transfer accuracyVSAvoidtraining complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and separates the expression information from the complex facial image processing by using a pre-trained face recognition model to obtain facial feature images. This extraction approach isolates the essential expression features (landmarks, contours) from the complex image data, allowing expression transfer without requiring complex image processing during training, thus reducing training complexity while maintaining accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces a facial feature image as an intermediary between the source and target facial images. This intermediary contains extracted expression information that mediates the transfer process, enabling the model to learn expression mapping without directly processing complex pixel-level image transformations, thereby reducing training difficulty

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If complex image processing is performed on facial images to achieve expression transfer, then the expression transfer accuracy is improved, but the training cost increases

Engineering Contradiction:
Improveexpression transfer accuracyVSAvoidtraining cost
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The patent performs preliminary action by pre-training a face recognition model to extract facial features before the main expression transfer training. This preliminary feature extraction creates ready-to-use facial feature images that contain essential expression information, eliminating the need for complex image processing during the subsequent expression transfer training, thus reducing training cost while maintaining accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts essential expression features from facial images in advance through face recognition model preprocessing, separating the essential information from complex image data. This extraction allows the main training to focus only on expression mapping without redundant complex image processing, reducing computational cost

Inventive Principle:
Principle #2Taking out (Extraction)

3Ease of operation

If traditional expression transfer methods are used, then the training process is simplified, but the realism of synthesized facial images deteriorates

Engineering Contradiction:
Improvetraining easeVSAvoidimage realism
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces a discriminative network model as an intermediary that provides feedback on the realism of synthesized images. This mediator evaluates whether the transferred expression maintains the target image's characteristics, guiding the generator to produce more realistic results without requiring complex manual processing or adjustment

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements a feedback mechanism where the discriminative network model evaluates the synthesized facial images and provides loss signals back to the expression transfer model. This feedback loop continuously improves the realism of generated images by penalizing unrealistic artifacts while maintaining training simplicity through automated evaluation

Inventive Principle:
Principle #23Feedback

4Device complexity

If the expression transfer model is trained without discriminative feedback, then the training process is simpler, but the model performance deteriorates

Engineering Contradiction:
Improvetraining process complexityVSAvoidmodel performance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent implements feedback through a discriminative network model that evaluates synthesized images and provides loss signals to guide the generator. This feedback mechanism improves model performance by continuously correcting unrealistic outputs while maintaining relatively simple training procedures through automated differentiation and gradient-based optimization

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces manual or mechanical adjustment methods with an automated discriminative feedback system. Instead of manually tuning the model to produce realistic images, the system uses automated discrimination and loss calculation to guide optimization, improving performance while keeping the training process manageable through algorithmic rather than manual control

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12361751B2Training method for expression transfer model, expression transfer method and apparatus
Publication Date: 2025.07.15 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12361751B2 patent drawing
  • US12361751B2 patent drawing
  • US12361751B2 patent drawing

AI summary

Disclosed is a method for training an expression transfer model performed by a computer device, the method including: obtaining a source domain facial image of a first object, a target domain facial image of a second object and a facial feature image associated with the target domain facial image; applying the facial feature image and the source domain facial image to an expression transfer model to obtain a synthesized facial image of the first object; applying the synthesized facial image and the target domain facial image to a discriminative network model obtain two discrimination results; applying the synthesized facial image and the target domain facial image to an image classification model to obtain a category feature vector for identifying a difference between the synthesized facial image and the target domain facial image; and updating the expression transfer model according to the category feature vector and the discrimination results.