Expression Driving Model Real-Time Video Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing generative adversarial network models require a large amount of data computation, resulting in poor real-time performance of generated videos when driving static character images.

Innovation Solution

An expression driving method and apparatus that utilizes a pre-trained expression driving model trained on target and first sample images, where facial images in the generated video are based on the target sample image, and the gesture expression features are preserved, allowing for real-time processing with reduced computational complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a generative adversarial network model is used to drive the static character image, then the generated video can be produced, but the real-time performance is poor due to large amount of data computation

Engineering Contradiction:
Improvereal-time performanceVSAvoiddata computation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the facial driving process into two independent parts: a pre-trained expression driving model that processes driving information, and a facial feature map generation module that applies the driving information to a target facial image. This segmentation allows the computationally intensive pre-training to be separated from real-time inference, improving real-time performance while maintaining generation quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by pre-training the expression driving model offline using a large dataset of facial images and driving information. This pre-training phase completes all complex computational tasks beforehand, leaving only lightweight inference operations for real-time video generation. The pre-trained model captures complex facial expression relationships without requiring complex computation during actual video generation.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If the expression driving model is trained with large datasets to maintain accuracy, then the gesture expression features are preserved, but the computational complexity increases

Engineering Contradiction:
Improvegesture expression feature accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent trains the expression driving model in advance using large datasets containing diverse facial expressions and corresponding driving information. This preliminary training allows the model to learn complex gesture expression relationships and preserve accuracy during inference. The training phase completes all data processing and pattern learning beforehand, so that real-time generation only requires applying the pre-learned relationships without reprocessing the large datasets.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a simplified copy of the complex facial driving process through the pre-trained model. Instead of recomputing complex facial expression relationships in real-time, the system uses the pre-trained model's learned representations to quickly generate driving information. This copying approach maintains gesture expression feature accuracy by preserving the learned relationships while significantly reducing computational complexity during video generation.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250078570A1Expression driving method and device, and expression driving model training method and device
Publication Date: 2025.03.06 LEMON INC(GB)
  • US20250078570A1 patent drawing
  • US20250078570A1 patent drawing
  • US20250078570A1 patent drawing

AI summary

The present disclosure provides an expression driving method and apparatus, and a training method and apparatus of an expression driving model. The expression driving method includes acquiring a first video; and inputting the first video into a pre-trained expression driving model to obtain a second video. The expression driving model is trained based on a target sample image and a plurality of first sample images. A facial image in the second video is generated based on the target sample image. A gesture expression feature of the facial image in the second video is the same as a gesture expression feature of a facial image in the first video.