Expression Driving Model Real-Time Video Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing generative adversarial network models require a large amount of data computation, resulting in poor real-time performance of generated videos when driving static character images.
Innovation Solution
An expression driving method and apparatus that utilizes a pre-trained expression driving model trained on target and first sample images, where facial images in the generated video are based on the target sample image, and the gesture expression features are preserved, allowing for real-time processing with reduced computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a generative adversarial network model is used to drive the static character image, then the generated video can be produced, but the real-time performance is poor due to large amount of data computation
Solution Approach 1:
The patent segments the facial driving process into two independent parts: a pre-trained expression driving model that processes driving information, and a facial feature map generation module that applies the driving information to a target facial image. This segmentation allows the computationally intensive pre-training to be separated from real-time inference, improving real-time performance while maintaining generation quality.
Solution Approach 2:
The patent applies preliminary action by pre-training the expression driving model offline using a large dataset of facial images and driving information. This pre-training phase completes all complex computational tasks beforehand, leaving only lightweight inference operations for real-time video generation. The pre-trained model captures complex facial expression relationships without requiring complex computation during actual video generation.
2Measurement precision
If the expression driving model is trained with large datasets to maintain accuracy, then the gesture expression features are preserved, but the computational complexity increases
Solution Approach 1:
The patent trains the expression driving model in advance using large datasets containing diverse facial expressions and corresponding driving information. This preliminary training allows the model to learn complex gesture expression relationships and preserve accuracy during inference. The training phase completes all data processing and pattern learning beforehand, so that real-time generation only requires applying the pre-learned relationships without reprocessing the large datasets.
Solution Approach 2:
The patent creates a simplified copy of the complex facial driving process through the pre-trained model. Instead of recomputing complex facial expression relationships in real-time, the system uses the pre-trained model's learned representations to quickly generate driving information. This copying approach maintains gesture expression feature accuracy by preserving the learned relationships while significantly reducing computational complexity during video generation.
Data Source
AI summary
The present disclosure provides an expression driving method and apparatus, and a training method and apparatus of an expression driving model. The expression driving method includes acquiring a first video; and inputting the first video into a pre-trained expression driving model to obtain a second video. The expression driving model is trained based on a target sample image and a plurality of first sample images. A facial image in the second video is generated based on the target sample image. A gesture expression feature of the facial image in the second video is the same as a gesture expression feature of a facial image in the first video.


