Face Video Generation Using Segmented Mouth Shape Drivers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current human face and mouth shape driving solutions generate videos in a generic style, failing to accurately reflect the personal mouth shape style of the target object, resulting in low accuracy of the generated human face video.
Innovation Solution
A method involving the use of a mouth-shape multimedia resource and a reference human face image to extract mouth-shape driving features, generate stylistic human face images, and create a stylistic human face video that reflects the personal mouth shape style of the target object, utilizing a trained human face and mouth shape driver model with a feature extraction network and a human face driving network.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a general-purpose human face and mouth shape driver model is used, then the device complexity is reduced, but the manufacturing precision of the generated video deteriorates
Solution Approach 1:
The driver model is segmented into multiple specialized sub-models: a mouth shape driver model for lip movements, an eye driver model for eye movements, and a facial expression driver model for facial expressions. Each sub-model specializes in a specific aspect, improving overall video accuracy without requiring a single overly complex general-purpose model.
Solution Approach 2:
Different parts of the face are processed with different levels of detail and specialized models. The mouth region uses a dedicated mouth shape driver model, eyes use an eye driver model, and other facial regions use the facial expression driver model. This local specialization ensures high precision for each facial component while maintaining manageable overall system complexity.
2Ease of operation
If a general-purpose human face and mouth shape driver model is used, then the ease of operation is improved, but the reliability of the generated video deteriorates
Solution Approach 1:
The driver model is divided into specialized sub-models for different facial components (mouth, eyes, expressions), where each sub-model is trained on specific data and handles particular aspects of facial animation. This segmentation improves reliability by ensuring each component is driven by a model optimized for its specific function, while the integrated system remains easy to operate through unified input processing.
Solution Approach 2:
The system employs multiple loss functions during training that provide feedback on different aspects of video quality: content loss for structural accuracy, perceptual loss for visual realism, and style loss for capturing individual characteristics. This multi-faceted feedback mechanism ensures high reliability and authenticity in the generated videos while maintaining ease of operation.
3Manufacturing precision
If feature extraction and style transfer processes are implemented, then the manufacturing precision of the video is improved, but the productivity of the generation process deteriorates
Solution Approach 1:
Style vectors are extracted and stored in advance from reference images of target objects. During video generation, these pre-extracted style vectors are directly applied without requiring real-time style transfer computations. This preliminary action significantly improves generation speed while maintaining high video quality through the use of pre-computed stylistic features.
Solution Approach 2:
The style transfer process is separated into a distinct preprocessing step where style vectors are extracted from reference images before the main video generation process. This extraction is performed once and reused across multiple video generations, reducing the computational burden during actual video synthesis and improving overall productivity while maintaining manufacturing precision.
Data Source
AI summary
A method for generating a human face video, an apparatus for generating a human face video, includes: obtaining a mouth-shape multimedia resource and a reference human face image of a target object; obtaining a reference style vector of the target object; for each resource frame in the mouth-shape multimedia resource, obtaining a mouth-shape driving feature by performing a feature extraction process on the resource frame; generating a stylistic human face image corresponding to the resource frame according to the mouth-shape driving feature, the reference human face image, and the reference style vector; and determining a stylistic human face video of the target object.


