Speech-Driven Lip Video Generation With Smooth Frame Transitions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video generation methods using speech-driven lip shapes result in abrupt transitions between frames, leading to poor lip shape driving effects in generated videos.
Innovation Solution
A video generation method that involves acquiring target audio and video data, performing mask processing, feature extraction using a target multimodal model, and predicting lip areas based on synchronized audio and video features to ensure natural transitions in lip shape across frames.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If image generation model is used to generate lip shape from target speech, then lip shape can be generated, but the transition of lip shape among different video frames is abrupt
Solution Approach 1:
The patent segments the lip shape generation process into multiple components: obtaining target audio data, obtaining first video data, performing mask processing to obtain second video data, extracting features from both video data sources, and finally predicting lip areas. This segmentation allows each component to be optimized independently, ensuring both accuracy and temporal smoothness.
Solution Approach 2:
The patent introduces an intermediary mechanism by using both first video data and second video data (with mask processing) as intermediate representations. These intermediaries capture different aspects of lip movement and are combined through feature extraction to produce smooth transitions while maintaining generation accuracy.
2Measurement precision
If mask processing is performed on lip area in video data, then lip area can be isolated for processing, but additional processing steps are required
Solution Approach 1:
The patent applies mask processing as a preliminary action to obtain second video data with isolated lip areas before the main prediction process. This preliminary segmentation improves the precision of lip area identification and enables more accurate feature extraction, justifying the additional processing step.
3Reliability
If synchronization alignment training is performed on paired sample audio and sample video, then audio-video synchronization can be improved, but training complexity increases
Solution Approach 1:
The patent performs synchronization alignment training as a preliminary action during the model training phase, before actual video generation. This preliminary training ensures that the target multimodal model learns to align audio and video features effectively, improving synchronization reliability in the generated videos while the actual generation process remains relatively simple.
Data Source
AI summary
The present disclosure relates to the technical field of video processing, and discloses a video generation method, apparatus, device, medium and program product. The method includes: acquiring target audio data and first video data of a target object; acquiring second video data, the second video data is obtained by performing mask processing on a lip area in video data of the target object; performing feature processing on the target audio data based on a target multimodal model to obtain a target audio feature; performing feature extraction on the first video data and the second video data to obtain a feature to be processed; and predicting a lip area in the second video data based on the target audio feature and the feature to be processed, to determine a target video corresponding to the target audio data.


