Digital Human Video Generation With Integrated Downsampling And Upsampling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech-to-lip conversion models, such as the wav2lip model, produce low-resolution blurred images, resulting in poor visual quality of digital human voiceover videos, and the use of super-resolution models for enhancement is inefficient and cannot meet real-time requirements.
Innovation Solution
A video generation method and apparatus that includes a video generation model with down-sampling and up-sampling layers to process high-resolution images without additional image processing models, enhancing the resolution and efficiency of digital human voiceover videos.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If super-resolution models are used to enhance the output of wav2lip model, then the resolution of digital human voiceover video is improved, but the generation efficiency deteriorates and real-time requirements cannot be met
Solution Approach 1:
The video generation model performs down-sampling on input images before processing, preparing the data in advance for efficient generation. This preliminary preparation allows the model to work with standardized dimensions, improving processing speed while maintaining output quality through subsequent up-sampling operations.
Solution Approach 2:
The model dynamically adjusts image resolution parameters by down-sampling input images to standardized dimensions (e.g., 96x96 or 192x192) during processing, then up-sampling the generated video frames to the target resolution. This parameter transformation approach enables efficient processing without sacrificing final video quality.
2Manufacturing precision
If another image processing model is added to process high-resolution images, then the visual quality is improved, but the device complexity increases
Solution Approach 1:
The video generation model is designed to handle multiple resolution requirements within a single unified architecture. By incorporating down-sampling and up-sampling capabilities, the model can process both low-resolution input images and generate high-resolution video frames, eliminating the need for separate specialized processing models.
Solution Approach 2:
The patent combines the image processing and video generation functions into a single integrated video generation model. The model merges down-sampling, feature extraction, video frame generation, and up-sampling operations into one cohesive system, reducing overall system complexity while maintaining high output quality.
Data Source
AI summary
The present disclosure relates to a video generation method, a readable medium, and an electronic device. The video generation method includes: obtaining a talking video of a target object and a target text for video generation; and generating, by using the talking video, the target text, and a video generation model, a target video of a digital human corresponding to the target object talking according to the target text.


