Video and picture generation method and system based on improved VAE

By introducing perceptual loss, GAN discriminator and time series similarity loss in VAE, the shortcomings of VAE in generating high-quality images and videos are solved, and better visual perceptual quality and timing consistency are achieved.

CN120147667APending Publication Date: 2025-06-13DATA TRANSMISSION GRP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510273785.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the image and video generation tasks, existing VAEs are difficult to generate images and videos with good visual perception quality, especially in dealing with complex textures and timing consistency.

Method used

Improve the generative model of VAE by introducing perceptual loss, GAN discriminator, and time series similarity loss. Perceptual loss extracts advanced features of images through pre-trained deep convolutional neural networks. GAN discriminator is used to judge the authenticity of the generated images, and time series similarity loss ensures timing consistency between video frames.

Benefits of technology

The quality of generated images and videos is significantly improved, ensuring that the generated results meet high standards in advanced features, authenticity and timing consistency, and the model training and inference efficiency are high.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147667A_ABST
    Figure CN120147667A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video and picture generation, and discloses a video and picture generation method and system based on improved VAE, the video and picture generation method based on improved VAE comprises the following steps: S1, sensing loss: using a pre-trained deep convolutional neural network to extract advanced features of an image, the perception loss is calculated by comparing the features; s2, a GAN discriminator: introducing a small network as a discriminator for judging whether the generated picture is true or false; and S3, time sequence similarity loss: additionally introducing a time sequence module to ensure the consistency of the generated video and the original video in time sequence. The method is reasonable in design, and the quality of generated images and videos is effectively improved by introducing the perception loss, the GAN discriminator and the time sequence similarity loss. By sensing the loss, the consistency of the generated image in the aspect of advanced features can be ensured; through the GAN discriminator, the generation effect of the VAE can be optimized, and the generated image is more real.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video and image generation, and particularly to a method and system for video and image generation based on an improved VAE. Background Art

[0002] The Variational Autoencoder (VAE) is a generative model widely used in tasks such as image generation and video generation. The core idea of the VAE is to generate new data samples by learning the latent representation of the data. In image generation, the VAE maps the input image to the latent space through an encoder and then reconstructs the image from this latent space through a decoder. Different from traditional autoencoders, the VAE learns the probability distribution of the latent space rather than a single deterministic mapping, which enables it to generate more diverse and variable images.

[0003] Many mainstream generative algorithms, such as SD, SDXL, SD3, etc., use the VAE as a key part for generating images. In video generation tasks, the application of the VAE is also very extensive, especially when dealing with temporal data, and usually 3DVAE is used at this time. Algorithms such as Sora and Cogvideox will directly compress the video into the features of the latent space. Considering the information of the video in space (height, width) and time (time dimension), the encoder of the VAE not only needs to capture image features but also learn temporal relationships; the decoder is responsible for generating new video frames from the latent space and ensuring the temporal consistency between these frames.

[0004] As an important part in image and video generation, the VAE plays a key role in generating high-quality and diverse samples. By learning the latent representation of the data, the VAE can deeply understand the distribution characteristics of the data and generate new samples based on this. In video generation, the VAE is particularly good at dealing with temporal data, thus ensuring the coherence and consistency between video frames. Therefore, the quality of the VAE directly affects the quality of the generated images and videos and becomes an indispensable part of the generation task; therefore, we propose a method and system for video and image generation based on an improved VAE to solve this problem. Summary of the Invention

[0005] The purpose of the present invention is to solve the above-mentioned disadvantages mentioned in the background art, and to propose a method and system for video and image generation based on an improved VAE.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] A method for video and image generation based on an improved VAE includes the following steps:

[0008] S1. Perceptual loss: Use a pre-trained deep convolutional neural network to extract high-level features of the image, and calculate the perceptual loss by comparing these features;

[0009] S2. GAN discriminator: Introduce a small network as the discriminator to judge whether the generated picture is real or fake;

[0010] S3. Time series similarity loss: Additionally introduce a time series module to ensure the temporal consistency between the generated video and the original video.

[0011] Preferably, in S1, the specific steps are as follows:

[0012] S11. Feature extraction: Use a pre-trained convolutional neural network to extract the features of the input image and the generated image respectively;

[0013] S12. Extraction layer selection: Select the middle layer of the network as the feature extraction layer;

[0014] S13. Function introduction: Introduce the perceptual loss function;

[0015] S14. Loss calculation: Calculate the feature difference between the input image and the generated image at the feature extraction layer. Usually, the L1 or L2 norm is used, and the feature difference is used as part of the loss function to optimize the VAE generation model.

[0016] Preferably, in S2, the specific steps are as follows:

[0017] S21. Discriminator design: When real data is input into the discriminator, it is expected that the discriminator judges it as true; when the generation of VAE is input into the discriminator, it is expected that the discriminator judges it as false;

[0018] S22. Joint training: The discriminator network and the VAE network are trained together, and the parameters are updated together.

[0019] Preferably, in S3, the specific steps are as follows:

[0020] S31. Time series module design: The time series module operates on the input video and the output video in the time dimension to capture the feature changes in time;

[0021] S32. Loss calculation: The loss function in the time dimension is the cosine similarity of the two features. The closer the similarity is to 1, the more similar the features of the input video and the output video are in the time dimension. Inject this loss function into the 3D VAE to improve the VAE generation effect;

[0022] Preferably, in S1, the pre-trained deep convolutional neural network includes but is not limited to VGG and ResNe.

[0023] Preferably, in S12, when selecting the feature extraction layer, select the features at the model resolution change point, such as the features after the third downsampling.

[0024] Preferably, in S31, the temporal module includes a recurrent neural network (RNN) and a convolutional neural network (CNN);

[0025] The recurrent neural network (RNN) adopts a bidirectional LSTM structure with a hidden layer dimension of 256, which is used to capture long-term temporal dependencies;

[0026] The convolutional neural network (CNN) consists of 3 3D convolutional layers with a convolutional kernel size of 3×3×3 and the number of channels being 64, 128, and 256 in sequence, which is used to extract local spatio-temporal features.

[0027] The present invention also provides an improved VAE image generation system, including:

[0028] An input module, responsible for receiving the input data of the image to be generated, converting the input original image data into a format suitable for network processing, and passing it to the encoder for subsequent processing;

[0029] An encoder module, used to map the input image to the latent space;

[0030] A perceptual loss module, based on a pre-trained convolutional neural network to extract the high-level features of the image and calculate the perceptual loss between the input image and the generated image;

[0031] A GAN discriminator module, used to judge the authenticity of the generated image;

[0032] A decoder module, which decodes the latent space representation generated by the encoder into a generated image;

[0033] An output module, responsible for converting the image generated by the decoder into the final output and providing it to the user or other systems for use;

[0034] A temporal module, used to process the temporal consistency between video frames;

[0035] An optimization module, used to coordinate and optimize the work of each module, combine objectives such as perceptual loss, GAN discriminator, and temporal loss to form a unified optimization objective, and use optimization algorithms such as gradient descent to train the entire system, and finally obtain a VAE model that generates high-quality images;

[0036] A training module, responsible for managing the training process of the model.

[0037] Preferably, the encoder module adopts an improved VAE structure, where the representation in the latent space is no longer a single deterministic vector but is represented by a probability distribution, allowing the model to generate more diverse images. Through optimized parameters, the encoder can extract high-level features of the input image and compress them into the latent space.

[0038] Preferably, the decoder module can restore the latent vector into an output image similar to the input image, and the decoder module consists of a multi-layer convolutional neural network (CNN).

[0039] Preferably, the training module includes a data loading unit, a model parameter initialization unit, a loss calculation unit, and a gradient update unit. By jointly training the VAE and GAN discriminator networks, and guided by perceptual loss and temporal loss, the model parameters can be effectively optimized to improve the quality of the generated images.

[0040] Compared with the prior art, the present invention provides a method and system for generating videos and pictures based on improved VAE, having the following beneficial effects:

[0041] (1) By introducing multiple additional loss functions, the learning ability of VAE is improved to ensure that the generated images or videos have good visual perception quality, with complex textures and shapes;

[0042] (2) By introducing perceptual loss, GAN discriminator, and time series similarity loss, the quality of the generated images and videos is effectively improved. Through perceptual loss, the consistency of the generated images in high-level features can be ensured; through the GAN discriminator, the generation effect of VAE can be optimized to make the generated images more realistic; through time series similarity loss, the temporal consistency of the generated videos can be guaranteed;

[0043] (3) The weighted summation strategy of multiple loss functions makes the model more stable during training, avoiding overfitting or underfitting problems that may be caused by a single loss function. It can be applied to different image and video generation tasks, with strong adaptability and scalability. By adjusting the weight coefficients of the loss functions and the network structure, different styles and types of image and video generation can be achieved.

[0044] (4) Although multiple additional loss functions are introduced, the present invention has high efficiency during training and inference. Through reasonable model design and selection of optimization algorithms, good generation efficiency can be obtained in a short time. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a flowchart of a method for generating videos and pictures based on improved VAE proposed by the present invention. DETAILED DESCRIPTION

[0046] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0047] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0048] Referring to Figure 1 , a method for generating videos and pictures based on an improved VAE, includes the following steps:

[0049] S1. Perceptual loss: Use a pre-trained deep convolutional neural network to extract high-level features of the image, and calculate the perceptual loss by comparing these features;

[0050] S2. GAN discriminator: Introduce a small network as a discriminator to judge whether the generated picture is real or fake;

[0051] S3. Temporal sequence similarity loss: Additionally introduce a temporal module to ensure the consistency of the generated video and the original video in terms of temporal sequence.

[0052] In this embodiment, in S1, the specific steps are as follows:

[0053] S11. Feature extraction: Use a pre-trained convolutional neural network to extract the features of the input image and the generated image respectively;

[0054] S12. Extraction layer selection: Select the middle layer of the network as the feature extraction layer;

[0055] S13. Function introduction: Introduce a perceptual loss function;

[0056] S14. Loss calculation: Calculate the feature difference between the input image and the generated image at the feature extraction layer. Usually, the L1 or L2 norm is used, and the feature difference is used as a part of the loss function to optimize the VAE generation model.

[0057] In this embodiment, in S2, the specific steps are as follows:

[0058] S21. Discriminator design: When real data is input into the discriminator, it is expected that the discriminator will judge it as true; when the generation of VAE is input into the discriminator, it is expected that the discriminator will judge it as false;

[0059] S22. Joint Training: The discriminator network and the VAE network are trained together, and the parameters are updated together.

[0060] In this embodiment, in S3, the specific steps are as follows:

[0061] S31. Temporal Module Design: The temporal module operates on the input video and the output video in the time dimension to capture the feature changes in time.

[0062] S32. Loss Calculation: The loss function in the time dimension is the cosine similarity of the two features. The closer the similarity is to 1, the more similar the features of the input video and the output video are in the time dimension. This loss function is injected into the 3D VAE to improve the VAE generation effect.

[0063] In this embodiment, in S1, the pre-trained deep convolutional neural network includes, but is not limited to, VGG and ResNe.

[0064] In this embodiment, in S12, when selecting the feature extraction layer, the features at the model resolution change are selected, such as the features after the third downsampling.

[0065] In this embodiment, in S31, the temporal module includes a recurrent neural network (RNN) and a convolutional neural network (CNN);

[0066] The recurrent neural network (RNN) adopts a bidirectional LSTM structure with a hidden layer dimension of 256, which is used to capture long-term temporal dependencies;

[0067] The convolutional neural network (CNN) consists of 3 3D convolutional layers with a convolutional kernel size of 3×3×3 and the number of channels being 64, 128, and 256 in sequence, which is used to extract local spatio-temporal features.

[0068] The present invention also provides an improved VAE-based image generation system, including:

[0069] An input module, responsible for receiving the input data of the image to be generated, converting the input original image data into a format suitable for network processing, and passing it to the encoder for subsequent processing;

[0070] An encoder module, used to map the input image to the latent space;

[0071] A perceptual loss module, which extracts the high-level features of the image based on the pre-trained convolutional neural network and calculates the perceptual loss between the input image and the generated image;

[0072] A GAN discriminator module, used to judge the authenticity of the generated image;

[0073] A decoder module, which decodes the latent space representation generated by the encoder into a generated image;

[0074] An output module, responsible for converting the images generated by the decoder into the final output and providing it to users or other systems for use;

[0075] A timing module, used to handle the temporal consistency between video frames;

[0076] An optimization module, used to coordinate and optimize the work of each module, combine objectives such as perceptual loss, GAN discriminator, and temporal loss to form a unified optimization objective, and use optimization algorithms such as gradient descent to train the entire system, ultimately obtaining a VAE model that generates high-quality images;

[0077] A training module, responsible for managing the training process of the model.

[0078] In this embodiment, the encoder module adopts an improved VAE structure, where the representation in the latent space is no longer a single deterministic vector but is represented by a probability distribution, allowing the model to generate more diverse images. Through optimized parameters, the encoder can extract high-level features of the input image and compress them into the latent space.

[0079] In this embodiment, the decoder module can restore the latent vector into an output image similar to the input image, and the decoder module consists of a multi-layer convolutional neural network (CNN).

[0080] In this embodiment, the training module includes a data loading unit, a model parameter initialization unit, a loss calculation unit, and a gradient update unit. The training module can effectively optimize the model parameters and improve the quality of the generated images by jointly training the VAE and GAN discriminator networks and being guided by perceptual loss and temporal loss.

[0081] In this embodiment, during use, a large-scale image and video dataset is collected for training and testing the model of the present invention. The dataset should cover a variety of scenarios and objects to ensure the generalization ability of the model. Initialize the parameters of the VAE model, including the weights of the encoder and decoder, and use stochastic gradient descent (SGD) or other optimization algorithms to train the model according to the total loss function. During the training process, update the model parameters according to the set learning rate, regularly evaluate the performance of the model on the validation set, and adjust hyperparameters such as the learning rate and batch size according to the evaluation results. After training is completed, save the best parameters of the model.

[0082] The standard parts used in the present invention can all be purchased from the market. The special-shaped parts can be customized according to the description in the specification and the drawings. The specific connection methods of each part all adopt conventional means such as bolts, rivets, and welding that are mature in the prior art. The machines, parts, and equipment all adopt conventional models in the prior art, and the circuit connections adopt conventional connection methods in the prior art, which will not be elaborated here.

Claims

1. A video and picture generation method based on improved VAE, characterized in that: The steps include: S1. Perceptual loss: Use a pre-trained deep convolutional neural network to extract high-level features of the image and calculate the perceptual loss by comparing these features; S2, GAN discriminator: introduce a small network as a discriminator to determine whether the generated image is real or fake; S3, time series similarity loss: an additional timing module is introduced to ensure the temporal consistency between the generated video and the original video.

2. The video and picture generation method based on improved VAE according to claim 1, characterized in that: In S1, the specific steps are as follows: S11, feature extraction: use the pre-trained convolutional neural network to extract the features of the input image and the generated image respectively; S12, extraction layer selection: select the middle layer of the network as the feature extraction layer; S13, function introduction: introduce the perceptual loss function; S14. Loss calculation: Calculate the feature difference between the input image and the generated image at the feature extraction layer, usually using the L1 or L2 norm, and use the feature difference as part of the loss function to optimize the VAE generation model.

3. The video and picture generation method based on improved VAE according to claim 1, characterized in that: In S2, the specific steps are as follows: S21. Discriminator design: When real data is input into the discriminator, we hope that the discriminator will judge it as true; when the VAE generation is used as the discriminator input, we hope that the discriminator will judge it as false; S22. Joint training: The discriminator network and the VAE network are trained together, and their parameters are updated together.

4. The video and picture generation method based on improved VAE according to claim 1, characterized in that: In S3, the specific steps are as follows: S31. Timing module design: The timing module operates on the input video and output video in the time dimension to capture the feature changes in time; S32. Loss calculation: The loss function in the time dimension is the cosine similarity of the two features. The closer the similarity is to 1, the more similar the features of the input video and the output video in the time dimension are. This loss function is injected into the 3D VAE to improve the VAE generation effect.

5. The video and picture generation method based on improved VAE according to claim 1, characterized in that: In S1, the pre-trained deep convolutional neural network includes but is not limited to VGG and ResNe.

6. The video and picture generation method based on improved VAE according to claim 1, characterized in that: In S12, when selecting the feature extraction layer, the features at the location where the model resolution changes are selected, such as the features after the third downsampling.

7. The video and picture generation method based on improved VAE according to claim 1, characterized in that: In S31, the timing module includes a recurrent neural network (RNN) and a convolutional neural network (CNN); The recurrent neural network (RNN) adopts a bidirectional LSTM structure with a hidden layer dimension of 256 to capture long-term temporal dependencies; The convolutional neural network (CNN) consists of three 3D convolutional layers with a convolution kernel size of 3×3×3 and channel numbers of 64, 128, and 256, respectively, for extracting local spatiotemporal features.

8. A system for generating images based on an improved VAE, characterized in that: include: The input module is responsible for receiving the input data of the image to be generated, converting the input raw image data into a format suitable for network processing, and passing it to the encoder for subsequent processing; The encoder module maps the input image to the latent space; Perceptual loss module, which extracts high-level features of images based on pre-trained convolutional neural networks and calculates the perceptual loss between input images and generated images; The GAN discriminator module is used to judge the authenticity of the generated images; The decoder module decodes the latent space representation generated by the encoder into a generated image; The output module is responsible for converting the image generated by the decoder into the final output and providing it to users or other systems; Timing module, used to process the timing consistency between video frames; The optimization module is used to coordinate and optimize the work of each module, combine the perceptual loss, GAN discriminator, and temporal loss objectives to form a unified optimization goal, and use optimization algorithms such as gradient descent to train the entire system, ultimately obtaining a VAE model that generates high-quality images; The training module is responsible for managing the model training process.

9. The video and picture generation method based on improved VAE according to claim 1, characterized in that: The encoder module adopts an improved VAE structure, in which the representation of the latent space is no longer a single deterministic vector, but is represented by a probability distribution, allowing the model to generate more diverse images. Through optimized parameters, the encoder is able to extract high-level features of the input image and compress it into the latent space.

10. The video and picture generation method based on improved VAE according to claim 1, characterized in that: The decoder module can restore the latent vector to an output image similar to the input image. The decoder module consists of a multi-layer convolutional neural network (CNN). The training module includes a data loading unit, a model parameter initialization unit, a loss calculation unit and a gradient update unit. The training module can effectively optimize the model parameters and improve the quality of generated images by jointly training the VAE and GAN discriminator networks and guiding them through perceptual loss and temporal loss.