Video generation method, apparatus and electronic device

Through sparsely distributed edge line diagrams and video generation models, combined with diffusion model and adaptation module, the video generation problems caused by high image quality requirements in the prior art are solved, and high-quality and rich content video generation is achieved.

CN119741222BActive Publication Date: 2025-06-17BEIJING SHENGSHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411938868.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-06-17
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

The existing image-generating video models have high requirements for image quality, which makes it difficult to generate rich and coherent videos when the image content is single or small, which limits the quality of the video finished products of the creator.

Method used

The sparsely distributed edge line diagram is used to generate videos. By obtaining the P-frame edge detection diagram and the Q-frame placeholder map, the pre-trained video generation model is input, and the target video is generated by combining the diffusion model and the adaptation module.

Benefits of technology

It realizes the generation of high-quality videos based on a small number of edge detection images, meets the video production needs of users and improves the richness and coherence of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741222B_ABST
    Figure CN119741222B_ABST
Patent Text Reader

Abstract

The present application relates to a video generation method, apparatus and electronic device. The method includes: obtaining a video segment to be processed, where the video segment includes a P-frame edge detection map and a Q-frame placeholder map, where P≥1 and Q≥0; inputting condition information into a pre-trained video generation model to generate a target video, and the condition information at least includes the video segment to be processed, and the target video includes visual content corresponding to the P-frame edge detection map; where the video generation model includes a diffusion model and an adaptation module, the diffusion model includes multiple layers of downsampling hidden layers, the adaptation module includes multiple layers of fine-tuning hidden layers, and the fine-tuning vector output by a single layer of fine-tuning hidden layer is used to be superimposed on the output vector of the mapped downsampling hidden layer to serve as the input vector of the next layer of downsampling hidden layer. The solution provided by the present application can generate a high-quality video finished product using a sparsely distributed edge line map, meeting the user's creation needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision processing technology, and particularly to a video generation method, apparatus, and electronic device. Background Art

[0002] With the rapid development of artificial intelligence technology, AIGC (Artificial Intelligence Generated Content) has been widely applied in various fields. Among them, the application of generating videos based on videos is booming.

[0003] However, the current image-to-video model has relatively high requirements for images. If the image content is single or the number of images is small, it is difficult to generate a video with rich content, or the video content is incoherent or unreasonable, thus restricting the quality of the video finished products of creators. Summary of the Invention

[0004] To solve or partially solve the problems existing in the related art, this application provides a video generation method, apparatus, and electronic device, which can generate rich video finished products using sparse distributed edge line graphs and meet the creative needs of users.

[0005] The first aspect of this application provides a video generation method, which includes:

[0006] Obtain a video segment to be processed, where the video segment includes a P-frame edge detection map and a Q-frame placeholder map, where P≥1 and Q≥0;

[0007] Input condition information into a pre-trained video generation model to generate a target video, where the condition information at least includes the video segment to be processed, and the target video includes the visual content corresponding to the P-frame edge detection map;

[0008] Among them, the video generation model includes a diffusion model and an adaptation module. The diffusion model includes multiple layers of downsampling hidden layers, and the adaptation module includes multiple layers of fine-tuning hidden layers. The fine-tuning vector output by a single layer of fine-tuning hidden layer is used to be superimposed with the output vector of the mapped downsampling hidden layer as the input vector of the next layer of downsampling hidden layer.

[0009] In some embodiments, the number of fine-tuning hidden layers of the adaptation module is less than or equal to the number of downsampling hidden layers of the encoder and corresponds layer by layer; the vector dimension of the fine-tuning hidden layer is the same as the vector dimension of the mapped downsampling hidden layer.

[0010] In some embodiments, the obtaining of the video segment to be processed includes:

[0011] Obtain the original video segment or the original sample image of the P-frame; generate an edge detection map from the P-frame original image in the original video segment, or generate an edge detection map from the P-frame original sample image; generate the video segment to be processed according to the edge detection map of the P-frame and the placeholder map of the Q-frame.

[0012] In some embodiments, the placeholder map is a solid-color map; preferably, the placeholder map is a pure black map.

[0013] In some embodiments, the conditional information further includes a prompt text;

[0014] Before inputting the conditional information into the pre-trained video generation model to generate the target video, it further includes: for the expected target, obtaining the prompt text based on the custom design and / or according to the visual content of the original video segment or the original sample image.

[0015] In some embodiments, the video generation model is obtained by training according to the following method:

[0016] Obtain multiple groups of training data; wherein, each group of the training data includes a first training video, a second training video and the corresponding text annotation, the first training video includes K-frame edge detection maps, the second training video includes a training video of P-frame edge detection maps and Q-frame placeholder maps and the corresponding text annotation, P≥1, Q≥0, P + Q = K;

[0017] In each round of training, input the second training video into the adaptation module, and save the fine-tuning vectors output by each layer of the fine-tuning hidden layer respectively;

[0018] Input the random noise, the first training video and the text annotation into the diffusion model to obtain the predicted noise image output by the diffusion model; wherein, the output vector of each layer of the downsampling hidden layer in the diffusion model is superimposed with the fine-tuning vector of the mapped fine-tuning hidden layer as the input vector of the next layer of the downsampling hidden layer;

[0019] Fix the network parameters of each layer of the diffusion model, and backpropagate according to the loss function of the diffusion model to adjust the network parameters of each layer of the fine-tuning hidden layer of the adaptation module to obtain the trained video generation model.

[0020] In some embodiments, the method further includes:

[0021] Before training the adaptation module, use the network parameters of each layer of the downsampling hidden layer as the initialization parameters of each layer of the mapped fine-tuning hidden layer in the adaptation module.

[0022] The second aspect of the present application provides a video generation device, which includes:

[0023] A video acquisition module, configured to acquire a video segment to be processed, where the video segment includes a P-frame edge detection map and a Q-frame placeholder map, where P≥1 and Q≥0;

[0024] A video generation module, configured to input condition information into a pre-trained video generation model to generate a target video, where the condition information at least includes the video segment to be processed, and the target video includes visual content corresponding to the P-frame edge detection map;

[0025] Wherein, the video generation model includes a diffusion model and an adaptation module, the diffusion model includes multiple layers of downsampling hidden layers, the adaptation module includes multiple layers of fine-tuning hidden layers, and the fine-tuning vector output by a single layer of fine-tuning hidden layer is used to be superimposed with the output vector of the mapped downsampling hidden layer as the input vector of the next layer of downsampling hidden layer.

[0026] A third aspect of the present application provides an electronic device, including:

[0027] A processor; and

[0028] A memory, on which executable code is stored, and when the executable code is executed by the processor, the processor is caused to execute the video generation method described in the first aspect above.

[0029] A fourth aspect of the present application provides a computer-readable storage medium, on which executable code is stored, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the video generation method described in the first aspect above.

[0030] A fifth aspect of the present application provides a computer program product, including a computer program, characterized in that the computer program is used to execute computer program code instructions corresponding to the video generation method described in the first aspect above.

[0031] The technical solution provided by the present application may include the following beneficial effects:

[0032] In the video generation method of the present application, by using a diffusion model with an adaptation module as a new video generation model, based on a small number of edge detection maps, on the basis of accurately capturing as many edges in the image as possible, the video content can be restored or reconstructed to generate a high-quality target video, meeting the user's video production requirements.

[0033] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings

[0034] The above and other objects, features, and advantages of the present application will become more apparent by describing exemplary embodiments of the present application in more detail with reference to the accompanying drawings. In the exemplary embodiments of the present application, the same reference numerals generally represent the same components.

[0035] Figure 1 is a schematic flowchart of a video generation method provided by an embodiment of the present application;

[0036] Figure 2 is a schematic structural diagram of an adaptation module and a diffusion model provided by an embodiment of the present application;

[0037] Figure 3 is a schematic flowchart of a training method of a video generation model provided by an embodiment of the present application;

[0038] Figure 4 is another schematic flowchart of a video generation method provided by an embodiment of the present application;

[0039] Figure 5 is a schematic structural diagram of a video generation device provided by an embodiment of the present application;

[0040] Figure 6 is another schematic structural diagram of a video generation device provided by an embodiment of the present application;

[0041] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0042] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0043] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0044] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, the meaning of "a plurality" is two or more, unless otherwise specifically defined.

[0045] In the related art, in the current AIGC field, users are limited to some high-quality image-to-video generation, with high requirements for the quality of the input images, which affects the creative needs of creators.

[0046] To address the above problems, this application provides a video generation method that can generate high-quality video finished products using sparsely distributed edge line drawings to meet the creative needs of users.

[0047] The technical solution of this application will be described in detail below with reference to the accompanying drawings.

[0048] See Figure 1 , a video generation method shown in this application includes:

[0049] S110, obtaining a video segment to be processed, where the video segment includes a P-frame edge detection map and a Q-frame placeholder map; where P≥1 and Q≥0.

[0050] In this application, the edge detection map can be an image obtained by processing the original image through an edge detection algorithm (such as the Canny algorithm); in a video segment, the frame numbers corresponding to multiple frames of edge detection maps can be consecutive or spaced apart, that is, the arrangement positions of multiple frames of edge detection maps in the video segment can be customized. Except for the edge detection map, other images in the video segment are placeholder maps. Among them, the edge detection map includes the edge information in the original image, which is convenient for accurately identifying object objects in the image. Therefore, on the basis of accurately capturing as many edges in the image as possible, more accurate information can be provided to the subsequent video generation model, and a better-targeted video can be generated. A placeholder map refers to an image that does not contain visual content. On the basis of combining the edge detection map, the subsequent video generation model can create based on the edge information and generate new images related to the edge information to enrich the visual content. In some embodiments, each frame of the placeholder map can be represented by a solid-color map. Different placeholder maps can be represented by different colors, or can be uniformly represented by the same color. Preferably, the placeholder map can be a pure black map. It can be understood that taking the gray value as an example, the gray value of each pixel of the solid-color map is the same; for example, the gray value of each pixel of the pure black map is 0.

[0051] S120. Input the conditional information into a pre-trained video generation model to generate a target video. The conditional information includes at least the video segment to be processed, and the visual content corresponding to the P-frame edge detection map is included in the target video.

[0052] Among them, the video generation model includes a diffusion model and an adaptation module. The diffusion model includes multiple layers of downsampling hidden layers and multiple layers of upsampling hidden layers. The adaptation module has multiple layers of fine-tuning hidden layers. The multiple layers of fine-tuning hidden layers of the adaptation module are respectively mapped one-to-one with the multiple layers of downsampling hidden layers of the diffusion model, and the structure of each layer of the fine-tuning hidden layer is the same as that of the corresponding downsampling hidden layer. The fine-tuning vector output by a single layer of the fine-tuning hidden layer of the adaptation module is used to be superimposed with the output vector of the downsampling hidden layer of the mapped diffusion model as the input vector of the next layer of the downsampling hidden layer.

[0053] The pre-trained video generation model is obtained by jointly training the diffusion model and the adaptation module. Optionally, during the joint training process, first: obtain a pre-trained diffusion model and an adaptation module to be trained. Among them, the pre-trained diffusion model has the ability to generate videos, but its video generation effect for a specific scenario is poor, and the network parameter magnitude of the pre-trained diffusion model is relatively large, such as in the order of billions or tens of billions. If the parameters of the pre-trained diffusion model are directly re-trained and adjusted for a specific scenario, it will not only affect the generality of the pre-trained diffusion model, but also require a very large amount of training resources and a long cycle. The adaptation module includes multiple layers of fine-tuning hidden layers, and the diffusion model includes multiple layers of downsampling hidden layers (i.e., the input layer) and multiple layers of upsampling hidden layers (i.e., the output layer). In this embodiment, the multiple layers of fine-tuning hidden layers of the adaptation module to be trained have the same layer structure as the multiple layers of downsampling hidden layers of the diffusion model, that is, the number of network layers of the adaptation module is the same as the number of network layers of the input layer of the diffusion model, and the network parameters of each network layer of the adaptation module are the same as those of the corresponding network layer of the input layer of the diffusion model. Second: fix the network parameters of each layer of the diffusion model, and adjust the network parameters of each fine-tuning hidden layer of the adaptation module according to the backpropagation of the loss function of the diffusion model, so as to obtain the pre-trained video generation model.

[0054] In the embodiments of the present invention, a pre-trained video generation model is obtained through the joint training of a diffusion model and an adaptation module. Without directly adjusting the parameters of the diffusion model, it can not only reduce the number of parameters required for fine-tuning, but also reduce the demand for computing resources during the training process and accelerate the training speed. At the same time, when the diffusion model performs a specific task, the vectors of the input layer of the diffusion model are fine-tuned through the pre-trained adaptation module, so as to achieve the efficient and accurate completion of the specific task without the need for large-scale overall training of the diffusion model. It can be seen that through the joint training of the diffusion model and the adaptation module, a pre-trained video generation model is obtained, achieving the goal of quickly generating a high-quality target video with visual content of the acquired video segments to be processed based on the AIGC technology, and moreover, it will not affect the generality of the diffusion model.

[0055] In the embodiments of the present invention, the diffusion model is a fusion architecture of Transformer and Diffusion, that is, after mapping the input data to the latent space representation through the encoder, the latent space representation of the target output corresponding to the input data is obtained through the process of adding noise and denoising, and the latent space representation of the target output is mapped to the corresponding data space through the decoder, and finally the target output is obtained. Specifically, the Transformer model architecture is a deep learning architecture that relies on the self-attention mechanism and is essentially composed of an encoder and a decoder.

[0056] Furthermore, the diffusion model is a generative model based on deep learning algorithms. It generates high-quality content through a process of gradually adding noise and denoising, and performs excellently in generating images, videos, audio, and other high-dimensional data, as well as being powerful in generative tasks. In this application, the diffusion model is a pre-trained model. After the above encoder maps the edge detection map to the latent space representation, guided by the latent space representation of the edge detection map, the random noise is denoised for T time steps, that is, downsampling and upsampling are performed at each time step to obtain the latent space representation of the target video. Then, the latent space representation of the target video is further mapped to the pixel space through the above decoder to obtain the target video. Among them, the latent space representation of the edge detection map is the latent space representation form obtained by encoding the edge detection map using the encoder. The encoder is responsible for mapping data from the original data space to the latent space, obtaining an implicit and continuous representation, which can make the model calculation more efficient; while the decoder can reconstruct the video from this latent space representation, mapping the generation result of the diffusion model from the latent space to the pixel space. In some embodiments, both the encoder and decoder used in the diffusion model of this application can be variational autoencoders (VAE, Variational Autoencoder). The latent space of VAE is continuous and suitable for generating continuous data, that is, making the generated target videos or images related to each other.

[0057] Exemplarily, such as Figure 2As shown, the core module in the diffusion model is the U-Net network. U-Net is a structure of a downsampling module - upsampling module with skip connections, used to denoise random noise and achieve the generation from noise to video or image. Among them, the downsampling module is used to gradually downsample the input noise, extract features and reduce the spatial resolution of the image. The upsampling module gradually restores the spatial resolution of the image through transposed convolution (or inverse convolution) and upsampling operations; through skip connections, the upsampling module can combine the low-level features (such as details like edges and textures) in the downsampling module with high-level features (such as semantic information of the prompt text). The fusion of features helps to reconstruct the details of the image; finally, a target image with the same resolution as the input image is output. In this application, based on the multi-layer downsampling hidden layer structure in the downsampling module of the diffusion model, an independent adaptation module is additionally designed to fine-tune the vectors generated during the downsampling process. That is to say, the adaptation module adapts to new tasks by inserting additional parameters into the pre-trained diffusion model, rather than directly adjusting the parameters of the diffusion model (the parameters of the diffusion model are on the order of billions or tens of billions), which can not only reduce the number of parameters required for fine-tuning, but also reduce the demand for computing resources during the training process and speed up the training speed. At the same time, when the diffusion model performs a specific task, through the trained adaptation module, the output vectors of each layer of the downsampling hidden layer are fine-tuned, so that it is possible to efficiently and accurately complete the specific task without large-scale overall training of the diffusion model. It can be seen that through the combination of the adaptation module and the diffusion model, based on the AIGC technology, a high-quality target video can be quickly generated through a video segment including an edge detection map, and the generality of the diffusion model will not be affected.

[0058] See also Figure 2 , the adaptation module has multiple fine-tuning hidden layers arranged in sequence, and the diffusion model includes multiple downsampling hidden layers arranged in sequence. In some embodiments, the number of layers of the fine-tuning hidden layer of the adaptation module is less than or equal to the number of layers of the downsampling hidden layer and corresponds layer by layer. For example, the adaptation module has M layers of fine-tuning hidden layers, and the downsampling module has N layers of downsampling hidden layers, then 1 ≤ M ≤ N, and both M and N are positive integers. Specifically, when M is equal to N, each layer of the fine-tuning hidden layer is mapped one-to-one with the downsampling hidden layer of the same level. When M is less than N, each layer of the fine-tuning hidden layer can be mapped one-to-one with the downsampling hidden layer of the same level, or can be mapped one-to-one layer by layer with the downsampling hidden layers of different levels, such as mapping with the downsampling hidden layers corresponding to odd layers, even layers, specified layers, or random layers. Such a design enables the input vector of at least one layer of the downsampling hidden layer to be fine-tuned by superimposing the fine-tuning vector, and then the corresponding output vector is changed, so that the general diffusion model is affected by the adaptation module for specific tasks and is more inclined to process specific tasks to promote the achievement of the expected goal.

[0059] Furthermore, in some embodiments, the vector dimension of the fine-tuning hidden layer is the same as that of the corresponding hierarchical downsampling hidden layer. Such a design facilitates the superposition of the fine-tuning vector output by the fine-tuning hidden layer and the output vector of the corresponding downsampling hidden layer in the same dimension, so as to achieve a comprehensive fine-tuning of the downsampling hidden layer.

[0060] It can be understood that after the conditional information is input into the video generation model, through the processing of the encoder, the latent space representation corresponding to the conditional information is obtained. The latent space representation of the video segment to be processed in the conditional information is input into the adaptation module, and each layer of the fine-tuning hidden layer of the adaptation module gradually downsamples the latent space representation; the fine-tuning vector output by each layer of the fine-tuning hidden layer is used as the input vector of the next layer of the fine-tuning hidden layer, and the fine-tuning vectors output by each layer of the fine-tuning hidden layer are saved. Correspondingly, random noise is input into the diffusion model for denoising processing, and downsampling is performed layer by layer through each layer of the downsampling hidden layer of the diffusion model; during the downsampling process, the output vector of each layer of the downsampling hidden layer is superimposed with the mapped fine-tuning vector to form the input vector of the next layer of the downsampling hidden layer. And so on, until the output vector of the last layer of the downsampling hidden layer is superimposed with the corresponding fine-tuning vector to form the final output vector of the downsampling module. The final output vector of the downsampling module is used as the input vector of the upsampling module for upsampling. Optionally, the output vector of the downsampling hidden layer and the corresponding fine-tuning vector can be directly superimposed or weighted superimposed.

[0061] It can be understood that one round of calculation by the above downsampling module and upsampling module is the noise prediction of one time step, that is, the denoising processing of one time step. The diffusion model can obtain the final target video after T rounds of iterative denoising. After the upsampling module outputs the output vector of the first round, it is used as the initial input vector of the second round and input into the downsampling module again. After repeating the above process T rounds, the upsampling module outputs the output vector of the Tth round to input the target video obtained by the decoder. During each round of iteration of the diffusion model, each layer of the downsampling hidden layer of the downsampling module needs to superimpose the fine-tuning vectors of the fine-tuning hidden layer mapped by the above adaptation module. The decoder can reconstruct the video from the latent space representation, map the generation result of the diffusion model from the latent space to the pixel space, and generate the target video.

[0062] It can be understood that the conditional information at least includes the video segment to be processed. Based on this, in the target video generated by the video generation model, it includes the visual content corresponding to the above P-frame edge detection map. That is to say, the conditional information includes the reference content expected by the user (the video segment includes the sparse edge detection map). Based on the guidance of the video generation model by this reference content, a high-quality target video that meets the user's expectations can be generated, thereby improving the applicability of the user to generate a high-quality video using the edge detection map corresponding to any video frame in the video segment, expanding the application field of the AIGC technology, and enhancing the interestingness of the application of the AIGC technology. Specifically, based on the adjacent edge detection maps arranged side by side, the Q-frame placeholder map can correspondingly generate visual content with continuous actions with the adjacent images, making the overall content of the target video complete and reasonable. On the basis of including the visual content corresponding to the edge detection map, the video generation model, based on the learned knowledge, enables the generated target video to also include other matching content, such as colors, background images, etc., making the content of the target video more colorful.

[0063] From this example, it can be seen that the video generation method of the present application, by using a diffusion model equipped with an adaptation module as a new video generation model, can combine the input conditional information and restore or reconstruct video content based on a small number of edge detection maps. Such a design can obtain a high-quality target video with the same duration and the same number of frames as the video segment, meeting the user's video production requirements.

[0064] See Figure 3 , an embodiment of the present application also provides a training method for the video generation model of the present application. The video generation model to be trained is composed of a pre-trained diffusion model and an adaptation module to be trained. Among them, the pre-trained diffusion model can implement the function of video generation, but the application scenario is relatively limited and it cannot generate a high-quality video that meets the user's expectations according to the edge detection map or a video segment. The video generation model of the present application can be trained according to the following method.

[0065] S210, obtain multiple groups of training data; among them, each group of training data includes a first training video and the corresponding second training video and text annotation. The first training video includes K-frame edge detection maps, and the second training video includes a training video of P-frame edge detection maps and Q-frame placeholder maps and the corresponding text annotation, where P≥1, Q≥0, and P + Q = K.

[0066] In the present application, the frame rate and duration of the target video generated by the video generation model can be preset, the total number of frames K of the target video is determined, and then it is determined that the total number of frames of the first training video in each group of training data is K.

[0067] In the same set of training data, each frame image of the first training video is an edge detection map. The second training video is obtained by replacing the Q-frame placeholder map based on the first training video. They have the same number of image frames, but the number of frames of their edge detection maps may be the same or different, that is, P ≤ K.

[0068] Furthermore, different second training videos can contain edge detection maps with the same number of frames or different numbers of frames, that is, the number of frames P of the edge detection maps in each second training video can be customized. That is to say, a second training video can contain multiple sparsely scattered frames of edge detection maps, and the remaining images are replaced with placeholder maps; each frame image of the second training video can also be entirely an edge detection map.

[0069] To obtain high-quality training videos to improve the training efficiency of the model, by screening sample videos with clear images and rich content, the sample videos have the same frame rate and duration as the first training video, and thus the same number of frames. And, after converting each frame of the sample images of the sample video into an edge detection map through an edge detection algorithm, the first training video is obtained, and then randomly select P frames of edge detection maps from it to form the second training video. Optionally, the P frames of edge detection maps can be arranged according to the corresponding frame order in the first training video, and the images of the remaining frame orders are replaced with placeholder maps.

[0070] Optionally, based on the differences in the edge detection maps randomly selected each time from the same first training video (such as different numbers or different frame orders), multiple different second training videos can be generated to form corresponding multiple sets of training data. It can be understood that based on a sample video having corresponding text annotations to describe the video content, when multiple sets of training data have the same first training video and different second training videos, each set of training data can use the same text annotation, so that the adaptation module can learn the association of visual content between different edge detection maps based on the same prompt words during training, and then more accurately restore the video content.

[0071] In some other embodiments, irrelevant K-frame sample images can be selected to form corresponding edge detection maps. The K-frame edge detection maps are combined to form a first training video, and then the P-frame edge detection maps and Q-frame placeholder maps therein are combined to form a second training video. It can be understood that for the same first training video, multiple frame sample images do not need to come from the same sample video. For example, they can be randomly selected from different sample videos, or existing images can be selected, and then these irrelevant sample images are respectively converted into edge detection maps and combined in the same first training video in an orderly or random manner. Correspondingly, the text annotation corresponding to the first training video is set according to the currently selected K-frame sample images. With such a design, by combining edge detection maps with unassociated content to form a training video, the adaptation module can learn the association between edge detection maps with a large content span during the training process, and then, under the guidance of the prompt words, can more accurately reconstruct the target video.

[0072] S220. In each round of training, the training data is input into the adaptation module of the video generation model, and the fine-tuning vectors output by each layer of the fine-tuning hidden layer are saved respectively.

[0073] In this step, the model architecture of the adaptation module is preset. As Figure 2 shown, a trained diffusion model is selected, which has a corresponding network structure U-Net. The network structure of the adaptation module is set to be the same as the network structure of the downsampling module in the diffusion model. That is, the number of layers of the fine-tuning hidden layer of the adaptation module is the same as the number of layers of the downsampling hidden layer, and the vectors of the fine-tuning hidden layer and the downsampling hidden layer have the same dimension. It should be noted that in other embodiments, the number of layers of the fine-tuning hidden layer of the trained adaptation module can also be less than the number of layers of the downsampling hidden layer, and this is not limited herein.

[0074] In some embodiments, before training the adaptation module, the network parameters of the adaptation module are initialized. Specifically, the network parameters of each layer of the downsampling hidden layer in the pre-trained diffusion model are used as the initialization parameters of each layer of the fine-tuning hidden layer of the mapping of the adaptation module.

[0075] After initializing the adaptation module, the training of the adaptation module officially starts. Among them, after each frame of edge detection image and placeholder map of the second training video are encoded and mapped to the latent space representation in the encoder, the latent space representation is input into the adaptation module, and each layer of the fine-tuning hidden layer outputs the corresponding fine-tuning vector and saves it.

[0076] It should be noted that this step S220 can be carried out synchronously and alternately with S230, that is, according to the number of fine-tuning hidden layers, every time S220 is executed, then S230 is executed, then S220 is executed again, and then S230 is executed, until the fine-tuning vector is output by the last fine-tuning hidden layer, that is, S220 ends. Or, first execute S220, and then execute S230 after saving the fine-tuning vectors output by all fine-tuning hidden layers.

[0077] S230, input the random noise, the first training video and the text annotation into the diffusion model of the video generation model to obtain the predicted noise image output by the diffusion model; wherein, the output vector of each downsampling hidden layer of the diffusion model is superimposed with the fine-tuning vector of the fine-tuning hidden layer mapped, as the input vector of the next downsampling hidden layer.

[0078] Optionally, the diffusion model and the adaptation module can share an encoder, or can use independent encoders respectively. In this step, each frame of edge detection image of the second training data in S210 is randomly noise-added and input into the diffusion model together with the text annotation. Referring to the operation principle of the diffusion model in step S120, the random noise is denoised for T time steps through the conditional guidance of the edge detection image to obtain the predicted noise image.

[0079] Taking the case where the number of fine-tuning hidden layers is the same as the number of downsampling hidden layers as an example, after the output vector of the first-layer downsampling hidden layer, it is added to the fine-tuning vector output by the first-layer fine-tuning hidden layer obtained in S220 to obtain the input vector of the second-layer downsampling hidden layer. After the output vector of the second-layer downsampling hidden layer, it is added to the fine-tuning vector output by the second-layer fine-tuning hidden layer to obtain the input vector of the third-layer downsampling hidden layer, and so on, until the output vector of the last downsampling hidden layer is added to the fine-tuning vector output by the last fine-tuning hidden layer, and then enters the upsampling module of the diffusion model for upsampling through forward propagation, and finally enters the decoder to output the predicted noise image corresponding to each frame of the input image. The specific principle can refer to the above introduction of the diffusion model and will not be elaborated here.

[0080] S240, fix the network parameters of each layer of the diffusion model, and backpropagate according to the loss function of the diffusion model to adjust the network parameters of each layer of the fine-tuning hidden layer of the adaptation module to obtain the trained video generation model.

[0081] In this application, during the training process, iterative parameter tuning is performed on the network parameters of the adaptation module, while the network parameters of each layer of the diffusion model remain fixed. Based on this, after calculating the noise loss function using the predicted noise image and the true pixels of the edge detection image corresponding to the first training video, the network parameters of each fine-tuning hidden layer of the adaptation module are adjusted through backpropagation. By iteratively looping in this way from S220 to S240 until the loss function converges, a trained adaptation module can be obtained, and each fine-tuning hidden layer has trained network parameters.

[0082] From this example, it can be seen that the training method of this application can form the first training video with edge detection maps that are content-related or unrelated, and obtain the second training video with sparse or continuous visual content expression. Based on text annotations to provide semantic features, an adaptation module is externally connected to the downsampling module of a pre-trained diffusion model to achieve the training of the adaptation module, thereby efficiently and at low cost obtaining a video generation model that can restore or reconstruct video content.

[0083] See Figure 4 , for the trained video generation model according to the above embodiments, the video generation method of this application will be further introduced in combination with embodiments below, which includes:

[0084] S310, generating an edge detection map from the P-frame original image in the original video segment, or generating an edge detection map from the P-frame original sample image, and generating a video segment to be processed according to the P-frame edge detection map and the Q-frame placeholder map, where P≥1 and Q≥0.

[0085] In this application, according to the frame rate and duration of the target video, the frame rate and duration of the video segment to be processed can be determined. The video segment to be processed includes K frames of images, P + Q = K, and P, Q, and K are all natural numbers. Among them, the number P of edge detection maps can be selected from 1 to K, that is, the number of edge detection maps is at least 1 frame, or each frame image of the video segment is an edge detection map. Conversely, according to the total number of frames K of the video segment, the number Q of placeholder maps is 0 to (K - 1). It can be understood that when the number of edge detection maps is greater than 1 frame, according to the total number of frames of the video segment, multiple edge detection maps can be arranged continuously or at intervals.

[0086] For some original video clips, there are situations where the resolution is low or some image content is damaged. When the expected goal is to restore the original video content, in some embodiments, the original video clip is obtained; when the total number of frames of the original video clip is equal to the number of frames of the video clip to be processed, the P-frame original images in the original video clip can be converted into corresponding edge detection maps, and the remaining original images are replaced with placeholder maps to generate the video clip to be processed. For example, the original images with higher quality in the original video clip can be selected and converted into edge detection maps, so as to improve the restoration degree of the video content. It can be understood that the edge detection maps of each frame can be arranged according to the frame numbers of the corresponding original images to help restore the visual content. Of course, in other embodiments, the edge detection maps of each frame can also be arranged in disorder and then used to reconstruct the visual content.

[0087] In some embodiments, when the total number of frames of the original video clip is less than the number of frames of the video clip to be processed, at least P-frame original images are converted into edge detection maps, and the missing image frames are replaced with placeholder maps, or the edge detection maps are copied for placeholder, or other edge detection maps with similar content are used for placeholder. In some embodiments, when the total number of frames of the original video clip is greater than the number of frames of the video clip to be processed, the original video clip can be clipped to match the total number of frames of the video clip.

[0088] When the user's expected goal is to reconstruct the content of the original video clip, the above processing methods for the original video clip are equally applicable and will not be elaborated here.

[0089] When the user's expected goal is to reconstruct the content of the original video clip, the user can also generate the video clip to be processed based on the existing P-frame original sample images. In some other embodiments, the P-frame original sample images are converted into edge detection maps; the P-frame original sample images and Q-frame placeholder maps are combined to form the video clip to be processed. That is to say, when the original video clip is missing, the user can select some original sample images to generate edge detection maps to achieve free creation of video content. For example, based on various visual contents of unrelated original sample images, the visual content of the generated target video can be made more abundant. For example, based on the content of related original sample images, the generated target video can restore the dynamic scene at the time of image shooting.

[0090] S320, input the video clip and the prompt text into a pre-trained video generation model to generate a target video, where the video generation model includes a diffusion model and an adaptation module.

[0091] In this embodiment, the conditional information further includes a prompt text. The prompt text at least includes the attribute information of the subject object in the target video. For example, it may include the type of the subject object and other attribute information of the subject object. The type of the subject object may include: human type, animal type, or other object (such as a certain object that needs to move in a movie or game scene) type. Other attribute information of the subject object may include the color information of the subject object, the structural information (such as the limb ratio of the subject object when the subject object is a human or an animal), and other information.

[0092] As a general video generation model, the diffusion model has relatively strong text understanding ability. In this embodiment, by using the prompt text as the input of the diffusion model, it can ensure that the generated target video better meets the user's expectations.

[0093] For the expected target, the prompt text is obtained based on custom design and / or according to the visual content of the original video clip or the original sample image. When the expected target is to restore the video content, the prompt text needs to be as consistent as possible with the visual content of the original video clip or the original sample image. When the expected target is to reconstruct the video content, the prompt text can be different from the visual content of the original video clip or the original sample image.

[0094] In this embodiment, the video generation model includes a diffusion model and an adaptation module. The diffusion model and the adaptation module may have the same encoder. The downsampling hidden layer of the diffusion model and the fine-tuning hidden layer of the adaptation module have the same number of layers and the same dimension.

[0095] After each frame image in the video clip is added with random noise, it is input into the video generation model together with the prompt text. According to the latent space expression corresponding to each encoded input image, each layer of the fine-tuning hidden layer in the adaptation module outputs a corresponding fine-tuning vector Δd one by one; for example, the first layer of the fine-tuning hidden layer outputs the fine-tuning vector Δd1, the second layer of the fine-tuning hidden layer outputs the fine-tuning vector Δd2, and the nth layer of the fine-tuning hidden layer outputs the fine-tuning vector Δd n 。

[0096] Correspondingly, the output vector of the first layer of the downsampling hidden layer in the diffusion model is d1. After superimposing d1 and Δd1, (d1 + Δd1) is used as the input vector of the second layer of the downsampling hidden layer. The output vector of the second layer of the downsampling hidden layer is d2. After superimposing d2 and Δd2, (d2 + Δd2) is used as the input vector of the third layer of the downsampling hidden layer. And so on, the output vector of the last layer of the downsampling hidden layer is d n And Δd nAfter superimposition, a final output vector is formed and input into the upsampling module for upsampling. Among them, the downsampling module and the upsampling module perform denoising according to the prompt information of the prompt text to obtain a latent space representation that conforms to the visual content of the target video. After T rounds of iterative denoising, it is input into the decoder to obtain the target video.

[0097] In summary, in the video generation method of the present application, users can select diversified sample source data according to the expected goal to obtain the video segment to be processed. The video generation model can restore or reconstruct the video content based on the edge detection map and the prompt text. With the assistance of the pre-trained adaptation module and the prompt text, the visual content of each frame of the target video is close to the user's expectation value, meeting the user's creative needs.

[0098] Corresponding to the foregoing application function implementation method embodiment, the present application also provides a video generation device and a corresponding embodiment.

[0099] Figure 5 It is a schematic structural diagram of the video generation device shown in the present application.

[0100] See Figure 5 , the video generation device shown in the present application includes a video acquisition module 510 and a video generation module 520; where:

[0101] The video acquisition module 510 is used to acquire the video segment to be processed. The video segment includes a P-frame edge detection map and a Q-frame placeholder map, where P≥1 and Q≥0.

[0102] The video generation module 520 is used to input the conditional information into the pre-trained video generation model to generate the target video. The conditional information at least includes the video segment to be processed, and the target video includes the visual content corresponding to the P-frame edge detection map.

[0103] Among them, the video generation model includes a diffusion model and an adaptation module. The diffusion model includes multiple layers of downsampling hidden layers, and the adaptation module includes multiple layers of fine-tuning hidden layers. The fine-tuning vector output by a single layer of fine-tuning hidden layer is used to be superimposed with the output vector of the mapped downsampling hidden layer as the input vector of the next layer of downsampling hidden layer.

[0104] See Figure 6, in some embodiments, the video acquisition module 510 includes a data acquisition module 511 and an image processing module 512. The data acquisition module 511 is used to acquire an original video segment or a P-frame original sample image. The image processing module 521 is used to generate an edge detection map from the P-frame original image in the original video segment, or generate an edge detection map from the P-frame original sample image; and generate a video segment to be processed according to the P-frame edge detection map and the Q-frame placeholder map. The video acquisition module 510 further includes a text acquisition module 513, and the text acquisition module 513 is used to acquire a prompt text corresponding to the video segment to be processed.

[0105] From this example, it can be seen that the video generation device of the present application, through the trained video generation model, can generate a target video based on the edge detection map of the sparse distribution / continuous distribution, where the input vector of the downsampling hidden layer of the diffusion model is fine-tuned by the adaptation module, and combined with the semantic features in the prompt text, so that the generated target video can achieve the effect of restoring or reconstructing the content and meet the creative needs of users.

[0106] An embodiment of the present application also discloses a training device for a video generation model, which includes:

[0107] A sample acquisition module, configured to acquire multiple groups of training data; wherein each group of training data includes a first training video, a second training video, and a corresponding text annotation. The first training video includes K-frame edge detection maps, the second training video includes P-frame edge detection maps and Q-frame placeholder maps, P≥1, Q≥0, and P + Q = K.

[0108] A fine-tuning vector module, configured to input the second training video into the adaptation module in each round of training, and respectively save the fine-tuning vectors output by each layer of the fine-tuning hidden layer.

[0109] A model training module, configured to input random noise, the first training video, and the text annotation into the diffusion model to obtain a predicted noise image output by the diffusion model; wherein the output vector of each layer of the downsampling hidden layer of the diffusion model is superimposed with the fine-tuning vector of the fine-tuning hidden layer mapped thereto as the input vector of the next layer of the downsampling hidden layer.

[0110] A parameter iteration module, configured to fix the network parameters of each layer of the diffusion model, and backpropagate according to the loss function of the diffusion model to adjust the network parameters of each layer of the fine-tuning hidden layer of the adaptation module to obtain a trained video generation model.

[0111] The training device of the video generation model of the present application adapts to a specific downstream task by inserting a small, trainable adaptation module into a pre-trained diffusion model, without the need to fine-tune the entire diffusion model. This allows the diffusion model to adapt to a new task by adjusting a small number of parameters while keeping the original pre-trained parameters unchanged, thereby improving the parameter training efficiency and training speed.

[0112] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0113] Figure 7 It is a schematic structural diagram of the electronic device shown in the present application.

[0114] See Figure 7 , the electronic device 1000 includes a memory 1010 and a processor 1020.

[0115] The processor 1020 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0116] The memory 1010 may include various types of storage units, such as system memory, read-only memory (ROP), and permanent storage devices. Among them, the ROP can store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage device can be a readable and writable storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device employs a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 1010 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAP, SRAP, SDRAP, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 1010 can include removable storage devices that are readable and / or writable, such as compact discs (CDs), read-only digital versatile discs (such as DVD-ROP, dual-layer DVD-ROP), read-only Blu-ray discs, ultra-density discs, flash memory cards (such as SD cards, PiQ SD cards, Picro-SD cards, etc.), magnetic floppy disks, etc. Computer-readable storage media do not include carrier waves and instantaneous electronic signals transmitted wirelessly or by wire.

[0117] Executable code is stored on the memory 1010, and when the executable code is processed by the processor 1020, it can cause the processor 1020 to execute some or all of the methods described above.

[0118] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps of the above method of the present application.

[0119] Alternatively, the present application can also be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium), on which executable code (or a computer program or computer instruction code) is stored. When the executable code (or the computer program or computer instruction code) is executed by a processor of an electronic device (or a server, etc.), it causes the processor to execute some or all of the steps of the above method according to the present application.

[0120] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A video generation method, characterized in that: include: Obtain a video clip to be processed, wherein the video clip includes a P frame edge detection map and a Q frame placeholder map, where P≥1 and Q≥0; Inputting condition information into a pre-trained video generation model to generate a target video, wherein the condition information at least includes the video segment to be processed, and the target video includes visual content corresponding to the P frame edge detection image; Among them, the video generation model includes a diffusion model and an adaptation module, the diffusion model includes multiple layers of downsampling hidden layers, and the adaptation module includes multiple layers of fine-tuning hidden layers. The multiple layers of the fine-tuning hidden layers are mapped one-to-one with the multiple layers of the downsampling hidden layers respectively, and the fine-tuning vector output by the single-layer fine-tuning hidden layer is used to superimpose with the output vector of the mapped downsampling hidden layer to serve as the input vector of the next downsampling hidden layer.

2. The method according to claim 1, characterized in that: The number of layers of the fine-tuning hidden layer of the adaptation module is less than or equal to the number of layers of the downsampling hidden layer of the diffusion model; The vector dimension of the fine-tuning hidden layer is the same as the vector dimension of the mapped down-sampling hidden layer.

3. The method according to claim 1, characterized in that The step of obtaining the video clip to be processed includes: Get the original video clip or P frame original sample image; Generate an edge detection map from the original image of the P frame in the original video clip, or generate an edge detection map from the original sample image of the P frame; A video segment to be processed is generated according to the edge detection map of the P frame and the placeholder map of the Q frame.

4. The method according to claim 1, characterized in that: The placeholder image is a solid color image.

5. The method according to claim 4, characterized in that: The placeholder image is a pure black image.

6. The method according to claim 1 or 3, characterized in that: The condition information also includes prompt text; Before inputting the conditional information into the pre-trained video generation model to generate the target video, it also includes: The prompt text is obtained based on the custom designed visual content for the expected target; and / or the prompt text is obtained according to the visual content of the original video clip or the original sample image.

7. The method according to claim 1, characterized in that The video generation model is trained according to the following method: Acquire multiple sets of training data; wherein each set of the training data includes a first training video and a second training video and corresponding text annotations, the first training video includes a K-frame edge detection image, the second training video includes a P-frame edge detection image and a Q-frame placeholder image training video and corresponding text annotations, P≥1, Q≥0, P+Q=K; In each round of training, the second training video is input into the adaptation module, and the fine-tuning vector output by each fine-tuning hidden layer is saved respectively; Inputting random noise, the first training video and the text annotation into the diffusion model to obtain a predicted noise image output by the diffusion model; wherein the output vector of each down-sampling hidden layer in the diffusion model is superimposed with the fine-tuning vector of the mapped fine-tuning hidden layer as the input vector of the next down-sampling hidden layer; The network parameters of each layer of the diffusion model are fixed, and the network parameters of each layer of the fine-tuning hidden layer of the adaptation module are adjusted according to the back propagation of the loss function of the diffusion model to obtain a trained video generation model.

8. The method according to claim 7, characterized in that The method further comprises: Before training the adaptation module, the network parameters of each down-sampling hidden layer are used as initialization parameters of each fine-tuning hidden layer of the mapping in the adaptation module.

9. A video generating device, characterized in that: include: A video acquisition module is used to acquire a video segment to be processed, wherein the video segment includes a P frame edge detection map and a Q frame placeholder map, where P≥1 and Q≥0; A video generation module, configured to input condition information into a pre-trained video generation model to generate a target video, wherein the condition information at least includes the video segment to be processed, and the target video includes visual content corresponding to the P frame edge detection map; Among them, the video generation model includes a diffusion model and an adaptation module, the diffusion model includes multiple layers of downsampling hidden layers, and the adaptation module includes multiple layers of fine-tuning hidden layers. The multiple layers of the fine-tuning hidden layers are mapped one-to-one with the multiple layers of the downsampling hidden layers respectively, and the fine-tuning vector output by the single-layer fine-tuning hidden layer is used to superimpose with the output vector of the mapped downsampling hidden layer to serve as the input vector of the next downsampling hidden layer.

10. An electronic device, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the video generating method according to any one of claims 1 to 8.

11. A computer-readable storage medium having executable codes stored thereon, which, when executed by a processor of an electronic device, causes the processor to execute the video generating method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Efficient video image editing method based on diffusion model

    CN117768678A

  • Video generation method and apparatus, and method and apparatus for training video generation model

    WO2024248736A1