Video generation method, device, electronic device, storage medium

By generating a first video including a target image and generating a second video based on the noise information and the first video using a video generation model, the problem of not being able to generate a specific image video in the prior art is solved, and the matching degree and user experience of video generation are improved.

CN119364134BActive Publication Date: 2025-07-01BEIJING SHENGSHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411935498.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-07-01
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

The existing Wensheng video model cannot generate videos including specific images, resulting in low matching between the generated videos and the videos needed by the user, affecting the user experience.

Method used

By generating a first video including a target image, the pre-trained video generation model uses the pre-trained video generation model to comprehensively consider the target image when generating the second video based on the preset noise information and the first video to ensure that the target image is included in the generated second video, and the target image is used as the first frame image of the first video to ensure that the first frame image of the second video is the target image.

Benefits of technology

The generated video includes the target image, which improves the matching degree between the video generated by the video generation model and the video required by the user and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119364134B_ABST
    Figure CN119364134B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a video generation method, apparatus, electronic device, and storage medium. The method includes: generating a first video based on a target image, where the first video includes the target image and multiple frames of all-black images, and the first frame image of the first video is the target image; inputting preset noise information and conditional guidance information into a pre-trained video generation model to obtain a second video, where the conditional guidance information includes the first video, the first frame of the second video is the target image, and the duration of the second video is equal to that of the first video. Thus, when the video generation model generates the second video based on the preset noise information and the first video, it can comprehensively consider the target image to ensure that the generated second video includes the target image. At the same time, the target image is used as the first frame image of the first video, thereby ensuring that the first frame image of the second video is the target image, and realizing video generation through the first frame image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of video generation, and in particular, to a video generation method, apparatus, electronic device, and storage medium. Background Art

[0002] In recent years, with the rapid development of text-to-video models, text-to-video models have shined brightly in the field of AIGC (Artificial Intelligence Generated Content). In related technologies, during the process of generating a video, it is often required that the generated video includes specific images, while text-to-video models cannot generate videos including specific images, thereby reducing the matching degree between the generated video and the video required by the user and affecting the user experience. Summary of the Invention

[0003] To solve the above technical problems, embodiments of the present disclosure provide a video generation method, apparatus, electronic device, storage medium, and program product.

[0004] In one aspect of the embodiments of the present disclosure, a video generation method is provided, including: generating a first video based on a target image, the first video including the target image and multiple frames of all-black images, and the first frame image of the first video being the target image; inputting preset noise information and conditional guidance information into a pre-trained video generation model to obtain a second video, the conditional guidance information including the first video, the first frame of the second video being the target image, and the duration of the second video being the same as the duration of the first video.

[0005] In another aspect of the embodiments of the present disclosure, a video generation method is provided, including: generating a first video based on a target image, the first video including the target image and multiple frames of all-black images, and the first frame image of the first video being the target image; inputting preset noise information and conditional guidance information into a pre-trained video generation model to obtain a second video, the conditional guidance information including the first video, the first frame of the second video being the target image, and the duration of the second video being the same as the duration of the first video.

[0006] In yet another aspect of the embodiments of the present disclosure, an electronic device is provided, including: a memory for storing a computer program; a processor for executing the computer program stored in the memory, and when the computer program is executed, implementing the above video generation method.

[0007] In still another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, implementing the above video generation method.

[0008] Another aspect of the embodiments of the present disclosure provides a computer program product, including computer program instructions, which implement the above video generation method when executed by a processor.

[0009] In the embodiments of the present disclosure, by generating a first video including a target image, when the video generation model generates a second video based on preset noise information and the first video, the target image can be comprehensively considered to ensure that the generated second video includes the target image; at the same time, in the embodiments of the present disclosure, the target image is also used as the first frame image of the first video, thereby ensuring that the first frame image of the second video is the target image, and realizing the generation of a video through the first frame image.

[0010] The technical solutions of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0011] The drawings forming a part of the specification depict the embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure.

[0012] Referring to the accompanying drawings, the present disclosure can be more clearly understood according to the following detailed description, where:

[0013] Figure 1 is a flowchart of a video generation method provided by an exemplary embodiment of the present disclosure;

[0014] Figure 2 is a flowchart of step S110 provided by an exemplary embodiment of the present disclosure;

[0015] Figure 3 is a schematic diagram of the input layer in an adapter network module and a video generation network module provided by an exemplary embodiment of the present disclosure;

[0016] Figure 4 is a flowchart of a video generation method provided by another exemplary embodiment of the present disclosure;

[0017] Figure 5 is a schematic structural diagram of an embodiment of a video generation device of the present disclosure;

[0018] Figure 6 is a schematic structural diagram of another embodiment of a video generation device of the present disclosure;

[0019] Figure 7 is a schematic structural diagram of an application embodiment of an electronic device of the present disclosure. Detailed Embodiments

[0020] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.

[0021] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices, or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.

[0022] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two, or more.

[0023] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present disclosure, in the absence of a clear limitation or a contrary indication in the context, it is generally understood as one or more.

[0024] In addition, the term "and / or" in the present disclosure is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.

[0025] It should also be understood that the present disclosure emphasizes the differences between various embodiments, and the same or similar parts thereof can be referred to each other. For the sake of brevity, they will not be elaborated one by one.

[0026] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0027] The following description of at least one exemplary embodiment is actually merely illustrative and in no way restricts the present disclosure and its application or use.

[0028] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.

[0029] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0030] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.

[0031] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, target programs, components, logics, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment, where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.

[0032] In the process of implementing the present disclosure, the inventors found that in practical applications, when generating a video, it is usually required that the generated video includes the specified image. However, the existing text-to-video models cannot generate a video including the specified image, which results in a low matching degree between the video generated by the text-to-video model and the video required by the user, affecting the user experience.

[0033] Figure 1 It is a schematic flowchart of a video generation method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, such as Figure 1 as shown, and includes the following steps:

[0034] Step S100, generate a first video based on the target image.

[0035] Among them, the first video includes the target image and multiple frames of all-black images, and the first frame image of the first video is the target image.

[0036] In one embodiment, multiple frames of all-black images and the target image can be combined according to preset timing information to obtain the first video. The preset timing information can include, for example: the target image is located at the first frame of the first video.

[0037] Step S110, input the preset noise information and conditional guidance information into a pre-trained video generation model to obtain a second video.

[0038] Among them, the conditional guidance information includes a first video, the first frame of the second video is the target image, and the duration of the second video is the same as that of the first video. The video generation model can be a deep learning model for extending the video based on video data.

[0039] In the embodiments of the present disclosure, by generating the first video including the target image, when the video generation model generates the second video based on the preset noise information and the first video, the target image can be comprehensively considered to ensure that the generated second video includes the target image; at the same time, in the embodiments of the present disclosure, the target image is also used as the first frame image of the first video, thereby ensuring that the first frame image of the second video is the target image, and realizing the generation of the video through the first frame image.

[0040] In some alternative embodiments, in the embodiments of the present disclosure, the conditional guidance information further includes video description information.

[0041] Among them, the video description information can be used to describe the video, and the video description information can include, for example, text information describing the video content.

[0042] Correspondingly, in step S110 of the embodiments of the present disclosure, it may include: processing the first video by the adapter network module in the video generation model to obtain a first vector output by each first network layer in the adapter network module; processing the multiple first vectors, the preset noise information, and the video description information by the video generation network module in the video generation model to obtain a second video.

[0043] Among them, the video generation model may include: an adapter network module and a video generation network module. The adapter network module and the video generation network module can be pre-trained deep learning models. Through the joint training of the video generation network module and the adapter network module to be trained, a pre-trained video generation model is obtained. Optionally, during the joint training based on the video generation network module and the adapter network module to be trained, first, the video generation network module and the adapter network module to be trained are obtained, where the network layer structure of the adapter network module to be trained is the same as the network structure of the input layer of the video generation network module, each network layer of the adapter network module to be trained corresponds one by one to the network layer of the input layer of the video generation network module, that is, the number of network layers of the adapter network module to be trained is the same as the number of network layers of the input layer of the video generation network module, and the network parameters of each network layer of the adapter network module to be trained are the same as those of the corresponding network layer of the input layer of the video generation network module. Secondly, through the loss function of the video generation network module, the network parameters of each network layer of the adapter network module to be trained are adjusted, and then the adapter network module in the video generation model is obtained.

[0044] In this embodiment, the network layer in the adapter network module can be referred to as the first network layer, and the adapter network module can include multiple first network layers. The input layer of the video generation network module includes multiple second network layers. The multiple first network layers of the adapter network module are respectively mapped to the multiple second network layers of the video generation network module in a one-to-one correspondence. Among them, the first vector output by each first network layer is used to be superimposed on the second vector output by the corresponding second network layer to serve as the input vector of the next second network layer.

[0045] Furthermore, in some embodiments, the structure of each first network layer is the same as that of the corresponding second network layer. Therefore, the dimension of the output vector of each first network layer is the same as that of the output vector of the corresponding second network layer. Thus, it is convenient for the output vector of the first network layer and the output vector of the corresponding second network layer to be superimposed in the same dimension, so as to achieve a comprehensive fine-tuning of the video generation network module through the adapter network module.

[0046] Exemplarily, the adapter network module can adopt an adapter model; the video generation network module is of a diffusion model structure. Among them, the diffusion model is a generative model based on deep learning algorithms, which generates high-quality content through a process of gradually denoising and adding noise, and performs well in generating images, videos, audio, and other high-dimensional data, and is powerful in generative tasks. In the embodiments of the present disclosure, the video generation network module can be obtained by training the diffusion model. This video generation network module is general in the field of video generation, but the matching effect between the generated video content and the input may not be very good. In the embodiments of the present disclosure, through the combination of the adapter network module and the video generation network module, on the basis of not changing the generality of the video generation network module, when the input includes a target image, a video of relatively high quality (that is, the matching effect between the generated video content and the input content is good) can be efficiently generated.

[0047] The preset noise information may include Gaussian noise. Exemplarily, a Gaussian noise matrix can be randomly generated in advance, and this Gaussian noise matrix can be determined as the preset noise information.

[0048] In one embodiment, the video generation network module has a diffusion model structure. The diffusion model mainly includes two processes: denoising and adding noise. In the process of adding noise, the original data is transformed into an expression form in the latent space through a variational autoencoder, which can make the model calculation more efficient. After adding noise for T (T≥1) steps, the noisy content is obtained. In the process of denoising, noise prediction is performed on the given noisy content, and conditional guidance information will be received during this process. The conditional guidance information includes the first video and video description information. Finally, it is restored from the latent space to the pixel space through a variational autoencoder to obtain a video that meets the expectations. In this embodiment, the video description information and the first video can be better referred to during the video generation process, so that the matching effect between the generated second video and the input content is better.

[0049] In one embodiment, when the preset noise information, the first video, and the video description information are input into the video generation model, the adapter network module processes the first video and records the first vectors output by each first network layer. Then, the video generation network module processes the first vectors output by each first network layer, the video description information, and the preset noise information to obtain a second video including the target image and outputs the second video.

[0050] In the embodiment of the present disclosure, the video generation model is designed to include an adapter network module and a video generation network module. By using the adapter network module to process the first video including the target image, multiple first vectors are obtained. Then, the video generation network module processes the multiple first vectors, the video description information, and the preset noise information, so that the video generation network module takes into account the first video and the video description information when generating the second video. As a result, the generated second video can not only include the target image, but also the video content conforms to the video description information. This not only improves the matching degree between the second video generated by the video generation model and the video required by the user, enhances the user experience, but also expands the application scope of the video generation model.

[0051] Figure 2 It is a schematic flowchart of step S100 provided by an exemplary embodiment of the present disclosure. In some alternative embodiments, as Figure 2 shown, step S100 may include the following steps:

[0052] Step S101, obtaining a preset video.

[0053] Among them, the preset video may be a video without content having a certain duration, and the preset video may include multiple frames of all-black images.

[0054] In one embodiment, a video including a preset number of all-black images may be created as the preset video. Exemplarily, a video including 160 frames of all-black images may be created as the preset video.

[0055] Step S102: Replace the first frame image of the preset video with the target image to obtain the first video.

[0056] In one embodiment, the first frame image in the preset video can be replaced with the target image to obtain the first video. Exemplarily, assume that the frame rate of the preset video is 160 FPS and the duration is 10 s. Then the preset video includes 160 all - black images. Video editing software such as Hand Brake, Freemake Video Converter, or Kuaijianji can be used to replace the first frame image (the first frame) in the preset video with the target image to obtain the first video. In this first video, the first frame image is the target image, and the remaining 159 frame images are all - black images.

[0057] Exemplarily, the obtained preset video is represented by 000000000000000000000, where each 0 represents an all - black video frame, and the target image is represented by 1. Replacing the first frame image of the preset video with the target image to obtain the first video is: 100000000000000000000. Inputting the preset noise information and the conditional guidance information including the conditional first video into the pre - trained video generation model, the final second video obtained is: 122222222222222222222, where 2 represents the video frame of the brand - new visual content generated by the pre - trained video generation model. The first frame of the second video is the target image 1, and the duration of the second video is the same as that of the first video. The temporal position of the target image in the first video is consistent with the temporal position of the target image in the second video. Therefore, through the embodiments of the present invention, video generation through the first frame can be achieved quickly and efficiently, meeting the current user's demand for generating a second video with the target image as the first frame of the video based on a target image, and enhancing the interest of video generation in the field of AIGC technology.

[0058] In the embodiments of the present disclosure, the preset video is designed to be composed of multiple all - black images, and at the same time, the first frame image in the preset video is replaced with the target image to obtain the first video. Thus, not only can the first video including the target image be generated quickly, but also since the first video only includes the target image and all - black images, interference caused by other content - containing images to the video generation model is avoided.

[0059] In some alternative embodiments, in the embodiments of the present disclosure, the adapter network module includes n first network layers, the input layer of the video generation network module includes n second network layers, and the n first network layers of the adapter network module correspond to the n second network layers of the video generation network module one by one.

[0060] Among them, the video generation network module includes an input layer and an output layer. At each time step, the video generation network module performs noise addition or noise removal for one time step through noise prediction of the input layer and the output layer. Then, after performing noise prediction for T (T≥1) time steps, the latent space representation of the finally generated video is obtained. In this embodiment, the network layer in the input layer of the video generation network module can be referred to as the second network layer, and any second network layer corresponds to a first network layer.

[0061] Among them, the input layer in the video generation network module includes n second network layers, where n is an integer greater than or equal to 1. The network structure of this input layer is the same as that of the adapter network module, that is, the adapter network module also includes n first network layers, and the first network layer with the same serial number corresponds to the second network layer. For example, the 1st first network layer corresponds to the 1st second network layer, and so on, the nth first network layer corresponds to the nth second network layer. The shape (dimension) of the data output by the first network layer is the same as that of the data output by the second network layer.

[0062] Exemplarily, both the first network layer and the second network layer can adopt the UNet network. The UNet can denoise random noise and achieve the generation from noise to video. In this embodiment, the input layer of the video generation network module is used to gradually perform downsampling on the input noise, extract features, and reduce the spatial resolution of the image. The output layer of the video generation network module gradually restores the spatial resolution of the video through deconvolution (or transposed convolution) and upsampling operations; through the skip connection structure of the UNet network, the output layer can combine the low-level features (such as details like edges and textures) in the input layer with the high-level features (such as semantic information of video description information), and the fusion of features helps to reconstruct the details of the video; finally, a second video with the same resolution as the input first video is output. In this application, based on the multi-layer second network layer structure in the input layer of the video generation network module, an independent adapter network module is additionally designed to fine-tune the vectors generated during the downsampling process of the input layer. That is to say, the adapter network module adapts to the new task by inserting additional parameters into the pre-trained video generation network module, rather than directly adjusting the parameters of the entire video generation network module (the parameter order of magnitude of the video generation network module is very large, such as in the order of billions or tens of billions, etc.), which can not only reduce the number of parameters required for fine-tuning, but also reduce the demand for computing resources during the training process and speed up the training speed. At the same time, when the video generation network module performs a specific task, through the trained adapter network module, the output vectors of each layer of the downsampling of the output layer are fine-tuned, so that it is possible to efficiently and accurately complete the specific task without large-scale overall training of the video generation network module. It can be seen that through the combination and joint training of the adapter network module and the video generation network module, based on the AIGC technology, high-quality videos can be quickly generated based on the target image, and the generality of the video generation network module will not be affected.

[0063] Correspondingly, in the embodiment of the present disclosure, the video generation network module can process multiple first vectors, preset noise information, and video description information in the following manner: in the denoising process of each time step of the video generation network module, based on the preset noise information and video description information, the n first vectors output by the n first network layers are respectively calculated with the n second vectors output by the corresponding n second network layers to obtain the second video.

[0064] It can be understood that one round of calculation by the input layer and the output layer of the video generation network module is the noise prediction for one time step, that is, the denoising process for one time step. After the video generation network module performs T rounds of iterative denoising, the finally generated second video can be obtained. After the output layer outputs the output vector of the first round, it is used as the input vector of the second round and input into the input layer again. After looping through the above process T times, the output layer outputs the output vector of the Tth round to input the finally generated video obtained by the decoder. The decoder can reconstruct the video from the latent space representation, map the generation result of the video generation network module from the latent space to the pixel space, and generate the final second video.

[0065] In the embodiments of the present disclosure, by setting the adapter network module to include n first network layers and setting the video network module to include n second network layers, each second network layer can process the first vector output by the corresponding first network layer, so that the video network module can comprehensively consider the first video when generating the second video, ensuring that the second video includes the target image.

[0066] In some alternative embodiments, in the embodiments of the present disclosure, using the n second network layers to process multiple first vectors, preset noise information, and video description information may include:

[0067] For the denoising process of each time step, the first second network layer processes the preset noise information based on the video description information to obtain a second vector output by the first second network layer; for the second second network layer to the nth second network layer, this second network layer processes the third vector based on the video description information to obtain a second vector output by this second network layer, and determines the second video based on the second vector output by the nth second network layer in the denoising process of the last time step.

[0068] Wherein, the third vector is obtained by adding the second vector output by the previous second network layer of this second network layer and the first vector output by the first network layer corresponding to the previous second network layer. In one embodiment, the second video is determined based on the second vector output by the nth second network layer and the first vector output by the nth first network layer.

[0069] Exemplarily, Figure 3 is a schematic diagram of the input layer in the adapter network module and the video generation network module provided by an exemplary embodiment of the present disclosure. In one embodiment, as Figure 3 shown, the first video can be encoded using a Variational Auto-Encoder (VAE), and then the encoded first video is input into the adapter network module, and the first vectors output by each first network layer in the adapter network module are recorded;

[0070] The VAE can be used to encode the video description information and the preset noise information. For the denoising process at each time step, the encoded video description information and the preset noise information are input into the first second network layer to obtain a second vector output by the first second network layer. The second vector is added to the first vector output by the first first network layer to obtain a third vector, and this third vector can be used as the input data for the second second network layer;

[0071] The third vector obtained by adding the second vector output by the (i - 1)-th second network layer to the first vector output by the (i - 1)-th first network layer and the encoded video description information are input into the i-th second network layer to obtain a second vector output by the i-th second network layer. The second vector is added to the first vector output by the i-th first network layer to obtain a third vector, and this third vector can be used as the input data for the (i + 1)-th second network layer, where 2 ≤ i ≤ n;

[0072] The result of adding the second vector output by the n-th second network layer in the denoising process of the last time step to the first vector output by the n-th first network layer is input into the VAE for decoding to obtain a second video.

[0073] In the embodiments of the present disclosure, the third vector obtained by adding the second vector output by each second network layer to the first vector output by the corresponding first network layer is used as the input data for the next second network layer. Thus, when the video generation network module generates the second video, the first video can be taken into account, ensuring that the generated second video includes the target image and the video content corresponds to the video description information.

[0074] Figure 4 It is a schematic flowchart of a video generation method provided by another exemplary embodiment of the present disclosure. In some alternative embodiments, as Figure 4 shown, the video generation model can be obtained in the following manner:

[0075] Step S200, obtain the model to be trained and the sample data.

[0076] Among them, the model to be trained includes: a video generation network module and an adapter network module to be trained.

[0077] The sample data includes multiple sample videos, as well as the corresponding label videos and label video description information for each sample video. For the multiple sample videos, each sample video includes a frame of an image with content and multiple frames of all-black images, and the image with content is the first frame image.

[0078] In some alternative embodiments, in the embodiments of the present disclosure, a sample video can be obtained in the following manner: Obtain a plurality of labeled videos; for the plurality of labeled videos, replace the remaining frames in each labeled video except the first frame image with all-black images to obtain the sample video corresponding to the labeled video.

[0079] Among them, a plurality of videos with the same duration and frame rate can be obtained as the labeled videos. Exemplarily, a complete video with a longer duration can be cut into multiple videos with the same duration as the labeled videos. For each labeled video, the remaining frames in the sample video except the first frame image can be replaced with all-black images to obtain the sample video corresponding to the labeled video. The above operation of replacing with all-black images is performed for each labeled video to obtain a plurality of sample videos.

[0080] Exemplarily, assume that the duration of each labeled video is 10s and the frame rate is 16FPS. Then each labeled video includes 160 frame images. For each labeled video, the first frame image of the labeled video is not processed, and the second frame image to the 160th frame image in the initial sample image are all replaced with all-black images to obtain the sample video corresponding to the labeled video. In this sample video, the first frame image is an image with content, and the remaining 159 frame images are all all-black images without content.

[0081] In an alternative embodiment, the labeled video description information can be obtained in the following manner: For each labeled video, input the labeled video into a Vision-Language Model (VLMs), and the video description information corresponding to the labeled video is output by the vision-language model, and this video description information is used as the labeled video description information corresponding to the sample video obtained based on the labeled video.

[0082] Step S210, for the plurality of sample videos, input the sample video, as well as the labeled video and the labeled video description information corresponding to the sample video, into the model to be trained. The sample video is processed by the adapter network module to be trained, and a fourth vector output by each first network layer in the adapter network module to be trained is obtained. The video generation network module processes the plurality of fourth vectors and the labeled video description information to obtain a predicted video.

[0083] Among them, the adapter network module to be trained includes n first network layers. Each sample video, as well as the labeled video and the labeled video description information corresponding to each sample video, can be sequentially input into the model to be trained, and the predicted video corresponding to each sample video is sequentially output by the model to be trained.

[0084] Step S220: Adjust the parameters of the adapter network module to be trained according to the predicted video and the labeled video corresponding to each sample video until the preset training end condition is met, and obtain a video generation model from the model to be trained.

[0085] Among them, in the model training stage, the parameters of the video generation network module can be frozen, and the parameters of the input layer of the video generation network module are used as the initial parameters of the adapter network module to be trained.

[0086] Specifically, the training process is to iteratively adjust the network parameters of the adapter network module, and the network parameters of each layer of the video generation network module remain fixed. Based on this, according to the loss function of the video generation network module, backpropagation is used to adjust the network parameters of each first network layer of the adapter network module. In this way, the loop iteration from S210 to S220 is performed until the loss function converges, and the trained adapter network module can be obtained, and each first network layer has trained network parameters.

[0087] The training method of the present disclosure can form sample data by using a sample video including a frame of an image with content and the description information of the labeled video, and externally connect an adapter network module during the downsampling process of the pre-trained video generation network module to implement the training of the adapter network module, so as to obtain a video generation model that can restore or reconstruct video content efficiently and at low cost.

[0088] In an embodiment, based on the difference between the labeled video and the predicted video corresponding to each sample video, a preset loss function can be used to determine the loss function value. Among them, the preset loss function can include, for example, but not limited to: cross-entropy error function or mean square error function, etc. The operation of determining the loss function value can be iteratively executed, and the parameters of the adapter network module to be trained are iteratively adjusted to continuously reduce the loss function value until the loss function value converges, determine that the preset training end condition is met, complete the training of the model to be trained, and use the trained model to be trained as the video generation model.

[0089] Parameter optimizers such as Stochastic Gradient Descent (SGD), Adagrad, Adaptive Moment Estimation (Adam), and Root Mean Square Prop (RMSprop) can be used to adjust the parameters of the adapter network module to be trained. For example, a parameter optimizer can be used to calculate the gradients of the parameters of the adapter network module to be trained, and the parameters can be adjusted along the direction of the gradients. The gradient represents the direction in which the loss function value decreases the most. The operation of iteratively determining the loss function value is performed until the loss function value no longer decreases, and the model to be trained is trained to obtain a video generation model.

[0090] In the embodiments of the present disclosure, a model to be trained including a video generation network module and an adapter network module to be trained is trained using multiple sample videos, as well as the corresponding label videos and label video description information of each sample video. During the model training process, only the parameters of the adapter network module to be trained need to be adjusted, which not only reduces the model training difficulty and improves the model training efficiency, but also enables the trained video generation model to generate a video including the image based on the image.

[0091] In some alternative embodiments, in the embodiments of the present disclosure, the video generation network module can process multiple fourth vectors and label video description information in the following manner:

[0092] For the denoising process at each time step, the first second network layer in the video generation network module processes the label video description information to obtain a fifth vector output by the first second network layer; for the second second network layer to the nth second network layer in the video generation network module, the second network layer processes the label video description information and a sixth vector to obtain a fifth vector output by the second network layer. Based on the fifth vector output by the nth second network layer in the denoising process at the last time step, the predicted video corresponding to the sample video is determined.

[0093] The sixth vector is obtained by adding the fifth vector output by the previous second network layer of the second network layer and the fourth vector output by the first network layer corresponding to the previous second network layer.

[0094] In one embodiment, the network structure of the adapter network module to be trained is the same as that of the input layer of the video generation network module, that is, the adapter network module to be trained also includes n first network layers. In the model to be trained, the first network layer with the same serial number corresponds to the second network layer, and the shape of the data output by the first network layer is the same as the shape of the data output by the second network layer. For example, the first network layer and the second network layer can adopt the UNet network.

[0095] In the forward propagation process of the video generation network module, for the denoising process of each time step, the labeled video description information is input into the first second network layer, and the fifth vector is output by the first second network layer. The fifth vector is added to the fourth vector output by the first first network layer to obtain a sixth vector, and the sixth vector is used as the input data of the second second network layer;

[0096] The sixth vector obtained by adding the fifth vector output by the (i - 1)-th second network layer and the fourth vector output by the (i - 1)-th first network layer and the labeled video description information are input into the i-th second network layer. The fifth vector is output by the i-th second network layer. The fifth vector is added to the fourth vector output by the i-th first network layer to obtain a sixth vector, and the sixth vector is used as the input data of the (i + 1)-th second network layer;

[0097] Based on the result of adding the fifth vector output by the n-th second network layer and the fourth vector output by the n-th first network layer in the denoising process of the last time step, the predicted video is determined.

[0098] In the embodiments of the present disclosure, by using the sixth vector obtained by adding the fifth vector output by each second network layer and the fourth vector output by the corresponding first network layer as the input data of the next second network layer, when the video generation network module generates the predicted video, it can take into account the data output by the adapter network module to be trained, which is convenient for the adapter network module to be trained to learn better.

[0099] Figure 5 It is a schematic structural diagram of an embodiment of the video generation device of the present disclosure. As Figure 5 shown, the video generation device of this embodiment may include: a first video generation module 310 and a second video generation module 320.

[0100] The first video generation module 310 is configured to generate a first video based on the target image. The first video includes the target image and multiple frames of all-black images, and the first frame image of the first video is the target image;

[0101] A second video generation module 320, configured to input preset noise information and conditional guidance information into a pre-trained video generation model to obtain a second video, where the conditional guidance information includes the first video, a first frame of the second video is the target image, and a duration of the second video is the same as a duration of the first video.

[0102] In some possible implementation manners of the present disclosure, the conditional guidance information in the embodiments of the present disclosure further includes video description information;

[0103] The inputting the preset noise information and the conditional guidance information into the pre-trained video generation model to obtain the second video is further configured to:

[0104] Process the first video by an adapter network module in the video generation model to obtain first vectors output by each first network layer in the adapter network module;

[0105] Process multiple first vectors, the preset noise information, and the video description information by a video generation network module in the video generation model to obtain the second video.

[0106] Figure 6 It is a schematic structural diagram of another embodiment of the video generation device of the present disclosure. In some possible implementation manners of the present disclosure, as Figure 6 shown, the first video generation module 310 includes:

[0107] A video acquisition sub-module 311, configured to acquire a preset video, where the preset video includes multiple frames of all-black images;

[0108] A video generation sub-module 312, configured to replace a first frame image of the preset video with the target image to obtain the first video.

[0109] In some possible implementation manners of the present disclosure, the adapter network module in the embodiments of the present disclosure includes n first network layers, an input layer of the video generation network module includes n second network layers, and the n first network layers of the adapter network module correspond to the n second network layers of the video generation network module one by one;

[0110] The processing the multiple first vectors, the preset noise information, and the video description information by the video generation network module in the video generation model to obtain the second video is further configured to:

[0111] In the denoising process of each time step of the video generation network module, based on the preset noise information and the video description information, the n first vectors output by the n first network layers are respectively calculated with the n second vectors output by the corresponding n second network layers to obtain the second video.

[0112] In some possible implementation manners of the present disclosure, in the denoising process of each time step of the video generation network module in the embodiments of the present disclosure, based on the preset noise information and the video description information, the n first vectors output by the n first network layers are respectively calculated with the n second vectors output by the corresponding n second network layers to obtain the second video, which is further used for:

[0113] For the denoising process of each time step, the first second network layer processes the preset noise information based on the video description information to obtain a second vector output by the first second network layer;

[0114] For the second second network layer to the nth second network layer, the second network layer processes a third vector based on the video description information to obtain a second vector output by the second network layer, where the third vector is obtained by adding the second vector output by the previous second network layer corresponding to the second network layer and the first vector output by the first network layer corresponding to the previous second network layer;

[0115] Based on the second vector output by the nth second network layer in the denoising process of the last time step, the second video is determined.

[0116] In some possible implementation manners of the present disclosure, the video generation device in the embodiments of the present disclosure further includes:

[0117] A training data acquisition module 330, configured to acquire a model to be trained and sample data, where the model to be trained includes a video generation network module and an adapter network module to be trained, the sample data includes a plurality of sample videos, as well as the corresponding labeled videos and labeled video description information of each sample video, and for the plurality of sample videos, the sample video includes a frame of image with content and multiple frames of all - black images;

[0118] A model training module 340, configured to, for the plurality of sample videos, input the sample videos, as well as the corresponding labeled videos and labeled video description information of the sample videos into the model to be trained, and the adapter network module to be trained processes the sample videos to obtain fourth vectors output by each first network layer in the adapter network module to be trained, and the video generation network module processes the multiple fourth vectors and the labeled video description information to obtain a predicted video;

[0119] A parameter adjustment module 350 is configured to adjust the parameters of the to-be-trained adapter network module according to the predicted videos and labeled videos corresponding to each sample video until a preset training end condition is met, and obtain the video generation model from the to-be-trained model.

[0120] In some possible implementation manners of the present disclosure, the processing of the plurality of fourth vectors and the labeled video description information by the video generation network module in the embodiments of the present disclosure is further configured to:

[0121] For the denoising process at each time step, the first second network layer in the video generation network module processes the labeled video description information to obtain a fifth vector output by the first second network layer;

[0122] For the second second network layer to the nth second network layer in the video generation network module, the second network layer processes the labeled video description information and a sixth vector to obtain a fifth vector output by the second network layer, where the sixth vector is obtained by adding the fifth vector output by the previous second network layer of the second network layer and the fourth vector output by the first network layer corresponding to the previous second network layer;

[0123] Based on the fifth vector output by the nth second network layer in the denoising process at the last time step, the predicted video is determined.

[0124] The video generation device in the embodiments of the present disclosure corresponds to the embodiments of the video generation method in the present disclosure above, and the relevant content can be referred to each other, and will not be elaborated here.

[0125] For the beneficial technical effects corresponding to the exemplary embodiments of the video generation device in the embodiments of the present disclosure, reference can be made to the corresponding beneficial technical effects in the corresponding exemplary method part above, and will not be elaborated here.

[0126] In addition, the embodiments of the present disclosure further provide an electronic device, including:

[0127] A memory for storing a computer program;

[0128] A processor for executing the computer program stored in the memory, and when the computer program is executed, implementing the video generation method in any of the above embodiments of the present disclosure.

[0129] Figure 7 It is a schematic structural diagram of an application embodiment of the electronic device of the present disclosure. Next, refer to Figure 7To describe an electronic device according to an embodiment of the present disclosure. The electronic device can be either the first device and / or the second device, or a stand-alone device independent of them, and the stand-alone device can communicate with the first device and the second device to receive the collected input signals from them.

[0130] As Figure 7 shown, the electronic device includes one or more processors and a memory.

[0131] The processor can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions.

[0132] The memory can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor can run the program instructions to implement the video generation methods of various embodiments of the present disclosure described above and / or other desired functions.

[0133] In one example, the electronic device can further include: an input device and an output device, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0134] In addition, the input device can further include, for example, a keyboard, a mouse, and so on.

[0135] The output device can output various information to the outside, including the determined distance information, direction information, etc. The output device can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0136] Of course, for simplicity, Figure 7 only some of the components related to the present disclosure in the electronic device are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device can further include any other appropriate components.

[0137] In addition to the above methods and devices, an embodiment of the present disclosure can also be a computer program product, which includes computer program instructions, and the computer program instructions, when run by a processor, cause the processor to execute the steps in the video generation methods according to various embodiments of the present disclosure described in the above part of this specification.

[0138] The computer program product can be written in any combination of one or more programming languages for executing the program code of the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0139] In addition, an embodiment of the present disclosure can also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the video generation method according to various embodiments of the present disclosure described in the foregoing part of this specification.

[0140] The computer-readable storage medium can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0141] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes various media such as ROM, RAM, magnetic disk, or optical disc that can store program code.

[0142] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for the purposes of illustration and easy understanding, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0143] In each embodiment described in this specification, a progressive approach is adopted. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple. For relevant parts, reference can be made to the partial description of the method embodiments.

[0144] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0145] The methods and apparatuses of the present disclosure can be implemented in many ways. For example, the methods and apparatuses of the present disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the method is only for illustration, and the steps of the method of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the methods according to the present disclosure.

[0146] It should also be noted that in the apparatuses, equipment, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.

[0147] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0148] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A video generation method, characterized in that: include: Based on the target image, a first video is generated, wherein the first video includes the target image and a plurality of frames of all-black images, and the first frame image of the first video is the target image; Inputting preset noise information and conditional guidance information into a pre-trained video generation model to obtain a second video, wherein the conditional guidance information includes the first video and video description information, the first frame of the second video is the target image, and the duration of the second video is the same as the duration of the first video, wherein the first video is processed by an adapter network module in the video generation model to obtain a first vector output by each first network layer in the adapter network module; the video generation network module in the video generation model processes multiple first vectors, the preset noise information and the video description information to obtain the second video; The adapter network module includes n first network layers, the input layer of the video generation network module includes n second network layers, and the n first network layers of the adapter network module correspond one-to-one to the n second network layers of the video generation network module respectively; the video generation network module in the video generation model processes multiple first vectors, the preset noise information and the video description information to obtain the second video, including: in the denoising processing of each time step of the video generation network module, based on the preset noise information and the video description information, the n first vectors output by the n first network layers are respectively calculated with the corresponding n second vectors output by the n second network layers to obtain the second video.

2. The method according to claim 1, characterized in that The step of generating a first video based on the target image includes: Acquire a preset video, wherein the preset video includes a plurality of frames of completely black images; The first frame image of the preset video is replaced with the target image to obtain the first video.

3. The method according to claim 1, characterized in that In the denoising process of each time step of the video generation network module, based on the preset noise information and the video description information, the n first vectors output by the n first network layers are respectively calculated with the corresponding n second vectors output by the n second network layers to obtain the second video, including: For the denoising process at each time step, the first second network layer processes the preset noise information based on the video description information to obtain a second vector output by the first second network layer; For the second second network layer to the nth second network layer, the second network layer processes the third vector based on the video description information to obtain a second vector output by the second network layer, wherein the third vector is obtained by adding a second vector output by a second network layer immediately preceding the second network layer to a first vector output by a first network layer corresponding to the immediately preceding second network layer; The second video is determined based on the second vector output by the nth second network layer in the denoising process of the last time step.

4. The method according to claim 1, characterized in that: The video generation model is obtained in the following way: Acquire a model to be trained and sample data, wherein the model to be trained includes a video generation network module and an adapter network module to be trained, and the sample data includes multiple sample videos, and label videos and label video description information corresponding to each sample video, and for the multiple sample videos, the sample videos include one frame of image with content and multiple frames of completely black images; For the multiple sample videos, the sample videos, the label videos corresponding to the sample videos, and the label video description information are input into the model to be trained, the adapter network module to be trained processes the sample videos to obtain fourth vectors output by each first network layer in the adapter network module to be trained, and the video generation network module processes the multiple fourth vectors and the label video description information to obtain a predicted video; The parameters of the adapter network module to be trained are adjusted according to the predicted video and the label video corresponding to each sample video until the preset training end condition is met, and the video generation model is obtained from the model to be trained.

5. The method according to claim 4, characterized in that The processing of the plurality of fourth vectors and the label video description information by the video generation network module includes: For the denoising process at each time step, the first second network layer in the video generation network module processes the label video description information to obtain a fifth vector output by the first second network layer; For the second second network layer to the nth second network layer in the video generation network module, the second network layer processes the label video description information and the sixth vector to obtain a fifth vector output by the second network layer, and the sixth vector is obtained by adding the fifth vector output by the second network layer above the second network layer and the fourth vector output by the first network layer corresponding to the above second network layer; The predicted video is determined based on the fifth vector output by the nth second network layer in the denoising process of the last time step.

6. A video generating device, characterized in that: include: A first video generating module, configured to generate a first video based on a target image, wherein the first video includes the target image and a plurality of frames of all-black images, and the first frame image of the first video is the target image; A second video generation module is used to input preset noise information and conditional guidance information into a pre-trained video generation model to obtain a second video, wherein the conditional guidance information includes the first video and video description information, the first frame of the second video is the target image, and the duration of the second video is the same as the duration of the first video, wherein the first video is processed by an adapter network module in the video generation model to obtain a first vector output by each first network layer in the adapter network module; and the video generation network module in the video generation model processes multiple first vectors, the preset noise information and the video description information to obtain the second video; Among them, the adapter network module includes n first network layers, the input layer of the video generation network module includes n second network layers, and the n first network layers of the adapter network module correspond one-to-one to the n second network layers of the video generation network module respectively; the video generation network module in the video generation model processes multiple first vectors, the preset noise information and the video description information to obtain the second video, and is further used for: in the denoising processing of each time step of the video generation network module, based on the preset noise information and the video description information, the n first vectors output by the n first network layers are respectively calculated with the corresponding n second vectors output by the n second network layers to obtain the second video.

7. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to execute the computer program stored in the memory, and when the computer program is executed, the video generation method described in any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the video generation method described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Event detection model training method and event detection method

    CN108334910A

  • Video generation method and device, electronic equipment and readable storage medium

    CN118042246A