Method, apparatus, device, and medium for generating a video using an image

By designing a video generation model that includes an adapter network module and a video generation network module, the problem that the prior art cannot generate related videos based on images is solved, and high-quality videos containing input images are generated, which improves the user experience.

CN119364135BActive Publication Date: 2025-07-01BEIJING SHENGSHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411935829.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-07-01
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

The existing Wensheng video model cannot generate videos related to the image based on the image, cannot meet the needs of users, and affects the user experience.

Method used

A video generation model is designed, including an adapter network module and a video generation network module. The adapter network module processes the input image and obtains multiple vectors. The video generation network module processes these vectors and preset noise information to generate a target video so that it contains the input image.

Benefits of technology

By combining the adapter network module and the video generation network module, video related video can be effectively generated, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119364135B_ABST
    Figure CN119364135B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method, apparatus, device, and medium for generating a video using an image. The method includes: obtaining conditional guidance information, where the conditional guidance information includes an input image; inputting preset noise information and the conditional guidance information into a pre-trained video generation model to obtain a target video, where the target video includes the input image. Among them, the input image is processed by an adapter network module in the video generation model to obtain a plurality of first vectors, and the plurality of first vectors and the preset noise information are processed by a video generation network module in the video generation model to obtain the target video. This enables the video generation network module to consider the input image when generating the target video, making the generated target video include the input image, thereby improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of video generation, and in particular, to a method, apparatus, device, and medium for generating a video using an image. Background Art

[0002] In recent years, with the rapid development of text-to-video models, text-to-video models have shone brightly in the field of AIGC (Artificial Intelligence Generated Content). In related technologies, a text-to-video model usually generates a corresponding video according to a text description. During the video generation process, sometimes it is necessary to generate a video based on a specified image. However, existing text-to-video models usually generate a corresponding video according to a text description and cannot generate a corresponding video based on an image, thus unable to meet user needs and affecting the user experience. Summary of the Invention

[0003] To solve the above technical problems, embodiments of the present disclosure provide a method, apparatus, device, and medium for generating a video using an image.

[0004] In one aspect of the embodiments of the present disclosure, a method for generating a video using an image is provided, including: obtaining conditional guidance information, where the conditional guidance information includes an input image; inputting preset noise information and the conditional guidance information into a pre-trained video generation model to obtain a target video, where the target video includes the input image; wherein, the input image is processed by an adapter network module in the video generation model to obtain first vectors output by each first network layer in the adapter network module, and the video generation network module in the video generation model processes a plurality of first vectors and the preset noise information to obtain the target video.

[0005] In another aspect of the embodiments of the present disclosure, a device for generating a video using an image is provided, including: an information acquisition module for obtaining conditional guidance information, where the conditional guidance information includes an input image; a video generation module for inputting preset noise information and the conditional guidance information into a pre-trained video generation model to obtain a target video, where the target video includes the input image; wherein, the input image is processed by an adapter network module in the video generation model to obtain first vectors output by each first network layer in the adapter network module, and the video generation network module in the video generation model processes a plurality of first vectors and the preset noise information to obtain the target video.

[0006] Another aspect of the embodiments of the present disclosure provides an electronic device, including: a memory for storing a computer program; a processor for executing the computer program stored in the memory, and when the computer program is executed, implementing the method for generating a video using an image as described above.

[0007] Another aspect of the embodiments of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, implementing the method for generating a video using an image as described above.

[0008] Another aspect of the embodiments of the present disclosure provides a computer program product, including computer program instructions, and when the computer program instructions are executed by a processor, implementing the method for generating a video using an image as described above.

[0009] In the embodiments of the present disclosure, the video generation model is designed to include an adapter network module and a video generation network module. The input image is processed by the adapter network module to obtain a plurality of first vectors, and then the plurality of first vectors and preset noise information are processed by the video generation network module, so that the video generation network module can consider the input image when generating the target video, and the generated target video includes the input image, thereby improving the user experience.

[0010] The technical solutions of the present disclosure will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings

[0011] The drawings constituting a part of the specification depict the embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure.

[0012] Referring to the drawings, the present disclosure can be more clearly understood according to the following detailed description, where:

[0013] Figure 1 is a schematic flowchart of a method for generating a video using an image provided by an exemplary embodiment of the present disclosure;

[0014] Figure 2 is a schematic flowchart of step S100 provided by an exemplary embodiment of the present disclosure;

[0015] Figure 3 is a schematic diagram of an input layer in an adapter network module and a video generation network module provided by an exemplary embodiment of the present disclosure;

[0016] Figure 4 is a schematic flowchart of generating a video using an image provided by another exemplary embodiment of the present disclosure;

[0017] Figure 5Schematic diagram of the structure of an embodiment of the apparatus for generating a video using an image according to the present disclosure;

[0018] Figure 6 Schematic diagram of the structure of another embodiment of the apparatus for generating a video using an image according to the present disclosure;

[0019] Figure 7 Schematic diagram of the structure of an application embodiment of an electronic device according to the present disclosure. Detailed implementation manners

[0020] Now, various exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.

[0021] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.

[0022] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.

[0023] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, without clear limitation or contrary indication in the context, it is generally understood to be one or more.

[0024] In addition, the term "and / or" in the present disclosure is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.

[0025] It should also be understood that the present disclosure emphasizes the differences between various embodiments, and their similarities or similarities can be referred to each other. For the sake of brevity, they will not be elaborated one by one.

[0026] At the same time, it should be understood that for the sake of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0027] The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present disclosure and its application or use.

[0028] Techniques, methods, and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and equipment should be considered as part of the specification.

[0029] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.

[0030] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.

[0031] Terminal devices, computer systems, servers, and other electronic devices can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.

[0032] In the process of implementing the present disclosure, the inventors found that in practical applications, during the video generation process, sometimes it is necessary to generate a video related to a specific image based on the specific image. For example, when creating an animated video, an animated video related to an image can be created based on the image. However, the existing text-to-video models generate videos based on text descriptions and cannot generate videos related to an image based on the image, which makes the text-to-image models unable to meet the needs of users and affects the user experience.

[0033] Figure 1 is a schematic flowchart of a method for generating a video using an image provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, such as Figure 1 shown, and includes the following steps:

[0034] Step S100, obtain conditional guidance information.

[0035] Among them, the conditional guidance information includes an input image, which can be used to generate a video. The input image can be, for example, a grayscale image or a true-color image, etc.

[0036] In step S110, input the preset noise information and the conditional guidance information into a pre-trained video generation model to obtain a target video.

[0037] Among them, the target video includes the input image. The conditional guidance information includes a plurality of individual input images. According to the preset order of the plurality of input images, input them into the pre-trained video generation model to obtain a target video including the plurality of input images, and the temporal order of the plurality of input images in the target video is the same as the preset order. Exemplarily, the conditional guidance information includes input image A, input image B, and input image C. The preset order of these 3 input images is: input image A, input image C, input image B. Then input the conditional guidance information into the pre-trained video generation model to obtain a target video, and the order in which the 3 input images appear in the target video is also input image A, input image C, input image B. Through this embodiment, the content of the finally generated target video can be controlled based on the business requirements or the preset order of the input images set by the user, which not only increases the fun of using video generation technology, but also meets more needs of users.

[0038] In step S110, the input image is processed by the adapter network module in the video generation model to obtain a first vector output by each first network layer in the adapter network module, and the video generation network module in the video generation model processes the plurality of first vectors and the preset noise information to obtain a target video.

[0039] The video generation model may include: an adapter network module and a video generation network module. The adapter network module and the video generation network module may be pre-trained deep learning models. In this embodiment, a pre-trained video generation model is obtained through the joint training of the video generation network module and the adapter network module to be trained. Optionally, during the process of joint training based on the video generation network module and the adapter network module to be trained, first, the video generation network module and the adapter network module to be trained are obtained, where the network layer structure of the adapter network module to be trained is the same as the network structure of the input layer of the video generation network module, and each network layer of the adapter network module to be trained corresponds one-to-one with the network layer of the input layer of the video generation network module, that is, the number of network layers of the adapter network module to be trained is the same as the number of network layers of the input layer of the video generation network module, and the network parameters of each network layer of the adapter network module to be trained are the same as the network parameters of the corresponding network layer of the input layer of the video generation network module. Secondly, through the loss function of the video generation network module, the network parameters of each network layer of the adapter network module to be trained are adjusted, and then the adapter network module in the video generation model is obtained.

[0040] In this embodiment, the network layer in the adapter network module may be referred to as the first network layer, and the adapter network module may include multiple first network layers. The input layer of the video generation network module includes multiple second network layers, and the multiple first network layers of the adapter network module are respectively mapped one-to-one with the multiple second network layers of the video generation network module. Among them, the output first vector of each first network layer is used to be superimposed with the output second vector of the corresponding second network layer to be used as the input vector of the next second network layer.

[0041] Furthermore, in some embodiments, the structure of each first network layer is the same as that of the corresponding second network layer. Therefore, the dimension of the output vector of each first network layer is the same as the dimension of the output vector of the corresponding second network layer. Thus, it is convenient for the output vector of the first network layer and the output vector of the corresponding second network layer to be superimposed in the same dimension, so as to realize the overall fine-tuning of the video generation network module through the adapter network module.

[0042] Exemplarily, the adapter network module may adopt an adapter model; the video generation network module may be a Diffusion model structure. Among them, the Diffusion model is a generative model based on deep learning algorithms. It generates high-quality content through a process of gradually denoising and adding noise, performs well in generating images, videos, audio, and other high-dimensional data, and has powerful performance in generative tasks. In the embodiments of the present disclosure, the video generation network module can be obtained by training the Diffusion model. The video generation network module is general in the field of video generation, but the matching effect between the generated video content and the input may not be very good. In the embodiments of the present disclosure, through the combination of the adapter network module and the video generation network module, on the basis of not changing the generality of the video generation network module, when the input includes an image, a relatively high-quality video can be efficiently generated, that is, the matching effect between the generated video content and the input content is good.

[0043] The preset noise information may include Gaussian noise. Exemplarily, a Gaussian noise matrix can be randomly generated in advance, and the Gaussian noise matrix can be determined as the preset noise information.

[0044] In one implementation, the video generation network module is a Diffusion model structure. The Diffusion model mainly includes two processes: denoising and adding noise. In the process of adding noise, the original data becomes the expression form in the latent space through a variational autoencoder, which can make the model calculation more efficient. After T (T≥1) steps of adding noise, the noisy content is obtained. In the process of denoising, noise prediction is performed on the given noisy content, and conditional guidance information will be received during the process. The conditional guidance information includes the input image. Finally, it is restored from the latent space to the pixel space through a variational autoencoder to obtain a video that meets the expectations. In this embodiment, the input image can be better referred to during the video generation process, thereby making the matching effect between the generated video and the input content better.

[0045] In one implementation, when the preset noise information and the input image are input into the video generation model, the adapter network module processes the input image and records the first vectors output by each first network layer; then the video generation network module processes the first vectors output by each first network layer and the preset noise information, and outputs a target video including the input image.

[0046] In the embodiments of the present disclosure, the video generation model is designed to include an adapter network module and a video generation network module. The input image is processed by the adapter network module to obtain multiple first vectors. Then, the video generation network module processes the multiple first vectors and the preset noise information, so that the video generation network module can consider the input image when generating the target video, and the generated target video includes the input image, thereby improving the user experience.

[0047] In some alternative embodiments, in the embodiments of the present disclosure, the conditional guidance information further includes video description information.

[0048] Among them, the video description information can be used to describe the video. For example, the video description information may include text information describing the video content.

[0049] Correspondingly, step S110 in the embodiments of the present disclosure may include: processing the input image by an adapter network module in the video generation model to obtain first vectors output by each first network layer in the adapter network module, and processing the multiple first vectors, preset noise information, and video description information by a video generation network module in the video generation model to obtain a target video.

[0050] In the embodiments of the present disclosure, by using the video description information and the input image together as the conditional guidance information, the video generation network module takes into account both the input image and the video description information when generating the target video, not only making the generated target video include the input image, but also making the video content conform to the video description information, thereby improving the user experience.

[0051] Figure 2 It is a schematic flowchart of step S100 provided by an exemplary embodiment of the present disclosure. In some alternative embodiments, as Figure 2 described above, step S100 may include the following steps:

[0052] Step S101, obtaining an image sequence.

[0053] Among them, the image sequence includes multiple frames of images.

[0054] Step S102, extracting an input image from a preset position in the image sequence based on preset position information.

[0055] Among them, the preset position information may include the extraction position of the image sequence. For example, the preset position information may include the p-th frame image, etc., where p is an integer greater than or equal to 1, that is, it means extracting the p-th frame image in the image sequence as the input image.

[0056] In one embodiment, the input image may also be processed such as denoising, sharpening, and smoothing to improve the image quality of the input image.

[0057] Step S103, determining conditional guidance information based on the input image.

[0058] Among them, the input image may be used as the conditional guidance information.

[0059] In an embodiment of the present disclosure, an input image is extracted from multiple frames of images in an image sequence according to preset position information, thereby enabling the rapid acquisition of the input image.

[0060] In some alternative embodiments, in an embodiment of the present disclosure, the adapter network module includes n first network layers, the input layer of the video generation network module includes n second network layers, and the n first network layers of the adapter network module correspond to the n second network layers of the video generation network module one by one.

[0061] In this embodiment, the network layers in the video generation network module can be referred to as second network layers, and any second network layer corresponds to a first network layer.

[0062] Among them, the video generation network module includes an input layer and an output layer. At each time step, the video generation network module performs noise addition or noise removal for one time step through noise prediction of the input layer and the output layer. Then, after performing noise prediction for T (T≥1) time steps, the latent space representation of the finally generated video is obtained. In this embodiment, the network layers in the input layer of the video generation network module can be referred to as second network layers, and any second network layer corresponds to a first network layer.

[0063] Among them, the input layer of the video generation network module includes n second network layers, where n is an integer greater than or equal to 1. The network structure of this input layer is the same as that of the adapter network module, that is, the adapter network module also includes n first network layers, and the first network layer with the same serial number corresponds to the second network layer. For example, the first first network layer corresponds to the first second network layer, and so on, the nth first network layer corresponds to the nth second network layer. The shape (dimension) of the data output by the first network layer is the same as that of the data output by the second network layer.

[0064] Exemplarily, both the first network layer and the second network layer can adopt the UNet network. UNet can denoise random noise to achieve the generation from noise to video. In this embodiment, the input layer of the video generation network module is used to gradually perform downsampling on the input noise, extract features, and reduce the spatial resolution of the image. The output layer of the video generation network module gradually restores the spatial resolution of the video through deconvolution (or transposed convolution) and upsampling operations; through the skip connection structure of the UNet network, the output layer can combine the low-level features (such as details like edges and textures) in the input layer with the high-level features (such as semantic information of video description information), and the fusion of features helps to reconstruct the details of the video; finally, the target video with the same resolution as the input video is output. In the present disclosure, based on the multi-layer second network layer structure in the input layer of the video generation network module, an independent adapter network module is additionally designed to fine-tune the vectors generated during the downsampling process of the input layer. That is to say, the adapter network module adapts to the new task by inserting additional parameters into the pre-trained video generation network module, rather than directly adjusting the parameters of the entire video generation network module (the parameter order of magnitude of the video generation network module is very large, such as in the order of billions or tens of billions, etc.). This can not only reduce the number of parameters required for fine-tuning, but also reduce the demand for computing resources during the training process and speed up the training speed. At the same time, when the video generation network module performs a specific task, through the trained adapter network module, the output vectors of each layer of the downsampling of the output layer are fine-tuned, so that it is not necessary to perform large-scale overall training on the video generation network module, and the specific task can be completed efficiently and accurately. It can be seen that through the combination and joint training of the adapter network module and the video generation network module, based on the AIGC technology, high-quality videos can be quickly generated based on the target image, and the generality of the video generation network module will not be affected.

[0065] Correspondingly, in the embodiment of the present disclosure, the video generation network module can process multiple first vectors, preset noise information, and video description information in the following manner: in the denoising process of each time step of the video generation network module, based on the preset noise information and video description information, the n first vectors output by the n first network layers are respectively calculated with the n second vectors output by the corresponding n second network layers to obtain the target video.

[0066] In the embodiment of the present disclosure, by setting the adapter network module to include n first network layers and setting the video network module to include n second network layers, each second network layer can process the first vector output by the corresponding first network layer, so that the video network module can comprehensively consider the input image when generating the target video, and the generated target video includes the input image.

[0067] In some alternative embodiments, in the embodiments of the present disclosure, processing the plurality of first vectors, the preset noise information, and the video description information by using n second network layers may include:

[0068] For the denoising process at each time step, the first second network layer processes the preset noise information based on the video description information to obtain a second vector output by the first second network layer; for the second second network layer to the nth second network layer, the second network layer processes a third vector based on the video description information to obtain a second vector output by the second network layer, and determines the target video based on the second vector output by the nth second network layer in the denoising process of the last time step.

[0069] Wherein, the third vector is obtained by adding the second vector output by the previous second network layer of the second network layer and the first vector output by the first network layer corresponding to the previous second network layer. In one embodiment, the target video is determined based on the second vector output by the nth second network layer and the first vector output by the nth first network layer.

[0070] Exemplarily, Figure 3 is a schematic diagram of an input layer in an adapter network module and a video generation network module provided by an exemplary embodiment of the present disclosure. In one embodiment, as Figure 3 shown, a variational auto-encoder (VAE) can be used to encode the input image, and then the encoded input image is input into the adapter network module to record the first vectors output by each first network layer in the adapter network module;

[0071] The VAE can be used to encode the video description information and the preset noise information. For the denoising process at each time step, the encoded video description information and the preset noise information are input into the first second network layer to obtain a second vector output by the first second network layer, and the second vector is added to the first vector output by the first first network layer to obtain a third vector, which can be used as the input data of the second second network layer;

[0072] The third vector obtained by adding the second vector output by the (i - 1)th second network layer and the first vector output by the (i - 1)th first network layer and the encoded video description information are input into the ith second network layer to obtain a second vector output by the ith second network layer, and the second vector is added to the first vector output by the ith first network layer to obtain a third vector, which can be used as the input data of the (i + 1)th second network layer, where 2 ≤ i ≤ n;

[0073] The result of adding the second vector output by the nth second network layer in the denoising process of the last time step to the first vector output by the nth first network layer is input into the VAE for decoding processing to obtain the target video.

[0074] In the embodiments of the present disclosure, the third vector obtained by adding the second vector output by each second network layer to the first vector output by the corresponding first network layer is used as the input data for the next second network layer. Thus, when the video generation network module generates the target video, it can take into account the input image, ensuring that the generated target video can include the input image and the video content corresponds to the video description information.

[0075] Figure 4 It is a schematic flowchart of generating a video using an image provided by another exemplary embodiment of the present disclosure. In some alternative embodiments, as Figure 4 shown, the video generation model can be obtained in the following manner:

[0076] Step S200, obtain the model to be trained and the sample data.

[0077] Among them, the model to be trained includes: a video generation network module and an adapter network module to be trained. The sample data includes: a plurality of sample images, and the corresponding labeled videos and labeled video description information for each sample image.

[0078] In some alternative embodiments, the sample images can be obtained in the following manner: obtain a plurality of labeled videos, and for the plurality of labeled videos, extract the preset frame images in the labeled video as the sample images corresponding to the labeled video.

[0079] Among them, the preset frame image can be, for example, the first frame image of the labeled video. Exemplarily, a complete video with a long duration can be cut into multiple videos with the same duration as the labeled videos. For each labeled video, extract the first frame image (the first frame image) in the labeled video as the sample image.

[0080] In an alternative embodiment, the labeled video description information can be obtained in the following manner: for each labeled video, input the labeled video into a vision-language model (Vision-Language Models, VLMs), and the vision-language model outputs the video description information corresponding to the labeled video, and use this video description information as the labeled video description information corresponding to the sample image obtained based on the labeled video.

[0081] Step S210: For multiple sample images, input the sample image, the corresponding labeled video and the labeled video description information of the sample image into the model to be trained. The adapter network module to be trained processes the sample image to obtain a fourth vector output by each first network layer in the adapter network module to be trained. The video generation network module processes the multiple fourth vectors and the labeled video description information to obtain a predicted video.

[0082] Among them, the adapter network module to be trained includes multiple first network layers. Each sample image, the corresponding labeled video and the labeled video description information of each sample image can be sequentially input into the model to be trained, and the model to be trained sequentially outputs the predicted videos corresponding to each sample image.

[0083] Step S220: Adjust the parameters of the adapter network module to be trained according to the predicted videos and the labeled videos corresponding to each sample image until the preset training end condition is met, and obtain a video generation model from the model to be trained.

[0084] Among them, in the model training stage, the parameters of the video generation network module can be frozen, and the parameters of the input layer of the video generation network module are used as the initial parameters of the adapter network module to be trained.

[0085] Specifically, the training process is to iteratively adjust the network parameters of the adapter network module, and the network parameters of each layer of the video generation network module remain fixed. Based on this, according to the loss function of the video generation network module, the network parameters of each first network layer of the adapter network module are adjusted by backpropagation. Iteratively execute the operations in S210 to S220 until the loss function converges, and a trained adapter network module can be obtained, and each first network layer has trained network parameters.

[0086] The training method of the present disclosure can form sample data from sample images and labeled text information. Due to the content sparsity characteristics of the sample images, an adapter network module is externally connected during the downsampling process of the pre-trained video generation network module to train the adapter network module, so as to efficiently and low-costly obtain a video generation model that can restore or reconstruct video content.

[0087] In one embodiment, based on the difference between the labeled video and the predicted video corresponding to each sample image, a preset loss function can be used to determine the loss function value. Among them, the preset loss function can include, for example, but is not limited to: cross-entropy error function or mean square error function, etc. The operation of determining the loss function value can be iteratively executed, and the parameters of the adapter network module to be trained can be iteratively adjusted to continuously reduce the loss function value until the loss function value converges, determine that the preset training end condition is met, complete the training of the model to be trained, and use the trained model to be trained as the video generation model.

[0088] Parameter optimizers such as Stochastic Gradient Descent (SGD), Adagrad, Adaptive Moment Estimation (Adam), and Root Mean Square Prop (RMSprop) can be used to adjust the parameters of the adapter network module to be trained. For example, a parameter optimizer can be used to calculate the gradients of the parameters of the adapter network module to be trained, and the parameters can be adjusted along the direction of the gradients. The gradient represents the direction in which the loss function value decreases the most. The operation of iteratively determining the loss function value is performed until the loss function value no longer decreases, and the training of the model to be trained is completed to obtain a video generation model.

[0089] In the embodiments of the present disclosure, a model to be trained including a video generation network module and an adapter network module to be trained is trained by multiple sample images, as well as the label videos and label video description information corresponding to the respective sample images. During the model training process, only the parameters of the adapter network module to be trained need to be adjusted, which not only reduces the model training difficulty and improves the model training efficiency, but also enables the trained video generation model to generate a video including the image based on the image.

[0090] In some alternative embodiments, in the embodiments of the present disclosure, the video generation network module can process multiple fourth vectors and label video description information in the following manner: the first second network layer in the video generation network module processes the label video description information to obtain a fifth vector output by the first second network layer; for the second second network layer to the nth second network layer in the video generation network module, the second network layer processes the label video description information and a sixth vector to obtain a fifth vector output by the second network layer, and a predicted video is determined based on the fifth vector output by the nth second network layer.

[0091] The sixth vector is obtained by adding the fifth vector output by the previous second network layer of the second network layer and the fourth vector output by the first network layer corresponding to the previous second network layer.

[0092] In one embodiment, the network structure of the adapter network module to be trained is the same as the network structure of the input layer of the video generation network module, that is, the adapter network module to be trained also includes n first network layers. In the model to be trained, the first network layer with the same serial number corresponds to the second network layer, and the shape of the data output by the first network layer is the same as the shape of the data output by the second network layer. For example, the first network layer and the second network layer can adopt a UNet network.

[0093] During the forward propagation of the video generation network module, the labeled video description information is input into the first second network layer, and the fifth vector is output by the first second network layer. The fifth vector is added to the fourth vector output by the first first network layer to obtain a sixth vector, which serves as the input data for the second second network layer;

[0094] The sixth vector obtained by adding the fifth vector output by the (i - 1)-th second network layer to the fourth vector output by the (i - 1)-th first network layer and the labeled video description information are input into the i-th second network layer. The fifth vector is output by the i-th second network layer, and the fifth vector is added to the fourth vector output by the i-th first network layer to obtain a sixth vector, which serves as the input data for the (i + 1)-th second network layer;

[0095] Based on the result of adding the fifth vector output by the n-th second network layer to the fourth vector output by the n-th first network layer, the predicted video is determined.

[0096] In the embodiments of the present disclosure, the sixth vector obtained by adding the fifth vector output by each second network layer to the fourth vector output by the corresponding first network layer is used as the input data for the next second network layer, enabling the video generation network module to consider the data output by the adapter network module to be trained when generating the predicted video, facilitating better learning of the adapter network module to be trained.

[0097] Figure 5 This is a schematic structural diagram of an embodiment of the apparatus for generating a video from an image according to the present disclosure. As Figure 5 shown, the apparatus for generating a video from an image in this embodiment may include: an information acquisition module 300 and a video generation module 310.

[0098] The information acquisition module 300 is configured to acquire conditional guidance information, where the conditional guidance information includes an input image;

[0099] The video generation module 310 is configured to input preset noise information and the conditional guidance information into a pre-trained video generation model to obtain a target video, where the target video includes the input image; wherein, the input image is processed by an adapter network module in the video generation model to obtain first vectors output by each first network layer in the adapter network module, and the video generation network module in the video generation model processes the multiple first vectors and the preset noise information to obtain the target video.

[0100] In some possible implementation manners of the present disclosure, the conditional guidance information in the embodiments of the present disclosure further includes video description information;

[0101] Inputting the preset noise information and the conditional guidance information into a pre-trained video generation model to obtain a target video, which is further used for:

[0102] Processing the input image by an adapter network module in the video generation model to obtain first vectors output by each first network layer in the adapter network module, and processing the multiple first vectors, the preset noise information, and the video description information by a video generation network module in the video generation model to obtain the target video.

[0103] Figure 6 This is a schematic structural diagram of another embodiment of the apparatus for generating a video using an image in the present disclosure. In some possible implementation manners of the present disclosure, as Figure 6 shown, the information acquisition module 300 includes:

[0104] An image sequence acquisition sub-module 301, configured to acquire an image sequence;

[0105] An image extraction sub-module 302, configured to extract the input image from the image sequence based on preset position information;

[0106] A generation sub-module 303, configured to determine the conditional guidance information based on the input image.

[0107] In some possible implementation manners of the present disclosure, the adapter network module in the embodiments of the present disclosure includes n first network layers, the input layer of the video generation network module includes n second network layers, and the n first network layers of the adapter network module are respectively in one-to-one correspondence with the n second network layers of the video generation network module;

[0108] The processing of the multiple first vectors, the preset noise information, and the video description information by the video generation network module in the video generation model is further used for:

[0109] In the denoising process of each time step of the video generation network module, based on the preset noise information and the video description information, calculating the n first vectors output by the n first network layers respectively with the n second vectors output by the corresponding n second network layers to obtain the target video. In some possible implementation manners of the present disclosure, in the denoising process of each time step of the video generation network module in the embodiments of the present disclosure, based on the preset noise information and the video description information, calculating the n first vectors output by the n first network layers respectively with the n second vectors output by the corresponding n second network layers to obtain the target video, which is further used for:

[0110] For the denoising process at each time step, the first second network layer processes the preset noise information based on the video description information to obtain a second vector output by the first second network layer;

[0111] For the second second network layer to the nth second network layer, the second network layer processes a third vector based on the video description information to obtain a second vector output by the second network layer. The third vector is obtained by adding the second vector output by the previous second network layer of the second network layer and the first vector output by the first network layer corresponding to the previous second network layer;

[0112] Based on the second vector output by the nth second network layer in the denoising process of the last time step, the target video is determined.

[0113] In some possible implementation manners of the present disclosure, the device for generating a video using an image in the embodiments of the present disclosure further includes:

[0114] A training data acquisition module 320, configured to acquire a model to be trained and sample data. The model to be trained includes a video generation network module and an adapter network module to be trained. The sample data includes a plurality of sample images, as well as a labeled video and labeled video description information corresponding to each sample image;

[0115] A model training module 330, configured to input, for the plurality of sample images, the sample images, the labeled video and the labeled video description information corresponding to the sample images into the model to be trained. The adapter network module to be trained processes the sample images to obtain fourth vectors output by each first network layer in the adapter network module to be trained. The video generation network module processes the plurality of fourth vectors and the labeled video description information to obtain a predicted video;

[0116] A parameter adjustment module 340, configured to adjust parameters of the adapter network module to be trained according to the predicted video and the labeled video corresponding to each sample image until a preset training end condition is satisfied, and obtain the video generation model by the model to be trained.

[0117] In some possible implementation manners of the present disclosure, the processing of the plurality of fourth vectors and the labeled video description information by the video generation network module further includes:

[0118] For the denoising process at each time step, the first second network layer in the video generation network module processes the labeled video description information to obtain a fifth vector output by the first second network layer;

[0119] For the second to the nth second network layers in the video generation network module, the second network layer processes the labeled video description information and the sixth vector to obtain a fifth vector output by the second network layer, where the sixth vector is obtained by adding the fifth vector output by the previous second network layer of the second network layer and the fourth vector output by the first network layer corresponding to the previous second network layer;

[0120] Based on the fifth vector output by the nth second network layer in the denoising process at the last time step, the predicted video is determined.

[0121] The apparatus for generating a video using an image according to an embodiment of the present disclosure corresponds to the embodiment of the method for generating a video using an image of the present disclosure above, and the relevant content can be referred to each other, and will not be elaborated here.

[0122] For the beneficial technical effects corresponding to the exemplary embodiments of the apparatus for generating a video using an image according to an embodiment of the present disclosure, reference may be made to the corresponding beneficial technical effects in the corresponding exemplary method part above, and will not be elaborated here.

[0123] In addition, an embodiment of the present disclosure further provides an electronic device, including:

[0124] a memory for storing a computer program;

[0125] a processor for executing the computer program stored in the memory, and when the computer program is executed, implementing the method for generating a video using an image according to any one of the above embodiments of the present disclosure.

[0126] Figure 7 FIG. is a schematic structural diagram of an application embodiment of the electronic device of the present disclosure. Next, reference is made to Figure 7 to describe the electronic device according to an embodiment of the present disclosure. The electronic device may be any one or both of the first device and the second device, or a stand-alone device independent of them, and the stand-alone device may communicate with the first device and the second device to receive the input signals collected from them.

[0127] As Figure 7 shown, the electronic device includes one or more processors and a memory.

[0128] The processor may be a central processing unit (CPU) or other form of processing unit having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0129] The memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor may run the program instructions to implement the method of generating a video using an image and / or other desired functions of the various embodiments of the present disclosure described above.

[0130] In one example, the electronic device may further include: an input device and an output device, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0131] In addition, the input device may further include, for example, a keyboard, a mouse, and so on.

[0132] The output device may output various information to the outside, including the determined distance information, direction information, etc. The output device may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.

[0133] Of course, for simplicity, Figure 7 only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.

[0134] In addition to the above methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the method of generating a video using an image according to various embodiments of the present disclosure described in the above part of this specification.

[0135] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0136] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium storing computer program instructions, which, when run on a processor, cause the processor to execute the steps in the method of generating a video using an image according to various embodiments of the present disclosure described in the foregoing part of this specification.

[0137] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but not be limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0138] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium, and when executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disk, or optical disk and other various media that can store program codes.

[0139] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for the purposes of illustration and facilitating understanding, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0140] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system embodiments, since they basically correspond to the method embodiments, the description is relatively simple, and reference can be made to the partial description of the method embodiments for the relevant parts.

[0141] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.

[0142] The methods and apparatuses of this disclosure can be implemented in many ways. For example, the methods and apparatuses of this disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the methods is for illustration only, and the steps of the methods of this disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, this disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to this disclosure. Therefore, this disclosure also covers the recording medium storing the programs for executing the methods according to this disclosure.

[0143] It should also be noted that in the apparatuses, equipment, and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.

[0144] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0145] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.

Claims

1. A method for generating a video using an image, characterized in that: include: Acquiring conditional guidance information, wherein the conditional guidance information includes an input image; Inputting the preset noise information and the conditional guidance information into a pre-trained video generation model to obtain a target video, wherein the target video includes the input image; wherein the input image is processed by an adapter network module in the video generation model to obtain a first vector output by each first network layer in the adapter network module, and a plurality of first vectors and the preset noise information are processed by a video generation network module in the video generation model to obtain the target video; The video generation model is obtained in the following manner: obtaining a model to be trained and sample data, the model to be trained comprising a video generation network module and an adapter network module to be trained, the sample data comprising a plurality of sample images, and a label video and label video description information corresponding to each sample image; freezing the parameters of the video generation network module, for the plurality of sample images, inputting the sample images, and the label video and label video description information corresponding to the sample images into the model to be trained, the adapter network module to be trained processes the sample images to obtain a fourth vector output by each first network layer in the adapter network module to be trained, the video generation network module processes the plurality of fourth vectors and the label video description information to obtain a predicted video; The parameters of the adapter network module to be trained are adjusted according to the predicted video and the label video corresponding to each sample image until the preset training end condition is met, and the video generation model is obtained from the model to be trained.

2. The method according to claim 1, characterized in that The conditional guidance information also includes video description information; The step of inputting the preset noise information and the conditional guidance information into a pre-trained video generation model to obtain a target video includes: The adapter network module in the video generation model processes the input image to obtain the first vector output by each first network layer in the adapter network module, and the video generation network module in the video generation model processes multiple first vectors, the preset noise information and the video description information to obtain the target video.

3. The method according to claim 1 or 2, characterized in that: The conditional guidance information for obtaining video generation includes: Get image sequence; Extracting the input image from the image sequence based on preset position information; Based on the input image, the conditional guidance information is determined.

4. The method according to claim 2, characterized in that: The adapter network module includes n first network layers, the input layer of the video generation network module includes n second network layers, and the n first network layers of the adapter network module correspond one-to-one to the n second network layers of the video generation network module respectively; The processing of the plurality of first vectors, the preset noise information and the video description information by the video generation network module in the video generation model includes: In the denoising process of each time step of the video generation network module, based on the preset noise information and the video description information, the n first vectors output by the n first network layers are respectively calculated with the corresponding n second vectors output by the n second network layers to obtain the target video.

5. The method according to claim 4, characterized in that In the denoising process at each time step of the video generation network module, based on the preset noise information and the video description information, the n first vectors output by the n first network layers are respectively calculated with the corresponding n second vectors output by the n second network layers to obtain the target video, including: For the denoising process at each time step, the first second network layer processes the preset noise information based on the video description information to obtain a second vector output by the first second network layer; For the second second network layer to the nth second network layer, the second network layer processes the third vector based on the video description information to obtain a second vector output by the second network layer, wherein the third vector is obtained by adding a second vector output by a second network layer immediately preceding the second network layer to a first vector output by a first network layer corresponding to the immediately preceding second network layer; The target video is determined based on the second vector output by the nth second network layer in the denoising process of the last time step.

6. The method according to claim 1, characterized in that The processing of the plurality of fourth vectors and the label video description information by the video generation network module includes: The first second network layer in the video generation network module processes the label video description information to obtain a fifth vector output by the first second network layer; For the denoising process at each time step, for the second second network layer to the nth second network layer in the video generation network module, the second network layer processes the label video description information and the sixth vector to obtain a fifth vector output by the second network layer, and the sixth vector is obtained by adding the fifth vector output by the second network layer immediately preceding the second network layer to the fourth vector output by the first network layer corresponding to the immediately preceding second network layer; The predicted video is determined based on the fifth vector output by the nth second network layer in the denoising process of the last time step.

7. A device for generating a video using an image, characterized in that: include: An information acquisition module, used to acquire conditional guidance information, wherein the conditional guidance information includes an input image; A video generation module, used for inputting preset noise information and the conditional guidance information into a pre-trained video generation model to obtain a target video, wherein the target video includes the input image; wherein the input image is processed by an adapter network module in the video generation model to obtain a first vector output by each first network layer in the adapter network module, and a plurality of first vectors and the preset noise information are processed by a video generation network module in the video generation model to obtain the target video; The video generation model is obtained in the following manner: obtaining a model to be trained and sample data, the model to be trained comprising a video generation network module and an adapter network module to be trained, the sample data comprising a plurality of sample images, and a label video and label video description information corresponding to each sample image; freezing the parameters of the video generation network module, for the plurality of sample images, inputting the sample images, and the label video and label video description information corresponding to the sample images into the model to be trained, the adapter network module to be trained processes the sample images to obtain a fourth vector output by each first network layer in the adapter network module to be trained, the video generation network module processes the plurality of fourth vectors and the label video description information to obtain a predicted video; The parameters of the adapter network module to be trained are adjusted according to the predicted video and the label video corresponding to each sample image until the preset training end condition is met, and the video generation model is obtained from the model to be trained.

8. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to execute the computer program stored in the memory, and when the computer program is executed, the method for generating a video using an image as described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a video using an image as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Video generation method and device, electronic equipment and readable storage medium

    CN118042246A

  • Video generation method and system, electronic equipment and storage medium

    CN118714417A