Method, apparatus, device and medium for generating video

The diffusion model-based video generation technology addresses the challenge of poor dynamics in existing methods by combining image and text instructions, resulting in videos with improved realism and reduced manual annotation requirements.

JP7785136B2Active Publication Date: 2025-12-12BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024114172
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-07-17
Publication Date
2025-12-12
Estimated Expiration
2044-07-17

AI Technical Summary

Technical Problem

Existing video generation methods using machine learning models often result in videos with poor dynamics and lack realistic visual effects, particularly in terms of object movement and complex camera movements, making it difficult to create videos with desired content.

Method used

A diffusion model-based video generation technology that combines image instructions of the first and last frames with text instructions, using a generative model trained on reference videos to generate target videos with improved dynamics and visual effects.

Benefits of technology

The proposed method generates videos with complex scenes and motions, aligning user expectations by ensuring variety and consistency, and reduces the need for manual annotation data, enhancing the efficiency and accuracy of video creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007785136000009
    Figure 0007785136000009
  • Figure 0007785136000010
    Figure 0007785136000010
  • Figure 0007785136000011
    Figure 0007785136000011
Patent Text Reader

Abstract

To provide a method, an apparatus, a device, and a medium for generating a video, in an object in the video.SOLUTION: A method includes: determining a first reference image and a second reference image from a plurality of reference images in a reference video; receiving a reference text for describing the reference video; and obtaining a generative model based on the first reference image, the second reference image, and the reference text. The generative model is used to generate a target video based on the first image, the second image, and the text, and the second reference image is used as guide data to determine the direction of story development in the video. In this way, the generative model clearly captures changes in contents of each image in the video, and helps generation of a richer and more compelling video.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD Exemplary embodiments of the present disclosure relate generally to computer vision, and more particularly to methods, apparatus, devices, and computer-readable storage media for generating video using machine learning models. [Background technology]

[0002] Machine learning techniques are widely used in many technical fields, and in the field of computer vision, many technical proposals have been proposed for automatically generating videos using machine learning models. For example, a corresponding video can be generated based on pre-specified images and text describing the video content. However, currently generated videos usually have poor dynamics (smoothness of movement). For example, objects in the video lack clear movement or dynamic effects, making it difficult to achieve realistic visual effects of movement. Therefore, a simpler and more effective method for generating videos containing desired content is desired. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for generating a video is provided, the method including: determining a first reference image and a second reference image from a plurality of reference images in a reference video; receiving reference text describing the reference video; obtaining a generative model based on the first reference image, the second reference image, and the reference text; and using the generative model to generate a target video based on the first image, the second image, and the text.

[0004] In a second aspect of the present disclosure, an apparatus for generating a video is provided, the apparatus including: an image determination module configured to determine a first reference image and a second reference image from a plurality of reference images in a reference video; a text determination module configured to receive reference text for describing the reference video; and an acquisition module configured to acquire a generative model based on the first reference image, the second reference image, and the reference text, wherein the generative model is utilized to generate a target video based on the first image, the second image, and the text.

[0005] In a third aspect of the present disclosure, there is provided an electronic device, the electronic device comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method of the first aspect of the present disclosure.

[0006] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program which, when executed by a processor, causes the processor to perform the method according to the first aspect of the present disclosure.

[0007] It should be understood that the contents described in this section are not intended to limit the main or important features of the embodiments of the present disclosure, nor do they limit the scope of the present disclosure. Other features of the present disclosure will be easily understood from the following description. [Brief explanation of the drawings]

[0008] These and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent hereinafter by reference to the following detailed description taken in conjunction with the accompanying drawings, in which the same or similar reference numerals indicate the same or similar elements. [Figure 1] 1 shows a block diagram of a technical solution for generating a video. [Figure 2]1 shows a block diagram of a process for generating video according to some embodiments of the present disclosure. [Figure 3] FIG. 1 shows a block diagram of a process for determining a second reference image according to some embodiments of the present disclosure. [Figure 4] FIG. 1 illustrates a block diagram of a process for determining a generative model according to some embodiments of the present disclosure. [Figure 5] 1 illustrates a block diagram of a process for determining reference features of a reference video according to some embodiments of the present disclosure. [Figure 6] FIG. 1 is a block diagram of a process for generating a target video according to some embodiments of the present disclosure. [Figure 7] 1 illustrates a block diagram of a process for manipulating a diffusion model in a process for generating a target video according to some embodiments of the present disclosure. [Figure 8] 1 illustrates a block diagram for generating a target video based on input data according to some embodiments of the present disclosure. [Figure 9] 1 illustrates a block diagram for generating a target video based on input data according to some embodiments of the present disclosure. [Figure 10] 1 illustrates a block diagram for generating a target video based on input data according to some embodiments of the present disclosure. [Figure 11] 1 illustrates a flowchart of a method for generating video according to some embodiments of the present disclosure. [Figure 12] 1 shows a block diagram of an apparatus for generating video according to some embodiments of the present disclosure. [Figure 13] 1 shows a block diagram of a device capable of implementing many embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Although the accompanying drawings show several embodiments of the present disclosure, it should be understood that the present disclosure can be realized in various forms and should not be construed as being limited to the embodiments described herein, but rather these embodiments are provided for a more thorough and complete understanding of the present disclosure. It should be understood that the accompanying drawings and embodiments of the present disclosure are intended to serve as examples and are not intended to limit the scope of protection of the present disclosure.

[0010] In describing embodiments of the present disclosure, the term "comprises" and similar terms should be understood as an open inclusion, i.e., "including, but not limited to." The term "based on" should be understood as "based at least in part on." The terms "one embodiment" or "this embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." Other explicit and implicit definitions may be included below. As used herein, the term "model" may indicate a relationship between data. For example, the relationship may be obtained based on numerous currently known and / or future developed technical solutions.

[0011] It should be understood that the data related to the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws and regulations and related provisions.

[0012] It should be understood that before using the technical solutions disclosed in each embodiment of this disclosure, the type, scope of use, usage scenario, etc. of personal information related to this disclosure should be notified to users by appropriate means in accordance with relevant laws and regulations, and user approval should be obtained.

[0013] For example, in response to receiving an unsolicited request from a user, a prompt message is sent to the user explicitly urging the user that the requested operation requires access to and use of the user's personal information, allowing the user to autonomously choose whether or not to provide personal information to software or hardware, such as an electronic device, application, server, or storage medium, that performs the operation of the technical solution of the present disclosure, based on the prompt message.

[0014] As an optional, non-limiting embodiment, the manner in which a prompt message is sent to the user in response to receiving an unsolicited request from the user may be, for example, via a pop-up window in which the prompt message may be presented in text, and the pop-up window may further include a selection control for the user to select "agree" or "disagree" to providing personal information to the electronic device.

[0015] It should be understood that the above process of notification and obtaining user authorization is merely exemplary and does not limit the embodiments of the present disclosure, and other methods that comply with relevant laws and regulations may also be applied to the embodiments of the present disclosure.

[0016] As used herein, the term "responsive to" refers to a state in which a corresponding event has occurred or a state in which a condition has been satisfied. It should be understood that there is not necessarily a strong correlation between the timing of the execution of a subsequent action executed in response to this event or condition and the time at which the event occurred or the condition was met. For example, in some cases, the subsequent action is executed immediately when the event occurs or the condition is met, while in other cases, the subsequent action is executed after a certain period of time has elapsed since the event occurred or the condition was met.

[0017] Example Environment Machine learning techniques are widely used in many technical fields, and in the field of computer vision, automatic video generation using machine learning models has been proposed. Traditional video generation methods focus on text-to-video generation or single-image video generation. Although the generated video may contain object movements, the movements are short-lived and dynamic for only a short period of time, resulting in a lack of dynamics and inability to present the desired information.

[0018] Referring to FIG. 1 for describing the generation method, FIG. 1 shows a block diagram 100 of a technical proposal for generating a video. As shown in FIG. 1, a machine learning model 120 is obtained, where the machine learning model 120 may be generated based on reference data in a pre-established training dataset. Text 110 may be used to specify the content of the video to be generated, and image 112 may be used to specify the environment of the video (e.g., as an image of the first frame or an image of another location). In this example, the text 110 may indicate, for example, "A cat is walking in the street." At this time, the machine learning model 120 may generate a video 130 having the content of a cat walking on the street based on the text 110 and 112.

[0019] However, currently generated video often faces difficulties in achieving realistic visual effects. Creating videos with high dynamic motion, complex camera movements, visual effects, facial close-ups, or shot conversions presents formidable challenges. Conventional video generation methods focus on text-to-video generation, tending to generate videos with minimal motion range and only capable of generating short video segments. While the morphology of an object (e.g., a cat) in a video may be similar to that of an image 112, the cat's movements may be stiff and have a small motion range. Therefore, it is expected that videos containing desired content can be generated more easily and efficiently.

[0020] Overview of generating videos Videos require an additional time dimension and consist of many keyframes, making it difficult to accurately describe each keyframe in simple language. Additionally, the types of motion in videos are highly diverse. Conventional approaches are not only source-intensive, but also pose significant challenges for generative models. Understanding complex text descriptions and generating matching videos dramatically increases the size of the model and the amount of annotation data required.

[0021] To at least partially address the shortcomings of the prior art, an exemplary embodiment of the present disclosure proposes a diffusion model-based video generation technology (e.g., called PixelDance). Generally, a diffusion model-based machine learning architecture is proposed that combines image instructions of the first and last frames in a video with text instructions for video generation. Experimental results show that the video generated using the proposed technology of the present disclosure exhibits excellent visual effects when synthesizing videos with complex scenes and complex motions.

[0022] According to an exemplary embodiment of the present disclosure, the image instructions for the first frame set the scene (and establish the characters) for video generation. The first frame also enables the model to generate successive videos, where the model generates subsequent videos using the last frame of the previous video together with the instructions for the first frame of the subsequent video. Furthermore, the instructions for the last frame, describing the end state of the video, can serve as an additional control mechanism. In this way, the alignment between user expectations and the text is strengthened (i.e., the generated video is aligned with the user's expectations), allowing the model to construct complex shots and ultimately generate rich video content, thereby ensuring variety and consistency.

[0023] Generally, the <text (text), first frame (first frame), last frame (last frame)> instructions can be used conditionally. When these three instructions are given, the model can intensively learn the dynamics of people, animals, objects, and other entities in the world during the training stage. In inference, the model can "generalize" the learned physical world movement rules to areas not covered during training, such as realizing the movement of anime characters or special effect shots.

[0024] Specifically, the above information can be integrated into the diffusion model. For example, text information is encoded by a pre-trained text encoder and embedded into the diffusion model using a cross-attention mechanism. Image instructions are encoded by a pre-trained variational autoencoder (VAE) and combined with the video content of the disturbance or Gaussian noise as the input to the diffusion model. In the training process, the image of the first frame of the video is directly used as the instruction for the first frame to force the model to strictly follow the instructions, thereby maintaining the continuity between consecutive videos. In the inference process, the instructions can be conveniently obtained from the model from text to image, or directly provided by the user.

[0025] Correspondingly, when generating a video, the model should avoid copying the instruction of the last frame. This is because it is difficult to obtain a perfect last frame in the inference process. Therefore, the model may be designed to adapt to the rough sketch provided by the user as a guide and generate the corresponding video. This rough sketch can be created using basic image editing tools. Also, the user may provide a hand-drawn contour diagram as a guide for the last frame. It should be understood that the sketches in this specification may also be sketches drawn using image editing tools.

[0026]

[0023] Referring to Figure 2 for an overview according to exemplary embodiments of the present disclosure, Figure 2 illustrates a block diagram 200 of a process for generating a video according to some embodiments of the present disclosure. As shown in Figure 2, a generative model 220 shown in Figure 2 can be constructed, which describes an association relationship between input data and output data. Here, the input data may include, for example, text for describing the content of the video, a start image (e.g., which may be referred to as a first image) for representing an image of a first frame in the video, and a guide image (e.g., which may be referred to as a second image) for representing an image of a last frame in the video. The output data may include, for example, a video.

[0027] According to an exemplary embodiment of the present disclosure, the generative model 220 may be constructed based on various currently known and / or future developed architectures. In the training phase, sample data may be extracted from a reference video 230 to further train the generative model 220. As shown in FIG. 2 , a reference video 230 including multiple reference images may be acquired, and a first reference image 212 and a second reference image 214 may be acquired from the multiple reference images. Here, the first reference image 212 (e.g., an image of the first frame) may be acquired from the head position of the reference video 230, and the second reference image 214 (e.g., an image of the final frame or an image within a predetermined range before the image of the final frame) may be acquired from the tail position of the reference video 230. Furthermore, reference text 210 for describing the reference video may be received.

[0028] 2, the reference video 230 includes a cat, and the cat in the first reference image 212 and the second reference image 214 is located at different positions in the screen. In this case, the text 110 may indicate "A cat is walking in the street." Furthermore, a generative model 220 may be trained based on the first reference image 212, the second reference image 214, and the reference text 210. In this case, the trained generative model 220 may generate a target video based on the first image, the second image, and the text.

[0029] In the process of processing the final frame instruction, multiple strategies may be adopted to adjust the influence of the final frame. During the training period, the final frame instruction may be randomly selected from the last three true frames of the reference video. To improve the robustness of the model against instruction dependency, noise may be introduced into the instruction. During the training phase, the final frame instruction may be randomly discarded with a certain probability. Correspondingly, a simple and effective inference strategy has been proposed. For example, in the τ denoising steps before the inverse process of the diffusion model, the final frame instruction may be utilized to guide the video generation to a desired end state. Following this, the remaining steps discard this final frame instruction, thereby generating more consistent video content over the model generation time. The strength of the final frame guide can be adjusted by τ.

[0030] It should be understood that FIG. 2 merely schematically illustrates the process of obtaining the generative model 220, and alternatively and / or additionally, multiple reference videos may be obtained and corresponding data extracted from the multiple reference videos to iteratively and continuously update the generative model 220. For example, reference videos in which different objects perform different actions may be obtained, and the generative model 220 may further grasp richer knowledge about video generation. According to exemplary embodiments of the present disclosure, the video generation process may be implemented in a multilingual environment. Although the reference text in FIG. 2 is expressed in English, alternatively and / or additionally, the reference text may be expressed based on other languages ​​(e.g., Chinese, French, Japanese, etc.).

[0031] Using the exemplary embodiments of the present disclosure, in the process of determining the generative model 220, the first reference image 212 may be used as the starting image of the video, and the second reference image 214 may be used as guide data to determine the direction of the story in the video. For example, the generated video may have the second reference image 214 as the final frame. In this way, the generative model 220 can clearly understand the changes in the content of each image in the video, which can help generate richer and more compelling videos.

[0032] The detailed process of generating a video

[0033] Having described an overview of generating a video, further details regarding generating a video are provided below. According to an exemplary embodiment of the present disclosure, to determine a first reference image, a reference image located at the beginning of a reference video may be determined as the first reference image. In other words, an image of a first frame in the reference video may be determined as the first reference image. In this way, the process of constructing training data is simplified, and training data can be obtained in a fast and efficient manner.

[0033] According to exemplary embodiments of the present disclosure, the second reference image may be determined based on various methods. For example, the image of the last frame in the reference video may be used as the second reference image. Alternatively and / or additionally, to introduce a disturbance factor into the training data, the second reference image may be determined from a set of reference images located within a predetermined range from the tail of the reference video.

[0034] For further details, refer to FIG. 3 , which shows a block diagram 300 of a process for determining a second reference video according to some embodiments of the present disclosure. As shown in FIG. 3 , the reference video 230 is assumed to include multiple reference images: a reference image 310 located at the position of the first frame, reference images 312, ..., 314, 316, and a reference image 318 located at the position of the last frame. For example, the predetermined range 320 may represent the last three frames (or other number) at the end of the reference video. Assuming that the reference video includes N image frames, the second reference image may be any one of frame N, frame N-1, or frame N-2 of the reference video. In this case, the second reference image may include any one of reference images 314, 316, or 318. In this way, the influence of the last frame on the training process can be somewhat weakened, resulting in a more consistent video.

[0035] For more information about the generative model, please refer to FIG. 4, which shows a block diagram 400 of a process for determining a generative model according to some embodiments of the present disclosure. As shown in FIG. 4, the generative model 220 may include an encoder model 410, a combination model 420, a diffusion model 430, and a decoder model 440. Here, the encoder model 410 and the decoder model 440 may have predetermined configurations and parameters. The parameters of the encoder model 410 and the decoder model 440 are not updated during subsequent updates of the parameters of the generative model 220. According to an exemplary embodiment of the present disclosure, a VAE and a corresponding decoder may be used.

[0036] 4, the generative model 220 may be obtained based on the first reference image 212, the second reference image 214, and the reference text 210. Specifically, the encoder model 410 is used to determine a first reference feature 412 (e.g., denoted as Z) of the reference video 230, where the first reference feature 412 may include multiple reference image features of multiple reference images. Furthermore, the encoder model 410 is used to determine a second reference feature 414 (e.g., c image, where the second reference features may include a first reference image feature of the first reference image and a second reference image feature of the second reference image.

[0037] According to an exemplary embodiment of the present disclosure, the first reference feature 412 includes multiple dimensions, each dimension corresponding to one reference image, i.e., in each dimension, a feature from one reference image (a predetermined dimension, e.g., R C*H*W where C represents the number of channels, H represents the height of the image, and W represents the width of the image. According to an exemplary embodiment of the present disclosure, the first reference feature and the second reference feature may have the same dimension, where the number of dimensions is equal to the number of images in the reference video. For example, it may be specified that the reference video includes 16 images (or another number of images).

[0038] For further details regarding feature determination, refer to FIG. 5, which illustrates a block diagram 500 of a process for determining reference features for a reference video according to some embodiments of the present disclosure. As illustrated in FIG. 5, each reference image in the reference video 230 may be processed one by one. For example, to generate a first reference image feature, an encoder model may be used to extract a feature for the first reference image 212. The first reference image feature may be placed at a first position (i.e., position 510) of the first reference feature 412. Subsequently, the encoder model may be used to extract a reference image feature for another reference image after the first reference image 212 and place the reference image feature at a second position of the first reference feature 412. Each reference image in the reference model 230 may be processed successively until the second reference image feature of the last reference image in the reference model 230 (i.e., second reference image 214) is placed at the last position (i.e., position 512) of the first reference feature 412.

[0039] Similarly, the encoder model 310 may be used to determine a second reference feature of the reference video. To generate the first reference image feature, the encoder model 410 may be used to extract features of the first reference image 212. The first reference image feature may be placed at the first position (i.e., position 520) of the second reference feature 442. The encoder model 410 may be used to place the second reference image feature of the last reference image (i.e., second reference image 214) in the reference video 230 at the last position 522 of the second reference feature 414.

[0040] 5, the position (i.e., the initial position) of the first reference image feature in the second reference feature 414 corresponds to the position (i.e., the initial position) of the first reference image in the reference video, and the position (i.e., the final position) of the second reference image feature in the second reference feature 414 corresponds to the position (i.e., the final position) of the second reference image in the reference video. In this way, it is convenient to align the first and second reference features in the subsequent combining process, which can facilitate improving the accuracy of the diffusion model.

[0041] According to an exemplary embodiment of the present disclosure, features at positions other than the first and second positions in the second reference features may be set to NULL. Specifically, position 530 in FIG. 5 may be set to NULL, e.g., filled with data "0." In this way, the second reference features can enhance the influence of the first and second reference images on the video generation process, thereby allowing the generative model to learn more knowledge about the transformation relationship between the first and second reference images.

[0042] Specifically, the reference video can be encoded using a VAE to obtain the first reference feature 412, that is, each frame of the reference video can be encoded by a VAE respectively. The first reference image and the second reference image can be encoded using a VAE, and then the positions of the intermediate frames are filled with 0, and the second reference feature 414 (having the same dimension as the first reference feature 412) can be obtained.

[0043] Returning to FIG. 4 , the diffusion model 430 may be determined based on the first reference feature 412, the second reference feature 414, and the reference text 210. Specifically, based on the principle of the diffusion model, noise may be added to the first reference feature 412 and the second reference feature 414, respectively (e.g., different levels of noise may be added to the two features using different methods), and a first noisy reference feature 412′ and a second noisy reference feature 414′ may be generated. Furthermore, a combination model 420 may be used to combine the first noisy reference feature 412′ and the second noisy reference feature 414′ to generate a feature 422. Furthermore, corresponding training data may be constructed based on the principle of the diffusion model, and the parameters of the diffusion model 430 may be updated.

[0044] According to exemplary embodiments of the present disclosure, the diffusion model 430 may be implemented based on various currently known and / or future developed structures. The diffusion process of the diffusion model 430 involves a forward noise generation process and a backward noise reduction process. In the forward noise generation process, noise data may be added to the features separately in multiple steps. For example, the initial feature may be denoted as Z_0. Noise data may be added continuously in each step, and the noise feature at step t may be denoted as Z_t, and the noise feature at step t+1 may be denoted as Z_(t+1). For example, in step t, noise data may be added to Z_t. This may be followed by random Gaussian data in step T.

[0045] In the backward denoising process, the inverse process of the above-mentioned noisy process may be performed in multiple steps to obtain initial features step by step. Specifically, the random Gaussian data, the reference text, and the corresponding second reference features may be input to the diffusion model to perform the denoising process step by step in multiple steps. For example, some noise in the noise feature Z_t may be removed in the t-th step to form a noise feature Z_(t-1) that is cleaner than the noise feature Z_t. Each denoising process may be performed iteratively to infer the initial features (i.e., video features without added noise).

[0046] According to an exemplary embodiment of the present disclosure, in the process of determining the diffusion model, a processing model 420 may be utilized to combine the first reference feature and the second reference feature to generate reference features of the reference video, perform a noisy processing on the reference features to generate noisy reference features 422 of the reference video, and then utilize a diffusion model 430 to reconstruct features 432 (e.g.,

number

[0047] According to an exemplary embodiment of the present disclosure, a 2D UNet (U-shaped network) can be used to implement a diffusion model. This model can be constructed by spatially downsampling and then spatially upsampling with the insertion of skip connections. Specifically, the model includes two basic blocks: a 2D convolutional block and a 2D attention block. The 2D UNet is extended to 3D by inserting a temporal layer, where a 1D convolutional layer follows a 2D convolutional layer along the time dimension, and a 1D attention layer follows a 2D attention layer along the time dimension. This model can be trained with images or videos to maintain high-fidelity spatial generative capabilities and disable 1D temporal manipulation of image inputs. Bidirectional self-attention can be used in all attention layers. Text instructions are encoded using a pre-trained text encoder and embedded in the c text is injected through the cross-attention layer of UNet, hiding the state as a query, and text Let be the key and value.

[0048] Regarding image command injection, the image commands of the first and last frames can be combined with text instructions. For example, first ,I last Given image instructions on the first and last frames, expressed as {f first ,f last} can be obtained, where

number

number

number

[0049] According to an exemplary embodiment of the present disclosure, the first reference image, the second reference image, and the reference text are determined from multiple reference videos, and then the first reference feature Z, the second reference feature c, and the like of each reference video are determined using the above-described aspects. image , and the corresponding reconstruction features

number

number

[0050] The diffusion model 430 may be iteratively and continuously updated using multiple reference videos until a desired stopping condition is met. It should be understood that in this process, no manual annotation data is required; rather, the first and last frames of the reference video and the corresponding text descriptions can be directly used as training data. In this way, the workload of obtaining manual annotation data can be eliminated, thereby improving the efficiency of obtaining the generative model 220.

[0051] The above is the second reference feature c image Although the example has been given in which the first reference image and the second reference image simultaneously contain related feature information, alternatively and / or additionally, in some cases, the second reference image may contain related feature information c image It should be understood that the feature information of the first reference image may include only the feature information of the first reference image. Specifically, the feature of the second reference image in the second reference feature may be set to NULL according to a predetermined condition. In other words, in FIG. 5, the feature at position 522 may be removed, for example, the value at that position may be set to 0.

[0052] Using exemplary embodiments of the present disclosure, the second reference image features may be removed by a predetermined percentage (e.g., 25% or other value). In this way, the process of updating the generative model can include application scenarios that do not consider the second reference image, making the generative model suitable for a wider range of application scenarios and improving the accuracy of the generative model in various situations.

[0053] According to exemplary embodiments of the present disclosure, the generative model 220 may be trained based on various currently known and / or future developed methods to capture associative relationships between the first image, the second image, the text describing the content of the video, and the corresponding video.

[0054] According to an exemplary embodiment of the present disclosure, the generative model 230 further includes a decoder model 440. After the generative model 220 is obtained, a need for generating a video may be input to the generative model 220, thereby generating a corresponding target video. According to an exemplary embodiment of the present disclosure, a first image and text may be input to the generative model, and the generative model may then be used to generate the corresponding target video. According to an exemplary embodiment of the present disclosure, after a first image, a second image, and text are input to the generative model, the generative model may then be used to generate the corresponding target video. Alternatively and / or additionally, the text may be null, and a target video may still be generated in that case.

[0055] To obtain a more expected target video, a first image and a second image for generating the target video and text describing the content of the target video may be received. Subsequently, a reconstructed feature of the target video may be generated based on the first image, the second image, and the text according to the trained diffusion model. Further, the target video may be generated based on the reconstructed feature according to the decoder model. In this way, the diffusion model 430 can be fully utilized to perform the backward denoising process, thereby obtaining the reconstructed feature of the target video in a more accurate manner.

[0056] For further details regarding video generation, please refer to FIG. 6 , which shows a block diagram 600 of a process for generating a target video according to some embodiments of the present disclosure. As shown in FIG. 6 , a first image 612, a second image 614, and text 610 describing the content of the images may be received. Here, the first image 612 may be, for example, the first frame of a target video 630 to be generated, including an ice surface. The second image 614 may be, for example, the ice surface and a polar bear on the ice surface. The second image 614 may be used as guide information to generate the final frame of the target video 630. In this case, the generative model 220 will generate the target video 630 under the conditions of the text 610, and the content of the video will be "A polar bear is walking on the ice surface."

[0057] The first image 612, the second image 614, and the text 610 may be input to the generative model 220. The generative model 220 then outputs a target video 630, where the first frame of the target video 630 corresponds to the first image 612 and the last frame corresponds to the second image 614. It should be understood that at this point, the target video 630 has not yet been generated, so noise features are used as the first features 620 of the target video 630. According to an exemplary embodiment of the present disclosure, Gaussian noise may be used to determine the first features 620.

[0058] Furthermore, a second feature 622 of the target video 630 may be generated based on the first image 612 and the second image 614 according to the encoder model 410. For example, the encoder model 410 may be used to determine a first image feature of the first image 612 and a second image feature of the second image 614, respectively. Subsequently, the first image feature may be placed at the first position of the second feature 622, and the second image feature may be placed at the last position of the second feature 622. It should be understood that the process of generating the second feature 622 is similar to that of the training stage, and therefore a detailed description thereof will be omitted.

[0059] If the first feature 620 and the second feature 622 have already been determined, the reconstructed feature 626 of the target video may be obtained based on the diffusion model using the first feature 620, the second feature 622, and the text 610. Similar to the process of determining the reconstructed feature in the training phase, the first noise feature 620′ and the second noise feature 622′ may be obtained, and the noise feature 624 may be formed by combining the first noise feature 620′ and the second noise feature 622′ using the combination model 420. Furthermore, the backward denoising stage of the diffusion model 430 may be used to generate the reconstructed feature 626.

[0060] It should be appreciated that, because the diffusion model 430 is trained based on the loss function described above, the reconstructed features 626 generated by the diffusion model 430 can then accurately describe various aspects of the information in the target video 630. Furthermore, the reconstructed features 626 can be input to the decoder model 440, which then outputs a corresponding target video 630 that accurately matches the first image 612, the second image 614, and the text 610 in the input data.

[0061] In the process of generating a video, the backward denoising process of the diffusion model 430 may be performed iteratively, and it should be understood that this may be referred to FIG. 7 for a more detailed description. FIG. 7 illustrates a block diagram 700 of a diffusion model operation process for generating a target video according to some embodiments of the present disclosure. As shown in FIG. 7, the backward denoising process may be performed iteratively in multiple steps. Assuming that the total number of multiple steps is T, in the first step (t=1), the first feature 620, the second feature 622, and the text 610 may be used to obtain reconstructed features 626 of the target video.

[0062] Following this, in the second step (t=2) after the first step in the multiple steps, the first feature is set as the reconstructed feature of the target video, and Z is

number

[0063] According to an exemplary embodiment of the present disclosure, the guiding role of the second image 614 can be further weakened. Specifically, the portion of the second feature 622 related to the second image 614 may be considered in only some of the steps. For example, in one set of steps of the multiple steps, the portion may be removed from the second feature, i.e., the portion of the second feature corresponding to the second image may be set to NULL. In FIG. 6, the feature of the second feature 622 in the last position may be set to 0.

[0064] Specifically, in τ steps prior to all T denoising steps, a final frame condition is applied to guide the video generation towards a desired end state, and the final frame condition is discarded in subsequent steps, thereby generating more plausible and temporally consistent video.

number

[0065] According to an exemplary embodiment of the present disclosure, the number of steps in which the second image 614 is not considered may be determined by a predetermined ratio. For example, the influence of the second image 614 may not be considered in τ=T*40% (or other value) of all T steps. For example, if T=50, the complete second features 622 may be used in the first 30 (50*(1-40%)) steps, and the value of the last position of the second features 622 may be set to 0 in the last 20 (50*40%) steps. In this way, the role of the first image 612 and the text 610 in generating the target video is strengthened, and as a result, the generated target video 630 may be more consistent with expectations.

[0066] According to an exemplary embodiment of the present disclosure, the second image is an image that designates the end image of the target video. For further details, refer to FIG. 8 , which shows a block diagram 800 for generating a target video based on input data according to some embodiments of the present disclosure. As shown in FIG. 8 , the input data 810 may include a first image 812, a second image 814, and text 816. Here, the second image 814 guides the generation of a final frame in the target video 820. For example, the final frame may be the same as the second image 814, where the first frame of the target video 820 corresponds to the first image 812, the final frame corresponds to the second image 814, and the content of the video is "A polar bear is walking on the ice surface."

[0067] FIG. 9 illustrates a block diagram 900 for generating a target video based on input data according to some embodiments of the present disclosure. As shown in FIG. 9, input data 910 may include a first image 912, a second image 914, and text 916. Here, the second image 914 may guide the generation of a final frame of a target video 920. For example, the final frame may be the same as the second image 914, in which case the first frame of the target video 920 corresponds to the first image 912, the final frame corresponds to the second image 914, and the video content is "a polar bear is walking on the ice surface, fireworks are bursting open, an astronaut is following the bear." Using exemplary embodiments of the present disclosure, the second image accurately depicts the ending screen of the target video, thereby allowing for more precise control of video generation.

[0068] According to an exemplary embodiment of the present disclosure, the second image may be a sketch that specifies the content of the end image of the target video. It should be understood that in some cases, it may be difficult to obtain the end image of the target video to be generated, and the content of the end image may then be specified in the form of a sketch (e.g., hand-drawn). FIG. 10 illustrates a block diagram 1000 for generating a target video based on input data according to some embodiments of the present disclosure. As shown in FIG. 10 , the input data 1010 may include a first image 1012, a second image 1014, and text 1016. Here, the second image 1014 can guide the generation of a final frame in the target video 1020.

[0069] In this case, the polar bear in the final frame may wave its hand in the pose specified in the second image 1014, where the first frame of the target video 1020 corresponds to the first image 1012, the final frame corresponds to the second image 1014, and the content of the video is "a polar bear walking on the ice surface gradually stands up and waves." Using exemplary embodiments of the present disclosure, the pose and position of objects in a video can be specified in a simpler and more efficient manner, and further facilitates the generation of videos with richer visual content.

[0070] According to an exemplary embodiment of the present disclosure, multiple videos may be generated continuously. Specifically, when a target video is generated using the above-described method, the end image of the target video may be set as the first image of another target video. Furthermore, a second image for generating another target video and another text describing the content of the other target video may be received, and the other target video may be generated based on the first and second images of the other target video and the another text according to the generative model.

[0071] Specifically, when the target video 1020 is acquired, the end image 1030 in the target video 1020 may be used as the first image for generating the next video. Furthermore, a second image (e.g., an image of a polar bear sleeping on an ice surface) may be input, and the text "a polar bear gradually lies down and falls asleep" may be input. In this case, the new video to be subsequently generated starts from the end image 1030 in the target video 1020, and the final frame of the new video corresponds to the image of the polar bear sleeping.

[0072] According to exemplary embodiments of the present disclosure, a target video may be combined with another target video. Using exemplary embodiments of the present disclosure, the ending image of a current video may be used as the starting image of the next video to be generated in a step-by-step and cumulative manner. In this way, a longer video with a richer plot may be gradually generated. Using exemplary embodiments of the present disclosure, the complexity of video creation may be greatly simplified, for example, by constructing shorter videos in the manner of sub-shots in a movie, and then editing these videos to generate a longer video.

[0073] To summarize the above, according to an exemplary embodiment of the present disclosure, a novel video generation architecture based on a diffusion model, PixelDance, is proposed, which combines image instructions and text instructions for the first and last frames. Corresponding training and inference processes can be implemented based on the PixelDance architecture. Exemplary embodiments of the present disclosure can be utilized to provide flexible control of the video generation process and a last frame instruction with varying intensity.

[0074] The proposed technology can be implemented on a large number of datasets, and experiments have shown that videos generated using PixelDance have strong advantages in synthesizing videos with complex scenes and / or motion. Specifically, a generative model can be trained using various available datasets, and each video in the dataset is associated with a pair of text, which usually provides a rough description and exhibits a weak correlation with the video content. Because image instructions are useful for learning complex video distributions, PixelDance can fully utilize various datasets without annotating the data and demonstrate excellent capabilities in generating videos with complex scenes and motion.

[0075] Example Process 11 shows a flowchart of a method 1100 for generating a video according to some embodiments of the present disclosure. In box 1110, a first reference image and a second reference image are determined from a plurality of reference images in a reference video. In box 1120, reference text for describing the reference video is received. In box 1130, a generative model is obtained based on the first reference image, the second reference image, and the reference text, and the generative model is used to generate a target video based on the first image, the second image, and the text.

[0076] According to an exemplary embodiment of the present disclosure, determining the first reference image includes determining a reference image located at the head of the reference video as the first reference image.

[0077] According to an exemplary embodiment of the present disclosure, determining the second reference image includes determining the second reference image from a set of reference images located within a predetermined range from the tail of the reference video.

[0078] According to an exemplary embodiment of the present disclosure, the generative model comprises an encoder network and a diffusion model, and obtaining the generative model based on the first reference image, the second reference image, and the reference text includes: using the encoder model to determine a first reference feature of the reference video, the first reference feature including multiple reference image features of the multiple reference images; using the encoder model to determine a second reference feature of the reference video, the second reference feature including the first reference image feature of the first reference image and the second reference image feature of the second reference image; and determining the diffusion model based on the first reference feature, the second reference feature, and the reference text.

[0079] According to an exemplary embodiment of the present disclosure, a first position of a first reference image feature in a second reference feature corresponds to a position of a first reference image in a reference video, and a second position of a second reference image feature in a second reference feature corresponds to a position of a second reference image in a reference video.

[0080] According to an exemplary embodiment of the present disclosure, the dimension of the second reference feature is equal to the dimension of the first reference feature, and features at positions other than the first position and the second position in the second reference feature are set to NULL.

[0081] According to an exemplary embodiment of the present disclosure, the method further includes setting the second reference image feature to NULL according to a predetermined condition.

[0082] According to an exemplary embodiment of the present disclosure, determining a diffusion model based on the first reference features, the second reference features, and the reference text includes: performing noise processing on the first reference features and the second reference features, respectively, to generate first noisy reference features and second noisy reference features of the reference video; combining the first noisy reference features and the second noisy reference features to generate noisy reference features of the reference video; using the diffusion model to determine reconstructed features of the reference video based on the noisy reference features and the reference text; and updating the diffusion model based on differences between the reconstructed features and the reference text.

[0083] According to an exemplary embodiment of the present disclosure, the generative model further comprises a decoder model, and the method further includes receiving a first image for generating a target video and text describing content of the target video, generating reconstructed features of the target video based on the first image and the text according to a diffusion model, and generating the target video based on the reconstructed features according to the decoder model.

[0084] According to an exemplary embodiment of the present disclosure, generating reconstruction features of the target video further includes receiving a second image for generating the target video, and generating reconstruction features of the target video based on the second image.

[0085] According to an exemplary embodiment of the present disclosure, generating reconstructed features of the target video includes using noise features as first features of the target video; generating second features of the target video based on the first image and the second image according to an encoder model; and obtaining reconstructed features of the target video using the first features, the second features, and the text according to a diffusion model.

[0086] According to an exemplary embodiment of the present disclosure, obtaining reconstructed features of a target video includes, in a first step of a plurality of steps, obtaining reconstructed features of the target video using a first feature, a second feature, and text; in a second step of the plurality of steps after the first step, setting the first feature as a reconstructed feature of the target video; and in the second step, obtaining reconstructed features of the target video based on the first feature, the second feature, and text according to a diffusion model.

[0087] According to an exemplary embodiment of the present disclosure, the method further includes, in a set of steps of the plurality of steps, setting a portion of the second feature corresponding to the second image to NULL.

[0088] According to an exemplary embodiment of the present disclosure, the second image includes at least one of an image for specifying an end image of the target video and a sketch for specifying the content of the end image of the target video.

[0089] According to an exemplary embodiment of the present disclosure, the method further includes setting an end image of the target video as a first image of another target video, receiving a second image for generating the other target video and another text describing the content of the other target video, and generating the other target video based on the first image and the second image of the other target video and the another text according to the generative model.

[0090] According to an exemplary embodiment of the present disclosure, the method further includes combining the target video with another target video.

[0091] Exemplary Devices and Equipment 12 shows a block diagram of an apparatus 1200 for generating a video according to some embodiments of the present disclosure. The apparatus includes an image determination module 1210 configured to determine a first reference image and a second reference image from a plurality of reference images in a reference video, a text receiving module 1220 configured to receive reference text for describing the reference video, and an acquisition module 1230 configured to acquire a generative model based on the first reference image, the second reference image, and the reference text, where the generative model is used to generate a target video based on the first image, the second image, and the text.

[0092] According to an exemplary embodiment of the present disclosure, the image determination module includes a first image determination module arranged to determine a first reference image, and determining the first reference image includes determining a reference image located at a head of the reference video as the first reference image.

[0093] According to an exemplary embodiment of the present disclosure, the image determination module comprises a second image determination module arranged to determine a second reference image from a set of reference images located within a predetermined range from the tail of the reference video.

[0094] According to an exemplary embodiment of the present disclosure, the generative model includes an encoder model and a diffusion model, the acquisition module is arranged to include a first encoding module, a second encoding module, and a determination module, the first encoding module is arranged to determine a first reference feature of the reference video using the encoder model, the first reference feature including a plurality of reference image features of a plurality of reference images, the second encoding module is arranged to determine a second reference feature of the reference video using the encoder model, the second reference feature including a first reference image feature of the first reference image and a second reference image feature of the second reference image, and the determination module is arranged to determine the diffusion model based on the first reference feature, the second reference feature, and the reference text.

[0095] According to an exemplary embodiment of the present disclosure, a first position of a first reference image feature in a second reference feature corresponds to a position of a first reference image in a reference video, and a second position of a second reference image feature in a second reference feature corresponds to a position of a second reference image in a reference video.

[0096] According to an exemplary embodiment of the present disclosure, the dimension of the second reference feature is equal to the dimension of the first reference feature, and features at positions other than the first position and the second position in the second reference feature are set to NULL.

[0097] According to an exemplary embodiment of the present disclosure, the apparatus further comprises a setting module arranged to set the second reference image feature to NULL according to a predetermined condition.

[0098] According to an exemplary embodiment of the present disclosure, the determination module includes a noise module configured to perform noise processing on the first reference feature and the second reference feature, respectively, to generate a first noise reference feature and a second noise reference feature of the reference video; a combination module configured to combine the first noise reference feature and the second noise reference feature to generate a noise reference feature of the reference video; a reconstruction module configured to determine a reconstructed feature of the reference video based on the noise reference feature and the reference text using a diffusion model; and an update module configured to update the diffusion model based on a difference between the reconstructed feature and the reference feature.

[0099] According to an exemplary embodiment of the present disclosure, the generative model further includes a decoder model, and the apparatus further comprises: a receiving module arranged to receive a first image for generating a target video and text describing content of the target video; a generation module arranged to generate reconstructed features of the target video based on the first image and the text in accordance with the diffusion model; and a video generation module arranged to generate the target video based on the reconstructed features in accordance with the decoder model.

[0100] According to an exemplary embodiment of the present disclosure, the receiving module is further configured to receive a second image for generating the target video, and the reconstruction module is further arranged to generate a reconstruction feature of the target video based on the second image.

[0101] According to an exemplary embodiment of the present disclosure, the reconstruction module includes a setting module configured to use noise features as first features of the target video; an encoding module configured to generate second features of the target video based on the first image and the second image according to an encoder model; and a feature acquisition module configured to acquire reconstructed features of the target video using the first features, the second features, and text according to a diffusion model.

[0102] According to an exemplary embodiment of the present disclosure, the feature acquisition module includes a first reconstruction module configured to acquire reconstructed features of the target video using the first feature, the second feature, and the text in a first step of the plurality of steps; a setting module configured to set the first feature as a reconstructed feature of the target video in a second step after the first step of the plurality of steps; and a second reconstruction module configured to acquire reconstructed features of the target video based on the first feature, the second feature, and the text according to a diffusion model.

[0103] According to an exemplary embodiment of the present disclosure, the apparatus further comprises a removal module arranged to set a portion of the second feature corresponding to the second image to NULL in a set of steps of the plurality of steps.

[0104] According to an exemplary embodiment of the present disclosure, the second image includes at least one of an image for specifying an end image of the target video and a sketch for specifying the content of the end image of the target video.

[0105] According to an exemplary embodiment of the present disclosure, the apparatus further includes an image setting module configured to set an end image of the target video as a first image of another target video, the receiving module is further configured to receive a second image for generating the other target video and another text describing the content of the other target video, and the generating module is further configured to generate the other target video based on the first and second images of the other target video and the another text according to the generative model.

[0106] According to an exemplary embodiment of the present disclosure, the apparatus further comprises a combining module arranged to combine the target video with another target video.

[0107] 13 illustrates a block diagram of an apparatus 1300 capable of implementing embodiments of the present disclosure. It should be understood that the computing device 1300 illustrated in FIG. 13 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The computing device 1300 illustrated in FIG. 13 may be used to implement the methods described above.

[0108] 13, computing device 1300 is a form of general-purpose computing device. Components of computing device 1300 may include, but are not limited to, one or more processors or processing units 1310, memory 1320, storage device 1330, one or more communication units 1340, one or more input devices 1350, and one or more output devices 1360. Processing unit 1310 may be a real processor or a virtual processor and is capable of performing various processes based on programs stored in memory 1320. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 1300.

[0109] Computing device 1300 typically includes multiple computer storage media. Such media may be any available media accessible by computing device 1300, including, but not limited to, volatile and nonvolatile media, removable and non-removable media. Memory 1320 may be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or a combination thereof. Storage device 1330 may be removable or non-removable media and may include machine-readable media, such as a flash drive, a disk, or any other medium that may be used to store information and / or data (e.g., training data for training) and that may be accessible within computing device 1300.

[0110] Computing device 1300 may also include other removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 13, a disk drive for reading from and writing to a removable, non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading from and writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data medium interfaces. Memory 1320 may include a computer program product 1325 having one or more program modules arranged to perform various methods or operations of various embodiments of the present disclosure.

[0111] The communications unit 1340 enables communication with other computing devices over a communications medium. Furthermore, the functionality of the components of the computing device 1300 may be implemented as a single computing cluster or multiple computing machines that can communicate over a communications connection. Thus, the computing device 1300 may use logical connections to one or more other servers, networked personal computers (PCs), or other network nodes to operate in a networked environment.

[0112] Input device(s) 1350 may be one or more input devices, such as, for example, a mouse, a keyboard, a tracking ball, etc. Output device(s) 1360 may be one or more output devices, such as, for example, a monitor, speakers, a printer, etc. Computing device 1300 may communicate, as needed, via communication unit 1340 with one or more external devices (not shown), such as a storage device, a display device, one or more devices that allow a user to interact with computing device 1300, or any device (e.g., a network card, a modem, etc.) that allows computing device 1300 to communicate with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).

[0113] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium having computer-executable instructions stored thereon is provided, the computer-executable instructions being executed by a processor to perform the above-described method. According to an exemplary embodiment of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and having computer-executable instructions being executed by a processor to perform the above-described method. According to an exemplary embodiment of the present disclosure, a computer program product is provided having a computer program stored thereon, the program being executed by a processor to perform the above-described method.

[0114] Aspects of the present disclosure have been described herein with reference to flowcharts and / or block diagrams of methods, apparatus, devices, and computer program products implemented in accordance with the present disclosure. It will be understood that each box in the flowcharts and / or block diagrams, and combinations of boxes in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0115] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to create a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, create an apparatus that performs the functions / acts specified in one or more boxes in the flowcharts and / or block diagrams. Also, by storing these computer-readable program instructions, which cause a computer, programmable data processing apparatus, and / or other device to function in a particular manner, on a computer-readable storage medium, the computer-readable medium having the instructions stored thereon has an article of manufacture including instructions that perform each aspect of the functions / acts specified in one or more boxes in the flowcharts and / or block diagrams.

[0116] The computer readable program instructions, when loaded onto a computer, other programmable data processing apparatus, or other device, may cause a series of operational steps to be executed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions executing on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowcharts and / or block diagrams.

[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operations that may be implemented in systems, methods, and computer program products implemented according to the present disclosure. In this regard, each box in a flowchart or block diagram may represent a module, program segment, or part of an instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the boxes may occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes may actually be executed substantially in parallel, or may be executed in reverse order depending on the functionality involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system that performs a given function or operation, or in a combination of dedicated hardware and computer instructions.

[0118] Although various implementations of the present disclosure have been described above, the foregoing description is illustrative and not exhaustive, and is not limited to the various implementations disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the various implementations described. The terminology used herein is chosen to best explain the principles of the various implementations, practical applications or improvements of the technology in the market, or to enable other ordinary skilled in the art to understand the various embodiments disclosed herein.

Claims

1. determining a first reference image and a second reference image from a plurality of reference images in a reference video; receiving a reference text describing the reference video; obtaining a generative model based on the first reference image, the second reference image, and the reference text; the generative model includes an encoder model and a diffusion model, and is used to generate a target video based on the first image, the second image, and the text; Obtaining the generative model includes: utilizing the encoder model to determine first reference features of the reference video, the first reference features including a plurality of reference image features of the plurality of reference images; determining second reference features of the reference video using the encoder model, the second reference features including first reference image features of the first reference image and second reference image features of the second reference image; determining the diffusion model based on the first reference feature, the second reference feature, and the reference text. A method for generating video.

2. The method of claim 1 , wherein determining the first reference image comprises determining a reference image located at a head of the reference video as the first reference image.

3. 2. The method of claim 1, wherein determining the second reference image comprises determining the second reference image from a set of reference images located within a predetermined range from a tail of the reference video.

4. 2. The method of claim 1, wherein a first position of the first reference image feature in the second reference feature corresponds to a position of the first reference image in the reference video, and a second position of the second reference image feature in the second reference feature corresponds to a position of the second reference image in the reference video.

5. The method of claim 4 , wherein a dimension of the second reference feature is equal to a dimension of the first reference feature, and features at positions other than the first position and the second position in the second reference feature are set to NULL.

6. The method of claim 4 , further comprising setting the second reference image feature to NULL according to a predetermined condition.

7. determining the diffusion model based on the first reference feature, the second reference feature, and the reference text, performing noise processing on the first reference feature and the second reference feature, respectively, to generate first and second noisy reference features of the reference video; combining the first and second noise reference features to generate a noise reference feature of the reference video; determining a reconstruction feature of the reference video based on the noise reference feature and the reference text using the diffusion model; and updating the diffusion model based on a difference between the reconstructed features and the first reference features.

8. the generative model further includes a decoder model; The method comprises: receiving a first image for generating a target video and text describing content of the target video; generating a reconstructed feature of the target video based on the first image and the text according to the diffusion model; The method of claim 1 , further comprising: generating the target video based on the reconstruction features according to the decoder model.

9. generating reconstructed features of the target video, receiving a second image for generating a target video; The method of claim 8 , further comprising: generating the reconstructed features of the target video based on the second image.

10. The reconstruction features for generating the target video include: a first feature using a noise feature as the target video; a second feature of generating the target video based on the first image and the second image according to the encoder model; and reconstructing features to obtain the target video using the first features, the second features, and the text according to the diffusion model.

11. The reconstruction features for obtaining the target video include: a first step of the plurality of steps for obtaining the target video by using the first feature, the second feature, and the text; In a second step after the first step among the plurality of steps, a reconstruction feature for setting the first feature as the target video; reconstructing features to obtain the target video based on the first features, the second features, and the text according to the diffusion model; The method of claim 10, comprising:

12. The method of claim 11 , further comprising, in one set of steps, setting a portion of the second feature corresponding to the second image to NULL.

13. The method of claim 8 , wherein the second image includes at least one of an image for specifying an end image of the target video and a sketch for specifying content of the end image of the target video.

14. setting an end image of the target video as a first image of another target video; receiving a second image for generating the other target video and other text describing content of the other target video; The method of claim 9 , further comprising: generating the other target video based on the first and second images of the other target video and the other text according to the generative model.

15. The method of claim 14 , further comprising combining the target video with the another target video.

16. an image determination module arranged to determine a first reference image and a second reference image from a plurality of reference images in the reference video; a text determination module configured to receive a reference text for describing the reference video; an acquisition module that acquires a generative model based on the first reference image, the second reference image, and the reference text; the generative model includes an encoder model and a diffusion model, and is used to generate a target video based on the first image, the second image, and the text; Obtaining the generative model includes: utilizing the encoder model to determine first reference features of the reference video, the first reference features including a plurality of reference image features of the plurality of reference images; determining second reference features of the reference video using the encoder model, the second reference features including first reference image features of the first reference image and second reference image features of the second reference image; determining the diffusion model based on the first reference feature, the second reference feature, and the reference text. A device for generating video.

17. at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit; An electronic device that performs the method of any one of claims 1 to 14, wherein the instructions, when executed by the at least one processing unit,

18. A computer-readable storage medium having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method of any one of claims 1 to 14.

Citation Information

Patent Citations

  • Video generation program, video generation device, and video generation method

    JP2021033961A