Method and device for generating transition video by using diffusion model, and medium
By utilizing a pre-trained image-to-video diffusion model and a denoising module to generate transition frames, the problem of insufficient quality in transition video generation in existing technologies is solved, achieving high-quality, smooth, and semantically consistent transition video generation.
Patent Information
- Application Number
- CN202511199545.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-04
- Filing Date
- 2025-08-26
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies struggle to generate high-quality transition videos, particularly in terms of maintaining semantic consistency, fidelity, smoothness, and alignment with text cues. Furthermore, their reliance on self-collected video datasets hinders progress in this field.
By acquiring the start and end frames of the video and their descriptions, a pre-trained image-to-video diffusion model is used to generate potential noise. Then, a denoising module such as the U-Net model is used to generate transition frames based on this noise and descriptions, thus avoiding the training process and preserving the appearance and motion priors of the pre-trained model.
The smoothness of transition frames was improved, randomness and discontinuity were reduced, and high-quality transition videos with semantic consistency, high fidelity and alignment with text prompts were generated, achieving a zero-sample, unified generation scheme.
Smart Images

Figure CN121639503A_ABST
Abstract
Description
BACKGROUND
[0001] Diffusion models are a type of generative model used in machine learning, particularly for tasks like image synthesis, denoising, and other data generation tasks. The core idea of diffusion models is to model the data generation process as a gradual transformation from noise to a structured data point, such as an image. This is achieved by simulating a process of diffusion, where data is gradually refined from a noisy state to a clear and recognizable state.
[0002] Transition videos refer to videos specifically designed to serve as a smooth link or bridge between two different scenes, shots, or segments of content. Transition videos are commonly used in video editing, film production, and multimedia presentations to enhance the visual flow and coherence between different elements, ensuring a seamless and visually appealing transition from one scene or concept to another. SUMMARY
[0003] In a first aspect according to some embodiments of the present disclosure, a method of video generation is provided. The method includes obtaining a start frame and an end frame for a video, a first caption for the start frame, and a second caption for the end frame. The method further includes generating a first latent noise in a latent space based on the start frame, and generating a second latent noise in the latent space based on the end frame. The method also includes generating a third latent noise in the latent space corresponding to a transition frame based on the first latent noise and the second latent noise. In addition, the method further includes generating the transition frame based on the third latent noise, the start frame, the end frame, the first caption, and the second caption by utilizing a pre-trained image-to-video diffusion model.
[0004] In a second aspect according to some embodiments of the present disclosure, an electronic device including a memory and a processor is provided. The memory is configured to store computer instructions that, when executed by the processor, cause the processor to obtain a start frame and an end frame for a video, a first caption for the start frame, and a second caption for the end frame. The instructions also cause the processor to generate a first latent noise in a latent space based on the start frame, and generate a second latent noise in the latent space based on the end frame. The instructions further cause the processor to generate a third latent noise in the latent space corresponding to a transition frame based on the first latent noise and the second latent noise. Additionally, the instructions also cause the processor to generate the transition frame based on the third latent noise, the start frame, the end frame, the first caption, and the second caption by utilizing a pre-trained image-to-video diffusion model.
[0005] In a third aspect according to some embodiments of the present disclosure, a non-transitory computer-readable medium is provided. The medium includes instructions stored thereon that, when executed by a processor, cause the processor to obtain a start frame and an end frame for a video, a first description of the start frame, and a second description of the end frame. The instructions further cause the processor to generate a first latent noise in a latent space based on the start frame, and a second latent noise in the latent space based on the end frame. The instructions further cause the processor to generate a third latent noise in the latent space corresponding to a transition frame based on the first latent noise and the second latent noise. Additionally, the instructions further cause the processor to generate the transition frame based on the third latent noise, the start frame, the end frame, the first description, and the second description by utilizing a pre-trained image-to-video diffusion model.
[0006] combinations of any one or more of the above described aspects. Any one or more of the aspects described herein. This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of examples are set forth in the detailed description that follows, and in part will be apparent to those with skill in the art based on the description and part will be learned by practice of the disclosure. Examples BRIEF DESCRIPTION OF DRAWINGS
[0007] Embodiments of the present disclosure can be understood from the following detailed description when read in conjunction with the appended drawings. In accordance with standard practice in the industry, various features are not drawn to scale. In fact, the dimensions of the various features can be arbitrarily increased or decreased for clarity of discussion. Some examples of the present disclosure are described with reference to the following figures.
[0008] Figure 1 An example environment in which example embodiments of the present disclosure can be implemented is shown;
[0009] Figure 2 is a flow diagram showing an example process for video generation according to some embodiments of the present disclosure;
[0010] Figure 3 is a schematic diagram showing an example framework for video generation according to some embodiments of the present disclosure;
[0011] Figure 4 is a schematic diagram showing an example of generating latent noise for a transition frame and feeding the latent noise into a denoising U-Net model according to some embodiments of the present disclosure;
[0012] Figure 5 is a schematic diagram showing an example process for generating latent noise for a transition frame according to some embodiments of the present disclosure;
[0013] Figure 6 FIG. 4 is a diagram illustrating an example of integrating a low-rank adaptation parameter set of a transition frame into a denoising U-Net model, according to some embodiments of the present disclosure;
[0014] Figure 7 FIG. 5 is a diagram illustrating an example of integrating a text embedding set of a transition frame into a denoising U-Net model, according to some embodiments of the present disclosure;
[0015] Figures 8A-8D FIG. 6 is a diagram illustrating an example of multiple conversion tasks, according to some embodiments of the present disclosure; and
[0016] Figure 9 FIG. 7 is a block diagram illustrating physical components (e.g., hardware) of a computing device with which aspects of the present disclosure can be practiced. DETAILED DESCRIPTION
[0017] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof, and in which are shown by way of illustration specific aspects or examples. These aspects can be combined, other aspects can be utilized, and structural changes can be made without departing from the scope of the present disclosure. Aspects can be practiced in various components such as hardware, software, or combinations thereof. The disclosure can be implemented in numerous ways, including as a process, an apparatus, or a system. In this specification, these implementations, or any other forms that the disclosure can take, can be referred to as “technology.” In particular, the disclosure relates to a method for converting text into images, and a system for converting text into images. Therefore, the following detailed description is not to be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims and their equivalents. The various steps in a method implementation of the present disclosure can be performed in the order shown or in a different order. Additionally, various steps can be performed concurrently or with partial concurrence. Various aspects of the disclosure are not limited to the order of the steps described. The scope of the disclosure is thus not limited to the specific implementations described herein.
[0018] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" is interpreted as "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be provided in the following description. Concepts such as "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules, or units and are not intended to limit the order or interdependence of the functions performed by these devices, modules, or units. Variations of "a" and "a plurality" mentioned in this disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless explicitly stated otherwise in the context, the modifier should be understood as "one or more". The names of messages or information exchanged between devices in embodiments of this disclosure are provided for illustrative purposes only and are not intended to limit the scope of such messages or information. Data involved in the technical solutions (including the data itself and data acquisition or use) shall comply with the requirements of applicable laws, regulations, and relevant provisions.
[0019] The success of diffusion models in image synthesis has spurred numerous diffusion-based schemes for video synthesis. Utilizing text cues, video frames, structure maps, and even motion patterns, some schemes have demonstrated impressive results in automatically generating realistic and high-fidelity videos. However, the generation of high-quality transition videos (including generating intermediate transition frames using given start and end frames and text cues as initial guidance) remains largely unexplored.
[0020] Creating realistic transition videos is a complex task. A high-quality transition generator should meet at least four criteria: 1) semantic consistency with the input frame; 2) high fidelity to the input frame; 3) smoothness across the generated frames; and 4) alignment with the provided text cues. Furthermore, research on transition generation often relies on self-collected, carefully curated, and not publicly accessible videos, which further hinders progress in the field.
[0021] Most existing related schemes address the challenge of transition generation using two approaches. The first approach focuses on deformation, given images of two topologically similar objects. Recent schemes employ various depth interpolation techniques to generate plausible object-level transitions. However, these schemes produce intermittent images rather than temporally coherent video frames, resulting in a loss of smoothness, especially when dealing with moving objects. The second approach focuses on video frame interpolation. Most related schemes in this category attempt to estimate intermediate optical flow during training or utilize frame conditioning. However, this approach often produces unreliable object transitions with abrupt content changes, struggles to generate long transition sequences, and requires time-consuming training on large-scale moving video datasets.
[0022] Therefore, embodiments of this disclosure provide a scheme for generating video. In this scheme, a computing device can acquire a start frame and an end frame for the video, a first description of the start frame, and a second description of the end frame. The computing device can generate first latent noise in a latent space based on the start frame and second latent noise in the latent space based on the end frame. Then, the computing device can generate a third latent noise in the latent space corresponding to a transition frame based on the first and second latent noise. Thus, the computing device can generate a transition frame based on the third latent noise, the start frame, the end frame, the first description, and the second description by utilizing a pre-trained image-to-video diffusion model.
[0023] In this way, the potential noise used to generate transition frames can include information from both the starting and ending frames. Therefore, the smoothness of the generated transition frames can be improved, and the randomness and discontinuity of the generated transition frames can be reduced. Furthermore, by utilizing a pre-trained image-to-video diffusion model, this scheme eliminates the need for a training or fine-tuning process, thus preserving the appearance and motion priors of the pre-trained model.
[0024] Figure 1 An example environment 100 in which exemplary embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, environment 100 includes computing device 102. Computing device 102 can be any device with computing capabilities. For example, computing device 102 can include, but is not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multiprocessor systems, consumer electronics, wearable computer devices, smart home devices, minicomputers, mainframe computers, edge computing devices, and distributed computing environments that include any of the above systems or devices.
[0025] like Figure 1As shown, a pre-trained image-to-video diffusion model 104 can be deployed on computing device 102. The diffusion model 104 is a generative model trained on a large dataset to generate a sequence of video frames based on one or more images. The diffusion model 104 utilizes the principles of diffusion models, which progressively refine noisy inputs into high-quality outputs. In environment 100, the diffusion model 140 can receive a start frame 112, a description 114 of the start frame 112, an end frame 122, and a description 124 of the end frame 122. The diffusion model 140 can then generate a sequence of transition frames 130-1, 130-2, ..., 130-N (also collectively referred to as transition frames 130) forming a coherent video 132. The video 132 can be formed from the start frame 112, transition frames 130, and end frame 122.
[0026] For example, starting frame 112 could be an image of a cat facing right, and description 114 could be "The cat is sitting and facing right." Similarly, ending frame 122 could be an image of a dog facing forward, and description 124 could be "The dog is sitting and facing forward." In transition frames 130-1, the cat's head can turn slightly to the left, and the cat's appearance can begin to morph into that of a dog. In transition frames 130-N, the cat's head can turn almost completely forward, and the cat's appearance can be almost completely morphed into that of a dog.
[0027] In some related schemes, image-to-video diffusion models apply a binary mask and broadcast the mask to the size of the latent code. Conditional input is given by concatenating the original latent code, the binary mask, and the masked latent code along the channel dimension. The conditional image is concatenated with initial noise for each frame to preserve visual detail. However, during the inference phase, the latent code is randomly initialized for denoising. This naive initialization strategy results in random and abrupt content flickering at the object level, compromising the fidelity of the generated frames.
[0028] In environment 100, computing device 102 can generate latent noise 116 in the latent space based on start frame 112 and latent noise 126 in the latent space based on end frame 122. The latent space refers to a low-dimensional representation of data in a machine learning model, particularly in the context of generative models such as autoencoders, generative adversarial networks (GANs), and variational autoencoders (VAEs). In the latent space, complex data such as images or text can be encoded into a more abstract, compressed form, capturing essential features while discarding less important details. Latent noise is noise residing in the latent space. In some implementations, the latent noise 116 for start frame 112 and the latent noise 126 for end frame 126 can be generated by reversing the denoising process used to generate start frame 112 and end frame 122. Because latent noise 116 and 126 are generated from start frame 112 and end frame 122, they can include information from start frame 112 and end frame 122.
[0029] In environment 100, computing device 102 can generate latent noise 128-1, 128-2, ..., 128-N (also collectively referred to as latent noise 128) based on latent noise 116 for start frame 112 and latent noise 126 for end frame 122. During the generation of latent noise 128, the contributions of start frame 112 and end frame 122 to latent noise 128 can be different. For example, during the generation of latent noise 128-1, the contribution of latent noise 116 can be greater than the contribution of latent noise 126. Furthermore, during the generation of latent noise 128-N, the contribution of latent noise 126 can be greater than the contribution of latent noise 116.
[0030] In environment 100, the denoising module 106 of diffusion model 104 can receive latent noise 128 and generate a transition frame 130 based on latent noise 128, start frame 112, description 114, end frame 122, and description 124. The denoising module 106 in diffusion model 104 is the component responsible for progressively removing noise from latent noise 128 during the reverse process of diffusion model 104. The denoising module 106 can be implemented as a neural network trained to predict and subtract noise added during the forward diffusion process, thereby recovering or generating a clean, high-quality image from a noisy input. Examples of the denoising module 106 may include, but are not limited to, U-Net models, residual networks, Transformer-based models, denoising autoencoders, etc. In embodiments of this disclosure, a denoising U-Net model can be used as an example of a denoising module.
[0031] In this way, the potential noise 128 used to generate the transition frame 130 can include information from both the start frame 112 and the end frame 122, improving the smoothness of the transition and reducing randomness and discontinuity. Furthermore, by using a pre-trained image-to-video diffusion model 104, no training or fine-tuning is required, preserving the appearance and motion knowledge from the pre-trained model.
[0032] Figure 2 This is a flowchart illustrating an example process 200 for video generation according to some embodiments of the present disclosure. Process 200 can be generated by a computing device (e.g., Figure 1 This is achieved through the computing device 102 in the computer. Figure 2 As shown, at box 202, the computing device can obtain the start and end frames of the video, a first description of the start frame, and a second description of the end frame. For example, in Figure 1 In environment 100, the computing device can acquire a start frame 112, a description 114 of the start frame 112, an end frame 122, and a description 124 of the end frame 122. For example, the start frame 112 could be an image of a cat facing right, and the description 114 could be "The cat is sitting and facing right." Similarly, the end frame 122 could be an image of a dog facing forward, and the description 124 could be "The dog is sitting and facing forward."
[0033] At box 204, the computing device can generate a first latent noise in the latent space based on the start frame and a second latent noise in the latent space based on the end frame. For example, in Figure 1 In environment 100, computing device 102 can generate latent noise 116 based on start frame 112. Furthermore, computing device 102 can generate latent noise 126 based on end frame 122. Because latent noise 116 and 126 are generated from start frame 112 and end frame 122, latent noise 116 and 126 can include information from start frame 112 and end frame 122.
[0034] At box 206, the computing device can generate a third latent noise in the latent space corresponding to the transition frame based on the first and second latent noise. For example, in Figure 1 In environment 100, computing device 102 can generate latent noise 128 based on latent noise 116 for start frame 112 and latent noise 126 for end frame 122. Latent noise 128 can be used to generate transition frame 130. For example, latent noise 128-1 can be used to generate transition frame 130-1, and latent noise 128-2 can be used to generate transition frame 130-2, etc. Because latent noise 116 and 126 include information about start frame 112 and end frame 122, the generated latent noise 128 can also include information about start frame 112 and end frame 122.
[0035] At box 208, the computing device can generate a transition frame based on a third latent noise, a start frame, an end frame, a first description, and a second description by utilizing a pre-trained image-to-video diffusion model. For example, in environment 100, denoising module 106 can generate transition frame 130 based on latent noise 128, start frame 112, description 114, end frame 122, and description 124. Start frame 112, transition frame 130, and end frame 122 can be video frames of video 132.
[0036] In this way, the potential noise used to generate transition frames can include information from both the starting and ending frames. Therefore, the smoothness of the generated transition frames can be improved, and the randomness and discontinuity of the generated transition frames can be reduced. Furthermore, by utilizing a pre-trained image-to-video diffusion model, this scheme eliminates the need for a training or fine-tuning process, thus preserving the appearance and motion priors of the pre-trained model.
[0037] Figure 3 This is a schematic diagram illustrating an example frame 300 for video generation according to some embodiments of the present disclosure. Figure 3 As shown, the framework 300 includes an image encoder 302, an image encoder 304, a latent noise generation module 306, a denoising module 308, a low-rank adaptation module 310, a text encoder 312, a text embedding generation module 314, and an image decoder 316. Figure 3 As shown, the start frame 320 may be provided with a description 322 (e.g., "The cat is sitting and facing right"). Furthermore, the end frame 330 may be provided with a description 332 (e.g., "The dog is sitting and facing forward").
[0038] In frame 300, image encoder 302 can encode the start frame 320 into an image embedding, and image encoder 304 can encode the end frame 330 into an image embedding. Image encoders 302 and 304 can be neural networks used to compress rich and high-dimensional information of the input image into a set of features that capture the basic features of the input image. For example, image encoders 302 and 304 can be convolutional neural networks (CNNs), residual networks, visual transformers, or autoencoders, etc.
[0039] The latent noise generation module 306 can generate latent noise corresponding to the starting frame 320 based on the image embedding of the starting frame 320 (e.g., Figure 1 The latent noise 116 in the image is generated based on the image embedding of the end frame 330, and the latent noise corresponding to the end frame 330 is generated (e.g., Figure 1The latent noise 126). Additionally, the latent noise generation module 306 can generate latent noise corresponding to the transition frame to be generated based on the latent noise corresponding to the start frame and the latent noise corresponding to the end frame (e.g., latent noise 126). Figure 1 Potential noise in (128).
[0040] like Figure 3 As shown, the start frame 320, description 322, end frame 330, and description 332 can be input into the low-rank adaptation module 310. Low-rank adaptation is a technique used in machine learning to efficiently fine-tune large pre-trained models with fewer parameters and computational resources. Low-rank adaptation can introduce low-rank approximations into changes in model weights during fine-tuning, allowing for more efficient adaptation to new tasks or datasets without needing to update all parameters of the model. The low-rank adaptation module 310 can be trained to determine the low-rank adaptation parameters for the transition frames to be generated. The low-rank adaptation parameters can be integrated into the denoising module 308. In this way, the semantic similarity between the generated transition frames and the input frames can be improved.
[0041] The text encoder 312 can encode text into a vector representation called a text embedding. Text embeddings capture the basic semantic and syntactic information of the input text. For example... Figure 3 As shown, the text encoder 312 can encode the description 322 of the start frame 320 into a text embedding, and also encode the description 332 of the end frame 330 into a text embedding. The text embedding can include information from the start frame 320 and the end frame 330. Then, the text embedding generation module 314 can generate the text embedding of the transition frame to be generated based on the text embedding corresponding to the description 322 and the text embedding corresponding to the end frame 330. The generated text embedding can be integrated into the denoising module 308. In this way, the alignment between the generated transition frame and the input description can be improved.
[0042] like Figure 3 As shown, a denoising module 308 (e.g., a denoising U-Net model) integrating low-rank adaptation parameters and text embeddings can generate image embeddings for transition frames based on latent noise generated by the latent noise generation module 306. These image embeddings can then be input into an image decoder 316. The image decoder 316 can transform the low-dimensional image embeddings back into a high-dimensional image. Figure 3 As shown, the image decoder 316 can embed and decode these images into transition frames 342, 344, and 346. Therefore, the start frame 320, transition frames 342, 344, and 346, and the end frame 330 can be used to generate a transition video.
[0043] In this way, the semantic similarity between the generated transition frames and the input frames can be improved, the fidelity of the generated transition frames can be improved, the smoothness across the generated transition frames can be improved, and the alignment of the generated transition frames with the provided description can be improved.
[0044] Figure 4 This is a schematic diagram illustrating example 400 of generating potential noise in a transition frame and feeding the potential noise into a denoising U-Net model according to some embodiments of this disclosure. Figure 4 As shown, Example 400 includes a potential noise generation module 410 (e.g., Figure 3 The potential noise generation module 306) and the denoising U-Net module 418 (e.g., Figure 3 The denoising module 308 in the model is used. The denoising U-Net model 418 is a neural network designed for image denoising tasks, and it is based on the U-Net architecture. The U-Net model is a convolutional neural network (CNN) architecture widely used in image processing tasks. The U-Net architecture is known for its ability to produce high-quality results with relatively little training data and is particularly effective when the task requires accurate localization and spatial information.
[0045] In example 400, the image embedding 402 is performed by an image encoder (e.g., Figure 3 The image encoder 302 in the image is generated based on the starting frame. Furthermore, the image embedding 404 is generated by the image encoder (e.g., ...). Figure 3 The image encoder 304 in the example generates noise based on the end frame. In example 400, the latent noise generation module 410 may include a denoising inversion module 406 and a denoising inversion module 408. The denoising inversion modules 406 and 408 can invert the denoising process used to generate image embeddings 402 and 404 to obtain latent noise 412 corresponding to image embedding 402 and latent noise 414 corresponding to image embedding 404.
[0046] In some implementations, the denoising inversion modules 406 and 408 can be denoising diffusion implicit model (DDIM) inversion modules. DDIM sampling is the process of generating an image from a trained DDIM by reversing the diffusion process. In the diffusion model, an image can be generated by starting with random noise and progressively refining it into a coherent output image through a series of denoising steps. DDIM inversion is the reverse process of DDIM sampling, where the goal is to acquire an existing image and reverse the generation process back to its corresponding latent noise. DDIM inversion allows the generation of an image from latent noise z using the following equation (1). t Obtain the latent noise z t+1 :
[0047]
[0048] Where α t Variance scheduling from the forward diffusion process, and Indicated by image conditions Denoising U-Net with text condition c as the condition.
[0049] like Figure 4 As shown, after generating latent noise 412 corresponding to the start frame and latent noise 414 corresponding to the end frame, the latent noise generation module 410 can generate latent noise 416-1, 416-2, ..., 416-N (also collectively referred to as latent noise 416) for the transition frame to be generated based on the latent noise 412 and 414. Then, the generated latent noise 416 can be concatenated with the latent noise 412 and 414 respectively. For example, as... Figure 4 As shown, the latent noise 416-1 for one of the transition frames to be generated can be concatenated with latent noises 412 and 414. Then, the concatenated latent noise can be fed into the denoising U-Net model 418 to generate the transition frame.
[0050] In some implementations, the potential noise for the transition frame can be generated by interpolating the potential noise corresponding to the start frame and the potential noise corresponding to the end frame. Figure 5 This is a schematic diagram illustrating an example process 500 for generating potential noise in a transition frame according to some embodiments of the present disclosure. Figure 5 As shown, in process 500, the computing device can generate potential noise for the transition frame based on the start frame 502 and the end frame 508. For example, the transition frame may include transition frame 504 and transition frame 506, wherein transition frame 504 is closer to the start frame 502 than transition frame 506.
[0051] In process 500, the computing device can generate latent noise 512 based on the start frame 502 and latent noise 518 based on the end frame 508. The computing device can generate latent noise 514 for the transition frame 504 and latent noise 516 for the transition frame 506. The computing device can perform interpolation on the latent noise 512 and latent noise 518 to generate latent noise 514 and latent noise 516. During the interpolation process, the contributions of latent noise 512 and 518 can be different. For example, in generating latent noise 514, the computing device can determine the weight 522 of latent noise 512 and the weight 524 of latent noise 518. Furthermore, in generating latent noise 516, the computing device can determine the weight 526 of latent noise 512 and the weight 528 of latent noise 518. Because the transition frame 504 is closer to the start frame 502 than the transition frame 506, the weight 522 can be greater than the weight 526.
[0052] In this way, the appearance of the object in transition frame 504 is more similar to the appearance of the object in the starting frame 502 than the appearance of the object in transition frame 506. Therefore, the smoothness of the generated transition frames can be improved, and the randomness and discontinuity of the generated transition frames can be reduced.
[0053] In some implementations, latent noise for transition frames can be generated by performing spherical interpolation on the latent noise corresponding to the start frame and the latent noise corresponding to the end frame. Spherical interpolation is a method of interpolation between two points on a sphere. Unlike linear interpolation, which operates in a straight line between two points in Euclidean space, spherical interpolation operates along the shortest path on the surface of a sphere. This ensures that the interpolation respects the spherical geometry of the data. By interpolating on a sphere, the transition between latent noises can be smoother and more realistic intermediate outputs can be produced. Spherical interpolation can be formulated as equation (2):
[0054]
[0055] Where z t1 z represents the potential noise corresponding to the starting frame. tN z represents the potential noise corresponding to the end frame. tn This represents the potential noise corresponding to the nth frame. And λ inject ∈[0,1] represents the parameters used for potential interpolation.
[0056] In this way, spherical interpolation can preserve the Euclidean norm of the interpolated potential noise, thereby improving the quality, consistency and realism of the generated transition frames.
[0057] In some implementations, the computing device may generate a first low-rank adaptation parameter based on a start frame and its description, and a second low-rank adaptation parameter based on an end frame and its description. Then, the computing device may generate a third low-rank adaptation parameter based on the first and second low-rank adaptation parameters. The computing device may generate a transition frame based on third latent noise and the third low-rank adaptation parameter. In some implementations, the computing device may generate the third low-rank adaptation parameter by performing linear interpolation on the first and second low-rank adaptation parameters. In some implementations, the computing device may generate a target denoising module by integrating the third low-rank adaptation parameter into the original denoising module. The computing device may generate a transition frame based on the third latent noise using the target denoising module.
[0058] Figure 6 This is a schematic diagram illustrating example 600 of integrating low-rank adaptation parameters of transition frames into a denoising U-Net model according to some embodiments of the present disclosure. Figure 6As shown, Example 600 includes a low-rank adapter module 610 (e.g., Figure 3 The low-rank adaptation module 610 can be configured to encapsulate high-level semantics into a low-rank parameter space. In example 600, a start frame 602, a description 604 of the start frame 602, an end frame 606, and a description 608 of the end frame 606 can be provided. The low-rank adaptation module 610 can be trained based on the start frame 602 and the description 604 to determine the low-rank adaptation parameters 612. In addition, the low-rank adaptation module 610 can be trained based on the end frame 606 and the description 608 to obtain the low-rank adaptation parameters 614. The objective function of the training process can be equation (3):
[0059]
[0060] Where Δθ n z represents the low-rank adaptation parameters for frame n. 0n Let c represent the encoded latent vector of frame n, and c n This indicates a text embedding associated with the transition description.
[0061] In Example 600, low-rank adaptation parameters 616-1, 616-2, ..., 616-N (also collectively referred to as low-rank adaptation parameters 616) of the transition frame can be generated based on low-rank adaptation parameters 612 and 614. In some implementations, low-rank adaptation parameters 616 can be generated by performing linear interpolation on low-rank adaptation parameters 612 and 614. If the transition frame to be generated is closer to the starting frame 602 than another transition frame, the contribution of low-rank adaptation parameters 612 to that transition frame can be greater than the contribution of low-rank adaptation parameters 612 to another transition frame. Linear interpolation can be formulated as equation (4):
[0062] Δθ=(1-λ adapt )Δθ1+λ adapt Δθ N (4)
[0063] Where Δθ represents the low-rank adaptation parameter for the transition frame, and Δθ1 represents the low-rank adaptation parameter for the start frame. N Let λ represent the low-rank adaptation parameters for the ending frame, and λ... adapt This represents the interpolation parameters during frame adaptation. As the transition frame gets closer to the starting frame, the interpolation parameters decrease.
[0064] like Figure 6As shown, the low-rank adaptation parameter 616 can be integrated into the denoised U-Net model 618. For example, the low-rank adaptation parameter 616-1 can be integrated into the denoised U-Net model 618 to obtain the denoised U-Net model 620. The denoised U-Net model 620 can be used to generate transition frames corresponding to the low-rank adaptation parameter 616-1.
[0065] In this way, by using a denoising U-Net model with integrated low-rank adaptation parameters as a noise prediction network, the generated transition frames can become semantically meaningful while maintaining temporal coherence.
[0066] In some implementations, the computing device may generate a first text embedding based on a description of a start frame and a second text embedding based on a description of an end frame. The computing device may then generate a third text embedding based on the first and second text embeddings. The computing device may then generate a transition frame based on the third text embedding and third latent noise. In some implementations, the computing device may generate the third text embedding by performing linear interpolation on the first and second text embeddings.
[0067] Figure 7 This is a schematic diagram illustrating example 700 of integrating text embedding for transition frames into a denoising U-Net model according to some embodiments of the present disclosure. Figure 7 As shown, Example 700 includes a text embedding generation module 710 (e.g., Figure 3 The text embedding generation module 314 can provide a start frame 702, a description 704 of the start frame 702, an end frame 706, and a description 708 of the end frame 706. The text embedding 712 can be generated by a text encoder (e.g., text encoder 312) based on the description 704, and the text embedding 714 can be generated by the text encoder based on the description 708.
[0068] In Example 700, text embeddings 716-1, 716-2, ..., 716-N (collectively referred to as text embedding 716) for a transition frame can be generated based on text embedding 712 and text embedding 714. In some implementations, text embedding 716 can be generated by performing linear interpolation on text embedding 712 and text embedding 714. If the transition frame to be generated is closer to the starting frame 702 than another transition frame, the contribution of text embedding 712 to that transition frame can be greater than the contribution of text embedding 712 to the other transition frame. Linear interpolation can be formulated as equation (5):
[0069] C λtext =(1-λ) text )c1+λ text c N (5)
[0070] Where c λtext c1 represents the text embedding for the transition frame, and c2 represents the text embedding for the start frame. N This represents the text embedding for the ending frame, and λ text ∈[0,1] is used as a frame-aware coefficient to control the transition sequence. The frame-aware coefficient decreases as the transition frame gets closer to the starting frame.
[0071] like Figure 7 As shown, text embedding 716 can be integrated into the cross-attention layer within the denoising U-Net model 718. For example, when generating a transition frame corresponding to text embedding 716-1, text embedding 716-1 can be integrated into the cross-attention layer within the denoising U-Net model 718.
[0072] In this way, meaningful transition frames can be generated by utilizing the interpolated text embedding used for DDIM sampling. For example, the interpolation between "lion" and "truck" can show a gradual transition, thus producing a truck with the shape and skin of a lion in the transition frame.
[0073] By leveraging latent noise interpolation, low-rank adaptation, and text embedding interpolation, the framework provided in this disclosure can be used for multiple transition generation tasks, including object morphing, concept mixing, motion prediction, and scene transitions. Furthermore, this framework is a zero-shot, unified, and plug-and-play solution that effectively generates semantically relevant, high-fidelity, and temporally coherent video transitions. Object morphing refers to input frames that depict the same or different objects in different poses, provided they are topologically similar. Concept mixing refers to input frames containing conceptually different objects (e.g., "airplane" and "cruise ship"). Motion prediction refers to input frames representing two moments in a video with one or more moving objects. Scene transitions refer to input frames of conceptually related scenes, but either belonging to different domains (e.g., "wooden cabin in the forest" and "wooden cabin in the snow") or representing different components of the scene (e.g., "erupting volcano" and "hot lava"). Figures 8A-8D This is a schematic diagram illustrating examples of multiple conversion tasks according to some embodiments of the present disclosure. In particular, Figure 8A Example 800 of an object deformation task is shown. Figure 8B Example 810 of a motion prediction task is shown. Figure 8C Example 820 of a conceptual hybrid task is shown. Figure 8D Example 830 of a scene transition task is shown.
[0074] Figure 9 This is a block diagram illustrating the physical components (e.g., hardware) of an electronic device 900 that can implement various aspects of this disclosure. For example, the electronic device 900 may be...Figure 1 The computing device 102 and the electronic device 900 can achieve the following: Figures 1-7 and Figures 8A-8D The process is illustrated. In a basic configuration, electronic device 900 may include at least one processing unit 902 and system memory 904. Depending on the configuration and type of computing device, system memory 904 may include, but is not limited to, volatile memory (e.g., random access memory), non-volatile memory (e.g., read-only memory), flash memory, or any combination of such memory.
[0075] System memory 904 may include operating system 905 and one or more program modules 906 adapted to perform the various aspects disclosed herein. For example, operating system 905 may be adapted to control the operation of electronic device 900. Furthermore, aspects of this disclosure may be practiced in conjunction with other operating systems or any other application, and are not limited to any particular application or system. This basic configuration is in Figure 9 The components within the dashed line 908 are shown. Electronic device 900 may have additional features or functions. For example, electronic device 900 may also include additional data storage devices (removable and / or non-removable), such as, for example, a disk, optical disk, or magnetic tape. Such additional storage... Figure 9 The image is shown by a removable storage device 909 and a non-removable storage device 910.
[0076] As described above, several program modules and data files can be stored in system memory 904. When executed on at least one processing unit 902, application 920 or program module 906 can perform processes including, but not limited to, one or more aspects as described herein. Application 920 may include application interface 921, which can communicate with, as previously discussed... Figures 1-7 and Figures 8A-8D The application interface 921 described in more detail is the same as or similar to that described herein. Other program modules that may be used according to various aspects of this disclosure may include email and contact applications, word processing applications, spreadsheet applications, database applications, PowerPoint presentation applications, drawing or computer-aided applications, and / or one or more components supported by the system described herein.
[0077] Furthermore, aspects of this disclosure can be practiced in discrete electronic components, packaged or integrated electronic chips containing logic gates, circuits utilizing microprocessors, or circuits on a single chip containing electronic components or microprocessors. For example, aspects of this disclosure can be practiced via a system-on-a-chip (SOC), wherein... Figure 9Each or many of the components shown herein can be integrated onto a single integrated circuit. Such a SOC device may include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all integrated (or “burned in”) onto a chip substrate as a single integrated circuit. When operating via the SOC, the functionality described herein regarding the client switching protocol can be operated via dedicated logic integrated on a single integrated circuit (chip) along with other components of the processing device 500. Aspects of this disclosure can also be practiced using other techniques capable of performing logical operations, such as, for example, AND, OR, and NOT, including but not limited to mechanical, optical, fluid, and quantum technologies. Furthermore, aspects of this disclosure can be practiced within a general-purpose computer or any other circuit or system.
[0078] Electronic device 900 may also have one or more input devices 912, such as a keyboard, mouse, pen, voice or speech input device, touch or swipe input device, etc. It may also include output devices 914, such as a monitor, speaker, printer, etc. The above devices are examples, and other devices may be used. Processing device 500 may include one or more communication connections allowing communication with other computing or processing devices 950. Examples of suitable communication connections include, but are not limited to, radio frequency (RF) transmitters, receivers, and / or transceiver circuitry; universal serial bus (USB), parallel and / or serial ports.
[0079] As used herein, the term computer-readable medium can include computer storage media. Computer storage media can include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, or program modules. System memory 904, removable storage device 909, and non-removable storage device 910 are examples of computer storage media (e.g., memory storage). Computer storage media can include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassette, magnetic tape, disk storage or other magnetic storage devices, or any other article of manufacture that can be used to store information and can be accessed by electronic device 900. Any such computer storage medium may be part of electronic device 900. Computer storage media does not include carrier waves or other propagated or modulated data signals.
[0080] Communication media can be embodied in computer-readable instructions, data structures, program modules, or other data in modulated data signals (such as carrier waves or other transmission mechanisms), and include any information delivery medium. The term "modulated data signal" can describe a signal whose one or more characteristics are set or altered in a manner that encodes information in the signal. By way of example and not limitation, communication media can include wired media, such as wired networks or direct wired connections, and wireless media, such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0081] Furthermore, the aspects and functions described herein can operate on distributed systems (e.g., cloud-based computing systems), where application functions, memory, data storage and retrieval, and various processing functions can operate remotely to each other via distributed computing networks (such as the Internet or intranets). User interfaces and various types of information can be displayed via onboard computing device displays or via remote display units associated with one or more computing devices. For example, various types of user interfaces and information can be displayed and interacted with. Interaction with numerous computing systems where embodiments of the invention can be practiced includes key input, touchscreen input, voice or other audio input, gesture input, wherein the associated computing device is equipped with detection (e.g., camera) functions for capturing and interpreting user gestures used to control the functions of the computing device, etc.
[0082] The phrases “at least one,” “one or more,” “or,” and “and / or” are open expressions that are both conjunction and disjunction in operation. For example, each of the expressions “at least one of A, B, and C,” “at least one of A, B, or C,” “one or more of A, B, and C,” “one or more of A, B, or C,” “A, B, and / or C,” and “A, B, or C” means A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.
[0083] The term "a" or "an" for an entity refers to one or more of that entity. Therefore, the terms "a" (or "an"), "one or more," and "at least one" are used interchangeably herein. It should also be noted that the terms "comprising," "including," and "having" are used interchangeably.
[0084] As used herein, the term "automatic" and its variations refer to any process or operation that, when performed, is typically continuous or semi-continuous without substantial human input. However, a process or operation can be automatic even if its execution uses substantial or non-substantial human input, provided that input is received prior to its execution. Human input is considered substantial if it influences how the process or operation will be performed. Human input that consents to the execution of a process or operation is not considered "substantial."
[0085] Any steps, functions, and operations discussed in this article can be performed continuously and automatically.
[0086] Exemplary systems and methods of this disclosure have been described with respect to computing devices. However, to avoid unnecessarily obscuring this disclosure, several known structures and devices have been omitted from the foregoing description. Such omissions should not be construed as limiting. Specific details have been set forth to provide an understanding of this disclosure. However, it should be understood that this disclosure can be practiced in a variety of ways beyond the specific details set forth herein.
[0087] Furthermore, while the exemplary aspects shown herein illustrate the juxtaposition of various system components, some components of the system may be located in a remote portion of a distributed network (such as a LAN and / or the Internet) or in a remote location within a dedicated system. Therefore, it should be understood that system components may be combined into one or more devices, such as servers, communication equipment, or deployed on specific nodes of a distributed network, such as analog and / or digital telecommunications networks, packet-switched networks, or circuit-switched networks. As can be understood from the foregoing description, and for computational efficiency reasons, system components may be deployed anywhere within the distributed network of components without affecting the operation of the system.
[0088] Furthermore, it should be understood that the various links connecting the elements can be wired or wireless links, or any combination thereof, or any other known or later-developed element capable of supplying and / or communicating data to the connected elements. These wired or wireless links can also be secure links and can be capable of communicating encrypted information. The transmission medium used as the link can be, for example, any suitable carrier for electrical signals, including coaxial cables, copper wires, and optical fibers, and can take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.
[0089] Although flowcharts have been discussed and illustrated with respect to a specific sequence of events, it should be understood that changes, additions, and omissions to this sequence may occur without substantially affecting the operation of the disclosed configurations and aspects.
[0090] Several variations and modifications of this disclosure may be used. Some features of this disclosure may be provided without providing others.
[0091] In another configuration, the systems and methods of this disclosure may be implemented using a combination of a dedicated computer, a programmable microprocessor or microcontroller and peripheral integrated circuit elements, an ASIC or other integrated circuit, a digital signal processor, hardwired electronic or logic circuitry (such as discrete component circuitry), a programmable logic device or gate array (such as a PLD, PLA, FPGA, PAL, dedicated computer, any similar means, etc.). Generally, any device or means capable of implementing the methods shown herein can be used to implement various aspects of this disclosure. Exemplary hardware that may be used in this disclosure includes computers, handheld devices, telephones (e.g., cellular phones, internet-enabled phones, digital phones, analog phones, hybrid phones, etc.), and other hardware known in the art. Some of these devices include processors (e.g., single or multiple microprocessors), memory, non-volatile memory, input devices, and output devices. Furthermore, alternative software implementations may be constructed, including but not limited to distributed processing or component / object distributed processing, parallel processing, or virtual machine processing, to implement the methods described herein.
[0092] In another configuration, the disclosed method can be readily implemented using software from an object-oriented or object-oriented software development environment that provides portable source code usable on various computer or workstation platforms. Alternatively, the disclosed system can be implemented partially or entirely in hardware using standard logic circuitry or VLSI design. Whether software or hardware is used to implement the system according to this disclosure depends on the system's speed and / or efficiency requirements, the specific functions utilized, and the particular software or hardware system or microprocessor or microcomputer system.
[0093] In another configuration, the disclosed method can be implemented in part in software, which can be stored on a non-transient storage medium and executed on a programmed general-purpose computer in cooperation with a controller and memory, a dedicated computer, a microprocessor, etc. In these cases, the systems and methods of this disclosure can be implemented as programs embedded in a personal computer, such as applets, Alternatively, it can be a CGI script, a resource residing on a server or computer workstation, or a routine embedded in a dedicated measurement system, system component, etc. The system can also be implemented by physically integrating the system and / or method into the software and / or hardware system.
[0094] If described herein, this disclosure is not limited to standards and protocols. Other similar standards and protocols not mentioned herein exist and are included in this disclosure. Furthermore, the standards and protocols mentioned herein, as well as other similar standards and protocols not mentioned herein, are regularly superseded by faster or more efficient equivalents having substantially the same functionality. Such alternative standards and protocols having the same functionality are considered equivalents included in this disclosure.
[0095] This disclosure includes, in various configurations and aspects, components, methods, processes, systems, and / or apparatus as depicted and described herein, including various combinations, sub-combinations, and subsets thereof. Upon understanding this disclosure, those skilled in the art will understand how to manufacture and use the systems and methods disclosed herein. This disclosure includes, in various configurations and aspects, providing apparatus and processes in the absence of items not depicted and / or described herein, or in various configurations or aspects thereof, including the absence of such items that might have been used in prior apparatus or processes, for example, to improve performance, ease of implementation, and / or reduce implementation costs.
[0096] The description and illustration of one or more aspects provided in this application are not intended to limit or constrain the scope of the claimed disclosure in any way. The aspects, examples, and details provided in this application are considered sufficient to convey ownership and enable others to make and use the best mode of the claimed disclosure. The claimed disclosure should not be construed as limited to any aspect, example, or detail provided in this application. Various features (both structural and methodological) are intended to be selectively included or omitted, whether shown and described in combination or separately, to produce embodiments with a particular set of features. Having been provided with the description and illustration of this application, those skilled in the art can conceive of variations, modifications, and alternatives falling within the spirit of the broader aspects of the general inventive concept embodied in this application, without departing from the broader scope of the claimed disclosure.
Claims
1. A method for generating a video, comprising: obtaining a start frame and an end frame for the video, a first description of the start frame, and a second description of the end frame; generating a first latent noise in a latent space based on the start frame, and a second latent noise in the latent space based on the end frame; generating a third latent noise in the latent space corresponding to a transition frame based on the first latent noise and the second latent noise; and generating the transition frame based on the third latent noise, the start frame, the end frame, the first description, and the second description by utilizing a pre-trained image-to-video diffusion model.
2. The method of claim 1, wherein generating the first latent noise in a latent space based on the start frame, and the second latent noise in the latent space based on the end frame comprises: generating a first image embedding based on the start frame; generating a second image embedding based on the end frame; generating the first latent noise by inverting a denoising process used to generate the first image embedding; and generating the second latent noise by inverting a denoising process used to generate the second image embedding.
3. The method of claim 1, wherein generating the third latent noise in the latent space corresponding to the transition frame based on the first latent noise and the second latent noise comprises: generating the third latent noise by performing an interpolation on the first latent noise and the second latent noise.
4. The method of claim 3, wherein generating the third latent noise by performing the interpolation on the first latent noise and the second latent noise comprises: generating the third latent noise by performing a spherical linear interpolation on the first latent noise and the second latent noise.
5. The method of claim 1, wherein generating the transition frame based on the third latent noise, the start frame, the end frame, the first description, and the second description by utilizing the pre-trained image-to-video diffusion model comprises: generating a first low-rank adaptation parameter based on the start frame and the first description; generating a second low-rank adaptation parameter based on the end frame and the second description; generating a third low-rank adaptation parameter based on the first low-rank adaptation parameter and the second low-rank adaptation parameter; and generating the transition frame based on the third latent noise and the third low-rank adaptation parameter.
6. The method of claim 5, wherein generating the third low-rank adaptation parameter based on the first low-rank adaptation parameter and the second low-rank adaptation parameter comprises: generating the third low-rank adaptation parameter by performing a linear interpolation on the first low-rank adaptation parameter and the second low-rank adaptation parameter.
7. The method of claim 5, wherein the pre-trained image-to-video diffusion model comprises an original denoising module, and generating the transition frame based on the third latent noise and the third low-rank adaptation parameter comprises: generating a target denoising module by integrating the third low-rank adaptation parameter into the original denoising module; and generating the transition frame based on the third latent noise by utilizing the target denoising module.
8. The method of claim 1, wherein generating the transition frame based on the third latent noise, the starting frame, the ending frame, the first description, and the second description by utilizing the pre-trained image-to-video diffusion model comprises: generating a first text embedding based on the first description of the starting frame; generating a second text embedding based on the second description of the ending frame; generating a third text embedding based on the first text embedding and the second text embedding; and generating the transition frame based on the third text embedding and the third latent noise.
9. The method of claim 8, wherein generating the third text embedding based on the first text embedding and the second text embedding comprises: generating the third text embedding by performing linear interpolation on the first text embedding and the second text embedding.
10. The method of claim 1, wherein the starting frame, the first description, the ending frame, and the second description are applied to any one of the following transformation tasks: object morphing, concept blending, motion prediction, and scene transition.
11. An electronic device, comprising: a memory and a processor; wherein the memory is configured to store one or more computer instructions that, when executed by the processor, cause the processor to: obtain a starting frame and an ending frame for the video, a first description of the starting frame, and a second description of the ending frame; generate a first latent noise in a latent space based on the starting frame, and a second latent noise in the latent space based on the ending frame; generate a third latent noise in the latent space corresponding to a transition frame based on the first latent noise and the second latent noise; and generate the transition frame based on the third latent noise, the starting frame, the ending frame, the first description, and the second description by utilizing a pre-trained image-to-video diffusion model.
12. The apparatus of claim 11, wherein the instructions that cause the processor to generate the first latent noise in a latent space based on the starting frame, and the second latent noise in the latent space based on the ending frame comprise instructions that cause the processor to: generate a first image embedding based on the starting frame; generate a second image embedding based on the ending frame; generate the first latent noise by reversing a denoising process used to generate the first image embedding; and generate the second latent noise by reversing a denoising process used to generate the second image embedding.
13. The apparatus of claim 11, wherein the instructions to cause the processor to generate, based on the first latent noise and the second latent noise, the third latent noise in the latent space corresponding to the transition frame comprise instructions to cause the processor to perform the following operations: generate the third latent noise by performing an interpolation on the first latent noise and the second latent noise.
14. The apparatus of claim 13, wherein the instructions to cause the processor to generate the third latent noise by performing the interpolation on the first latent noise and the second latent noise comprise instructions to cause the processor to perform the following operations: generate the third latent noise by performing a spherical linear interpolation on the first latent noise and the second latent noise.
15. The apparatus of claim 11, wherein the instructions to cause the processor to generate, based on the third latent noise, the start frame, the end frame, the first description, and the second description, the transition frame by utilizing the pre-trained image-to-video diffusion model comprise instructions to cause the processor to perform the following operations: generate a first low-rank adaptation parameter based on the start frame and the first description; generate a second low-rank adaptation parameter based on the end frame and the second description; generate a third low-rank adaptation parameter based on the first low-rank adaptation parameter and the second low-rank adaptation parameter; and generate the transition frame based on the third latent noise and the third low-rank adaptation parameter.
16. The apparatus of claim 15, wherein the instructions to cause the processor to generate the third low-rank adaptation parameter based on the first low-rank adaptation parameter and the second low-rank adaptation parameter comprise instructions to cause the processor to perform the following operations: generate the third low-rank adaptation parameter by performing a linear interpolation on the first low-rank adaptation parameter and the second low-rank adaptation parameter.
17. The apparatus of claim 15, wherein the pre-trained image-to-video diffusion model comprises a raw denoising module, and the instructions to cause the processor to generate the transition frame based on the third latent noise and the third low-rank adaptation parameter comprise instructions to cause the processor to perform the following operations: generate a target denoising module by integrating the third low-rank adaptation parameter into the raw denoising module; and generate the transition frame based on the third latent noise by utilizing the target denoising module.
18. The apparatus of claim 11, wherein the instructions to cause the processor to generate, based on the third latent noise, the start frame, the end frame, the first description, and the second description, the transition frame by utilizing a pre-trained image-to-video diffusion model comprise instructions to cause the processor to perform the following operations: generate a first text embedding based on the first description of the start frame; generate a second text embedding based on the second description of the end frame; generate a third text embedding based on the first text embedding and the second text embedding; and generate the transition frame based on the third latent noise and the third text embedding. generating the transition frame based on the third text embedding and the third latent noise.
19. The apparatus of claim 18, wherein the instructions to cause the processor to generate the third text embedding based on the first text embedding and the second text embedding comprise instructions to cause the processor to: generate the third text embedding by performing linear interpolation on the first text embedding and the second text embedding.
20. A non-transitory computer-readable medium comprising instructions stored thereon that, when executed by a processor, cause the processor to: obtain a starting frame and an ending frame for the video, a first description for the starting frame, and a second description for the ending frame; generate a first latent noise in a latent space based on the starting frame and a second latent noise in the latent space based on the ending frame; generate a third latent noise in the latent space corresponding to a transition frame based on the first latent noise and the second latent noise; and generate the transition frame based on the third latent noise, the starting frame, the ending frame, the first description, and the second description by utilizing a pre-trained image-to-video diffusion model.