Video interpolation with diffusion models

Cascaded diffusion models address the challenge of generating high-quality, high-resolution video interpolations by first producing low-resolution outputs and then conditioning on original high-resolution frames to achieve smooth and detailed transitions between start and end frames.

WO2025117737A1PCT designated stage expired Publication Date: 2025-06-05GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/057744
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-11-27
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current video interpolation techniques struggle to generate high-quality, high-resolution interpolated frames, especially when the start and end frames are distinct, and the motion is complex, nonlinear, or ambiguous.

Method used

The use of cascaded diffusion models for video interpolation, where a first denoising diffusion model generates low-resolution synthetic interpolated images, and a second model, conditioned on the original high-resolution frames and the low-resolution outputs, produces high-resolution interpolated images.

Benefits of technology

This approach effectively handles complex motion scenarios and achieves high-fidelity results by generating high-quality, high-resolution interpolated videos that smoothly transition between the start and end frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024057744_05062025_PF_FP_ABST
    Figure US2024057744_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are generative models for video interpolation. The proposed models are able to create short videos given a start and end frame. In order to achieve high fidelity and generate motions unseen in the input data, the present disclosure employs the use of cascaded diffusion models to first generate the target video at low resolution, and then generate the high-resolution video conditioned on the low-resolution generated video.
Need to check novelty before this filing date? Find Prior Art

Description

VIDEO INTERPOLATION WITH DIFFUSION MODELSRELATED APPLICATIONS

[0001] This application claims priority to and the benefit of United States Provisional Patent Application Number 63 / 603,439, filed November 28, 2023. United States Provisional Patent Application Number 63 / 603,439 is hereby incorporated by reference in its entirety.FIELD

[0002] The present disclosure relates to the field of video processing, and more specifically, to video interpolation methods using diffusion models.BACKGROUND

[0003] Video interpolation is a fundamental technology in numerous applications, including video editing, slow-motion video generation, frame-rate up-sampling, and others. The task of video interpolation typically involves generating intermediate frames between two given frames of a video sequence. For instance, it can be used to transform a video captured at 30 frames per second (fps) into a video sequence playing at 60 fps, thereby enhancing the smoothness and quality of the video playback.

[0004] Despite the extensive research and numerous methods proposed in the art, the current state-of-the-art video interpolation techniques face considerable challenges, especially when the start and end frames are distinct. Existing methods that are based on linear or unambiguous motion assumptions fail to generate plausible interpolations in such scenarios. In addition to these limitations, there is a significant technical challenge relating to high- resolution video generation. Existing approaches which employ generative models struggle to achieve satisfactory sample quality when generating high-resolution videos.SUMMARY

[0005] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or can be learned from the description, or can be learned through practice of the embodiments.

[0006] One example aspect of the present disclosure is directed to computing system configured to perform video interpolation via cascaded diffusion models, the computing system comprising: one or more processors; and one or more non-transitory computer- readable media that collectively store computer-executable instructions for performingoperations. The operations comprising: obtaining a pair of images comprising input versions of a start image and an end image that depict a scene, wherein the input versions of the start image and the end image have a input resolution; downsampling the input versions of the start image and the end image to generate downsampled versions of the start image and the end image that have a reduced resolution, the reduced resolution being smaller than the input resolution; processing a first noisy input with a first denoising diffusion model that is conditioned on the downsampled versions of the start image and the end image to generate, as an output of the first denoising diffusion model, one or more first synthetic interpolated images that depict content of the scene temporally between the start image and the end image, wherein the one or more first synthetic interpolated images have the reduced resolution; and processing a second noisy input with a second denoising diffusion model that is conditioned on the input versions of the start image and the end image and the one or more first synthetic interpolated images to generate, as an output of the second denoising diffusion model, one or more second synthetic interpolated images that depict content of the scene temporally between the start image and the end image, wherein the one or more second synthetic interpolated images have the input resolution.

[0007] Another example aspect of the present disclosure is directed to computer- implemented method for training diffusion models to perform video interpolation, the method comprising: obtaining, by a computing system comprising one or more computing devices, a training tuple comprising input versions of a start image and an end image that depict a scene and one or more ground truth images that depict content of the scene temporally between the start image and the end image, wherein the input versions of the start image and the end image have a input resolution; downsampling, by the computing system, the input versions of the start image and the end image to generate downsampled versions of the start image and the end image that have a reduced resolution, the reduced resolution being smaller than the input resolution; processing, by the computing system, a first noisy input with a first denoising diffusion model that is conditioned on the downsampled versions of the start image and the end image to generate, as an output of the first denoising diffusion model, one or more first synthetic interpolated images that depict content of the scene temporally between the start image and the end image, wherein the one or more first synthetic interpolated images have the reduced resolution; evaluating, by the computing system, a first loss function that compares: first added noise added to the ground truth images to generate the first noisy input, and first predicted noise removed from the first noisy input to generate the one or more first synthetic interpolated images; and modifying, by the computing system, one or more firstvalues of one or more first parameters of the first denoising diffusion model based on the first loss function.

[0008] Other aspects of the present disclosure are directed to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.

[0009] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the related principles.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Detailed discussion of embodiments directed to one of ordinary skill in the art is set forth in the specification, which makes reference to the appended figures, in which:

[0011] Figure 1 depicts a graphical diagram of an example inference scheme to perform video interpolation according to example embodiments of the present disclosure.

[0012] Figure 2 depicts a graphical diagram of an example training scheme to learn to perform video interpolation according to example embodiments of the present disclosure.

[0013] Figure 3A depicts a block diagram of an example computing system according to example embodiments of the present disclosure.

[0014] Figure 3B depicts a block diagram of an example computing device according to example embodiments of the present disclosure.

[0015] Figure 3C depicts a block diagram of an example computing device according to example embodiments of the present disclosure.

[0016] Reference numerals that are repeated across plural figures are intended to identify the same features in various implementations.DETAILED DESCRIPTION

[0017] Overview

[0018] The present disclosure provides generative models for video interpolation. The proposed models are able to create short videos given a start and end frame. In order to achieve high fidelity and generate motions unseen in the input data, the present disclosure employs the use of cascaded diffusion models to first generate the target video at low resolution, and then generate the high-resolution video conditioned on the low-resolution generated video. While prior work fail in most settings where the underlying motion iscomplex, nonlinear, or ambiguous, the proposed models can easily handle such cases. Furthermore, some example implementations of the present disclosure perform classifier-free guidance on the start and end frame and / or condition the super-resolution model on the original high-resolution frames without additional parameters to achieve high-fidelity results. The proposed models are fast to sample from as they jointly denoise all the frames to be generated, require less than a billion parameters per diffusion model to produce compelling results, and still enjoy scalability and improved quality at larger parameter counts.

[0019] More particularly, one example aspect is directed to computing systems and methods for performing video interpolation with cascaded diffusion models. As an example, a computing system can first obtain a pair of images, which serve as the start and end frames of the video. These images depict a scene and come in a specific input resolution. For instance, the start and end frames could be frames from a video of a sporting event or a nature scene, captured at a resolution of 256x256 pixels.

[0020] Once the input frames have been obtained, the system can downsample these images to create versions with a reduced resolution. This reduced resolution is smaller than the input resolution. For example, the system might downsample the input frames from a resolution of 256x256 pixels to a reduced resolution of 64x64 pixels.

[0021] The system then processes a noisy input with a first denoising diffusion model, which can also be referred to a base model. This first denoising diffusion model is conditioned on the downsampled versions of the start and end frames. The output of this model is one or more first synthetic interpolated images that depict content of the scene between the start and end frames. These first synthetic interpolated images have the reduced resolution. For instance, if the start and end frames respectively show a racecar approaching a corner and exiting a corner, then the first synthetic interpolated images might show the racecar in various stages of passing through the corner.

[0022] Next, the system processes a second noisy input with a second denoising diffusion model, which can also be referred to as a super-resolution model. This second denoising diffusion model is conditioned on the original, high-resolution versions of the start and end frames, as well as the first synthetic interpolated images generated by the first model. The output of this second denoising diffusion model is one or more second synthetic interpolated images that also depict the scene between the start and end frames, but at the original, high-resolution. These high-resolution interpolated images provide a detailed and smooth transition between the start and end frames.

[0023] Another aspect of the present disclosure is directed to techniques for a method for conditioning the diffusion models on the input frames. In one or both of the first and second denoising diffusion models, the start and end frames can be concatenated with the noisy input along the channel axis before being input to the model. This process allows information from the input frames to propagate through the network, aiding in the generation of the interpolated images, but without requiring additional parameters, thereby conserving computational requirements.

[0024] Another aspect of the present disclosure is directed to the use of classifier-free guidance (CFG) on the conditioning frames. CFG can assist in achieving the best quality in the generated images. For instance, CFG can help ensure that the interpolated images accurately represent the motion of the racecar in the example above, leading to a more realistic and high-quality video.

[0025] Another aspect of the present disclosure is directed to the use of shared convolution and self-attention blocks in the diffusion models. These blocks are shared over frames, and feature maps are only allowed to mix over frames via temporal attention blocks. This architecture can help to reduce the complexity of the models and improve their efficiency. For example, by sharing these blocks over frames, the system can reduce the number of parameters it needs to learn, speeding up the training process and improving the quality of the generated images.

[0026] Another aspect of the present disclosure is directed to systems and methods for training diffusion models for video interpolation. The two models described above can be trained using a training tuple that includes input versions of a start image and an end image, along with one or more ground truth images. The input images depict a scene and the ground truth images depict the content of the scene temporally between the start and end images.

[0027] The input versions of the start and end images are first downscaled by the computing system to a reduced resolution. The downscaled versions of the start and end images are then used to condition the first denoising diffusion model, also referred to as the base model, as described above.

[0028] The base model processes a first noisy input to generate one or more synthetic interpolated images. The first noisy input is generated by adding noise to the ground truth images. The base model then removes this noise to generate the synthetic interpolated images. These synthetic images depict the content of the scene temporally between the start and end images and have the same reduced resolution as the downscaled start and end images.

[0029] The present disclosure then evaluates a first loss function that compares the first added noise and the first predicted noise. The result of this evaluation is used to modify the parameters of the base model.

[0030] The present disclosure also includes a second denoising diffusion model, which can also be referred to as the super-resolution model. This model is conditioned on the original high-resolution versions of the start and end images, as well as the synthetic interpolated images generated by the base model. The super-resolution model processes a second noisy input to generate one or more high-resolution synthetic interpolated images. The second noisy input is generated by adding noise to the ground truth images.

[0031] The present disclosure then evaluates a second loss function that compares the second added noise and the second predicted noise removed from the second noisy input to generate the high-resolution synthetic interpolated images. The result of this evaluation is used to modify the parameters of the super-resolution model.

[0032] The proposed techniques for video interpolation using diffusion models provide a number of technical effects and benefits. The problem addressed by the present disclosure is a technical one: the generation of high-quality, high-resolution interpolated video frames from two distinct start and end frames. This is a significant challenge in the field of video processing, and the current state-of-the-art methods fail to produce satisfactory results, especially in scenarios where the start and end frames are highly distinct or where the underlying motion is complex, nonlinear, or ambiguous.

[0033] The proposed techniques provide a technical solution to this problem. By employing cascaded diffusion models, the techniques first generate a low-resolution target video, then create a high-resolution video conditioned on the low-resolution output. This stepwise approach allows the generation of high-quality videos that can handle complex and nonlinear motion scenarios. Furthermore, the techniques include additional innovations like the use of classifier-free guidance and the direct conditioning the super-resolution model on the original high-resolution frames without additional parameters, which can enhance the quality of the output.

[0034] The proposed techniques for video interpolation using diffusion models can find technical applications in various fields where video processing is beneficial. One significant application is in the domain of video capture systems, where the quality and clarity of recorded footage is paramount. The techniques can be used to enhance the frame rate of the recorded video, making the captured footage smoother and more detailed. For example, in the field of live sports broadcasting, these techniques can be employed to generate slow-motionreplays from real-time footage by creating intermediate frames, thereby enhancing the viewer experience. As another example, in cinematography and video editing, the proposed method can be used for generating high-resolution slow-motion scenes from standard video footage, contributing to the artistic and technical quality of the production. As yet another example, in the medical field, these techniques could be applied in endoscopy or microscopy videos to generate high-resolution, smooth video sequences from low-resolution, jerky inputs, aiding in better diagnosis and understanding.

[0035] With reference now to the Figures, example embodiments of the present disclosure will be discussed in further detail.

[0036] Example Inference Scheme

[0037] Figure 1 provides a schematic representation of an example method for video interpolation using cascaded diffusion models, in accordance with the present disclosure. The figure illustrates the sequence of operations that take place during the video interpolation process.

[0038] The process starts with the obtaining of a pair of images that are input versions of a start image (12) and an end image (14). These images depict a specific scene, for instance, a car moving along a track. The start and end images have an input resolution, which could, as one example, be a high resolution such as 256x256 pixels.

[0039] Following the acquisition of the start and end images, the system downsamples these images to create versions with a reduced resolution (16 and 18). This reduced resolution is smaller than the input resolution. As one example, the system might downsample the input frames from a resolution of 256x256 pixels to a reduced resolution of 64x64 pixels.

[0040] The next step involves processing a first noisy input (20) with a first denoising diffusion model (22). The first denoising diffusion model, also referred to as the base model, is conditioned on the downsampled versions of the start and end frames (16 and 18). The output of this model is one or more first synthetic interpolated images (24). These images show the scene's content between the start and end frames and have the reduced resolution.For instance, if the start and end frames respectively show a racecar approaching a corner and exiting a corner, then the first synthetic interpolated images might show the racecar in various stages of passing through the corner.

[0041] Following the generation of the first synthetic interpolated images (24), the system processes a second noisy input (26) with a second denoising diffusion model (28). This model, also referred to as the super-resolution model, is conditioned on both the original high-resolution versions of the start and end frames (12 and 14) and the first syntheticinterpolated images (24) generated by the base model (22). The second denoising diffusion model (28) generates one or more second synthetic interpolated images (30). These images also depict the scene between the start and end frames, but at the original high resolution.

[0042] Thus, Figure 1 illustrates an example embodiment of the present disclosure's video interpolation method. It presents a sequence of operations that transform a pair of start and end images into a series of interpolated images that smoothly transition from the start to the end image. The cascaded application of two diffusion models, each with its unique conditioning and resolution characteristics, allows for the generation of high-quality, high- resolution interpolated images.

[0043] Example Training Scheme

[0044] Figure 2 provides a schematic representation of an example method for training diffusion models for video interpolation, in accordance with the present disclosure. The figure illustrates the sequence of operations that take place during the training process.

[0045] In a first stage of the process, the present disclosure obtains a training tuple. This training tuple includes input versions of a start image (212) and an end image (214), along with one or more ground truth images (232). The start and end images depict a specific scene, and the ground truth images depict the content of the scene temporally between the start and end images. The start and end images have an input resolution, which could be a high resolution such as, for example, 256x256 pixels.

[0046] Following the acquisition of the training tuple, the system downsamples the input versions of the start and end images (212 and 214) to create versions with a reduced resolution (216 and 218). This reduced resolution is smaller than the input resolution. As one example, the system might downsample the input frames from a resolution of 256x256 pixels to a reduced resolution of 64x64 pixels.

[0047] The next step involves processing a first noisy input (220) with a first denoising diffusion model (222). The first denoising diffusion model (222), also referred to as the base model, is conditioned on the downsampled versions of the start and end frames (216 and 218). The output of this model is one or more first synthetic interpolated images (224). These images show the scene's content between the start and end frames and have the reduced resolution.

[0048] Following the generation of the first synthetic interpolated images (224), the system evaluates a first loss function (234). This loss function compares first added noise, which was added to the ground truth images (232) to generate the first noisy input (220), and first predicted noise, which was removed from the first noisy input (220) to generate the firstsynthetic interpolated images (224). The result of this evaluation is used to modify the parameters of the base model (222).

[0049] Figure 2 also illustrates a second denoising diffusion model (228), which can also be referred to as the super-resolution model. This model is conditioned on both the original high-resolution versions of the start and end frames (212 and 214) and (upsampled versions of) the first synthetic interpolated images (224) generated by the base model (222). The super-resolution model (228) processes a second noisy input (226) to generate one or more second synthetic interpolated images (230). These images also depict the scene between the start and end frames, but at the original high resolution.

[0050] Following the generation of the second synthetic interpolated images (230), the system evaluates a second loss function (236). This loss function compares second added noise, which was added to the ground truth images (232) to generate the second noisy input (226), and second predicted noise, which was removed from the second noisy input (226) to generate the second synthetic interpolated images (230). The result of this evaluation is used to modify the parameters of the super-resolution model (228).

[0051] Thus, Figure 2 illustrates an example method for training diffusion models for video interpolation. It presents a sequence of operations that transform a pair of start and end images and ground truth images into a series of interpolated images that smoothly transition from the start to the end image. The cascaded application of two diffusion models, each with its unique conditioning and resolution characteristics, allows for the generation of high- quality, high-resolution interpolated images. The evaluation of loss functions at each stage ensures that the models are trained to generate images that are as close as possible to the ground truth images, thereby enhancing the quality of the video interpolation.

[0052] The training process can be implemented on a range of computing systems, from personal computers to dedicated servers, depending on the scale of the task. The flexibility of the process means it can be adapted to different types of video content and different resolutions, making it applicable to a wide variety of scenarios.

[0053] Example Implementation Details

[0054] This section provides example details for example implementations of the systems and methods described herein. The systems and methods described herein are not limited to the example details contained in this section.

[0055] Prior work has shown that diffusion models do not achieve good sample quality for high-resolution generation with a single model without revising several hyperparameters and architecture details. In contrast, example implementations of the present disclosureleverage a cascaded model strategy. As one example, this can include training separate base and super-resolution models. While there is additional overhead to maintaining multiple models (e.g., diffusion models), this still avoids several complexities of latent diffusion models, such as finding an optimal encoder-decoder model, and having to address temporal inconsistency in the decoder with other training or fine-tuning procedures.

[0056] Some example implementations train two video diffusion models: first a base model can be trained. The base model can be conditioned on two 64x64 frames and generate seven 64x64 in-between frames. Next, a super-resolution model can be trained. The superresolution model can be conditioned on two 256x256 frames and seven 64x64 frames. The super-resolution model can generate the seven corresponding 256x256 frames. Some example implementations generate an odd number of frames to allow evaluating the middle frame. However, the number of frames is a hyperparameter that can be modified to meet various different objectives.

[0057] Example Model Architecture: In some implementations, a UNet architecture can be adapted by sharing all convolution and self-attention blocks over frames. In some implementations, feature maps are only allowed to mix over frames with the addition of temporal attention blocks where the query-key -value sequence lengths are the number of frames. In some implementations, simple positional encodings (e.g., differing over frames) for video timestamps normalized to [0,1] can be summed to the usual noise level embeddings (identical for all noisy frames). In some implementations, these embeddings can be propagated to each UNet block using FiLM. Some example implementations do the same for the super-resolution model to condition on the high-resolution start and end frames. In some implementations, the super-resolution model only differs from the base model in that (1) it concatenates each (e.g., naively upsampled) low-resolution conditioning frame to the noisy high-resolution frames along the channel axis, and (2) it downsamples before the first convolutional residual block to reduce memory usage. For more stable and efficient training, some example implementations additionally use attention blocks, which employ query-key normalization and an MLP block that runs in parallel to the attention block.

[0058] Example parameter-free frame conditioning: Some example implementations perform frame-conditioning without any additional parameter. As an example, in both the base and super-resolution models, some example implementations condition on the start and end frames simply by feeding these two additional frames to the entire UNet (e.g., concatenating along the frame axis). Because the feature maps for each frame additionally depends on the noise levels, some example implementations simply set fake noise levels forthe conditioning frames as the minimum noise level (maximum log-signal-to-noise-ratio). This adds two new sequence elements to the temporal attention layer, which lets information from the conditioning frames propagate to the rest of the network without any additional parameters. This contrasts other diffusion architectures, where the usual choice is additional cross-attention layers which doesn’t scale to condition on more frames and make the parameter count dependent on the number of frames. Other work has found that concatenating extra frames along the channel axis leads to worse sample quality when the generated and conditioning frames are not perfectly aligned.

[0059] Example guidance on conditioning frames: In some implementations, classifier- free guidance (CFG) on the conditioning frames improves ample quality. Similarly to the parameter-free frame conditioning strategy, some example implementations benefit from masked conditioning frames for CFG playing naturally with parameter sharing across frames. To achieve this, instead of zeroing-out the conditioning frames, some example implementations replace them with isotropic Gaussian noise and set their corresponding noise levels to the maximum value. It would also be confounding if the timestamps for the conditioning frames are set to zero, so some example implementations instead replace them with a learned null token.

[0060] Example diffusion modeling choices: This subsection now provides a brief overview of an example training objective formulation that can be used for example implementations of the present disclosure. The description begins with the simpler continuous-time objective for learning a data distribution p(x|c), where c are the start and end frames and x are the middle frames. Now define a forward process at every possible log signal-to-noise-ratio (a.k.a. “log-SNR”) in the usual manner via= ^ / sigmoid I) and O = / sigmoid(— ). One example training objective is then

[0061] Some example implementations use the LI loss as, in some settings, it helps produce better high-frequency details in samples compared to the standard L2 loss. Some example implementations use a cosine log-SNR schedule At, with maximum log-SNR of 20 at t = 0 and minimum log-SNR of -20 at t = 1.

[0062] Example Devices and Systems

[0063] Figure 3A depicts a block diagram of an example computing system 100 that performs video interpolation according to example embodiments of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled over a network 180.

[0064] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0065] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 which are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0066] In some implementations, the user computing device 102 can store or include one or more machine-learned models 120. For example, the machine-learned models 120 can be or can otherwise include various machine-learned models such as neural networks (e.g., deep neural networks) or other types of machine-learned models, including non-linear models and / or linear models. Neural networks can include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks or other forms of neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models). Example machine-learned models 120 are discussed with reference to Figures 1-2.

[0067] In some implementations, the one or more machine-learned models 120 can be received from the server computing system 130 over network 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 can implement multiple parallel instances of a single machine-learned model 120 (e.g., to perform parallel video interpolation across multiple instances of input images).

[0068] Additionally or alternatively, one or more machine-learned models 140 can be included in or otherwise stored and implemented by the server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine-learned models 140 can be implemented by the server computing system 140 as a portion of a web service (e.g., a video interpolation service). Thus, one or more models 120 can be stored and implemented at the user computing device 102 and / or one or more models 140 can be stored and implemented at the server computing system 130.

[0069] The user computing device 102 can also include one or more user input components 122 that receives user input. For example, the user input component 122 can be a touch-sensitive component (e.g., a touch-sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component can serve to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.

[0070] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 which are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0071] In some implementations, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. In instances in which the server computing system 130 includes plural server computing devices, such server computing devices can operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.

[0072] As described above, the server computing system 130 can store or otherwise include one or more machine-learned models 140. For example, the models 140 can be or can otherwise include various machine-learned models. Example machine-learned models include neural networks or other multi-layer non-linear models. Example neural networks include feed forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine-learned models can leverage anattention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models). Example models 140 are discussed with reference to Figures 1-2.

[0073] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 via interaction with the training computing system 150 that is communicatively coupled over the network 180. The training computing system 150 can be separate from the server computing system 130 or can be a portion of the server computing system 130.

[0074] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 which are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0075] The training computing system 150 can include a model trainer 160 that trains the machine-learned models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as, for example, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations.

[0076] In some implementations, performing backwards propagation of errors can include performing truncated backpropagation through time. The model trainer 160 can perform a number of generalization techniques (e.g., weight decays, dropouts, etc.) to improve the generalization capability of the models being trained.

[0077] In particular, the model trainer 160 can train the machine-learned models 120 and / or 140 based on a set of training data 162. The training data 162 can include, for example, a series of real images contained in a movie.

[0078] In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 on user-specific data received from the user computing device 102. In some instances, this process can be referred to as personalizing the model.

[0079] The model trainer 160 includes computer logic utilized to provide desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into a memory and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions that are stored in a tangible computer-readable storage medium such as RAM, hard disk, or optical or magnetic media.

[0080] The network 180 can be any type of communications network, such as a local area network (e.g., intranet), wide area network (e.g., Internet), or some combination thereof and can include any number of wired or wireless links. In general, communication over the network 180 can be carried via any type of wired and / or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0081] Figure 3 A illustrates one example computing system that can be used to implement the present disclosure. Other computing systems can be used as well. For example, in some implementations, the user computing device 102 can include the model trainer 160 and the training dataset 162. In such implementations, the models 120 can be both trained and used locally at the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the models 120 based on user-specific data.

[0082] Figure 3B depicts a block diagram of an example computing device 10 that performs according to example embodiments of the present disclosure. The computing device 10 can be a user computing device or a server computing device.

[0083] The computing device 10 includes a number of applications (e.g., applications 1 through N). Each application contains its own machine learning library and machine-learned model(s). For example, each application can include a machine-learned model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0084] As illustrated in Figure 3B, each application can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0085] Figure 3C depicts a block diagram of an example computing device 50 that performs according to example embodiments of the present disclosure. The computing device 50 can be a user computing device or a server computing device.

[0086] The computing device 50 includes a number of applications (e.g., applications 1 through N). Each application is in communication with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and model(s) stored therein) using an API (e.g., a common API across all applications).

[0087] The central intelligence layer includes a number of machine-learned models. For example, as illustrated in Figure 3C, a respective machine-learned model can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine-learned model. For example, in some implementations, the central intelligence layer can provide a single model for all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by an operating system of the computing device 50.

[0088] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for the computing device 50. As illustrated in Figure 3C, the central device data layer can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0089] The technology discussed herein makes reference to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein canbe implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0090] While the present subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of explanation, not limitation of the disclosure. Those skilled in the art, upon attaining an understanding of the foregoing, can readily produce alterations to, variations of, and equivalents to such embodiments. Accordingly, the subject disclosure does not preclude inclusion of such modifications, variations and / or additions to the present subject matter as would be readily apparent to one of ordinary skill in the art. For instance, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the present disclosure cover such alterations, variations, and equivalents.

Claims

WHAT IS CLAIMED IS:

1. A computing system configured to perform video interpolation via cascaded diffusion models, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store computerexecutable instructions for performing operations, the operations comprising: obtaining a pair of images comprising input versions of a start image and an end image that depict a scene, wherein the input versions of the start image and the end image have a input resolution; downsampling the input versions of the start image and the end image to generate downsampled versions of the start image and the end image that have a reduced resolution, the reduced resolution being smaller than the input resolution; processing a first noisy input with a first denoising diffusion model that is conditioned on the downsampled versions of the start image and the end image to generate, as an output of the first denoising diffusion model, one or more first synthetic interpolated images that depict content of the scene temporally between the start image and the end image, wherein the one or more first synthetic interpolated images have the reduced resolution; and processing a second noisy input with a second denoising diffusion model that is conditioned on the input versions of the start image and the end image and the one or more first synthetic interpolated images to generate, as an output of the second denoising diffusion model, one or more second synthetic interpolated images that depict content of the scene temporally between the start image and the end image, wherein the one or more second synthetic interpolated images have the input resolution.

2. The computing system of any preceding claim, wherein the one or more first synthetic interpolated images comprises a plurality of first synthetic interpolated images and the one or more second synthetic interpolated images comprises a plurality of second synthetic interpolated images.

3. The computing system of any preceding claim, wherein the second denoising diffusion model is conditioned on upsampled versions of the one or more first synthetic interpolated images.

4. The computing system of any preceding claim, wherein the first denoising diffusion model is conditioned on the downsampled versions of the start image and the end image by concatenating the downsampled versions of the start image and the end image with the first noisy input on a channel axis before input to the first denoising diffusion model.

5. The computing system of claim 4, wherein minimum noise levels are set for the downsampled versions of the start image and the end image.

6. The computing system of any preceding claim, wherein the second denoising diffusion model is conditioned on the input versions of the start image and the end image by concatenating the input versions of the start image and the end image with the second noisy input on a channel axis before input to the second denoising diffusion model.

7. The computing system of claim 6, wherein minimum noise levels are set for the input versions of the start image and the end image.

8. The computing system of any preceding claim, wherein one or both of the first denoising diffusion model and the second denoising diffusion model share one or more convolution and self-attention blocks over frames and feature maps only mix over frames via temporal attention blocks.

9. The computing system of any preceding claim, wherein the input resolution comprises 256x256 pixels and the reduced resolution comprises 64x64 pixels.

10. The computing system of any preceding claim, wherein the operations further comprise performing classifier-free guidance on when conditioning one or both of the first denoising diffusion model and the second denoising diffusion model.

11. A computer-implemented method for training diffusion models to perform video interpolation, the method comprising: obtaining, by a computing system comprising one or more computing devices, a training tuple comprising input versions of a start image and an end image that depict a scene and one or more ground truth images that depict content of the scene temporally between the start image and the end image, wherein the input versions of the start image and the end image have a input resolution; downsampling, by the computing system, the input versions of the start image and the end image to generate downsampled versions of the start image and the end image that have a reduced resolution, the reduced resolution being smaller than the input resolution; processing, by the computing system, a first noisy input with a first denoising diffusion model that is conditioned on the downsampled versions of the start image and the end image to generate, as an output of the first denoising diffusion model, one or more first synthetic interpolated images that depict content of the scene temporally between the start image and the end image, wherein the one or more first synthetic interpolated images have the reduced resolution; evaluating, by the computing system, a first loss function that compares: first added noise added to the ground truth images to generate the first noisy input, and first predicted noise removed from the first noisy input to generate the one or more first synthetic interpolated images; and modifying, by the computing system, one or more first values of one or more first parameters of the first denoising diffusion model based on the first loss function.

12. The computer-implemented method of claim 11, further comprising: processing a second noisy input with a second denoising diffusion model that is conditioned on the input versions of the start image and the end image and the one or more first synthetic interpolated images to generate, as an output of the second denoising diffusion model, one or more second synthetic interpolated images that depict content of the scene temporally between the start image and the end image, wherein the one or more second synthetic interpolated images have the input resolution; evaluating a second loss function that compares: second added noise added to the ground truth images to generate the second noisy input, and second predicted noise removedfrom the second noisy input to generate the one or more second synthetic interpolated images; and modifying one or more second values of one or more second parameters of the second denoising diffusion model based on the second loss function.

13. The computer-implemented method of claim 11 or 12, wherein the one or more first synthetic interpolated images comprises a plurality of first synthetic interpolated images and the one or more second synthetic interpolated images comprises a plurality of second synthetic interpolated images.

14. The computer-implemented method of any of claims 11-13, wherein the second denoising diffusion model is conditioned on upsampled versions of the one or more first synthetic interpolated images.

15. The computer-implemented method of any of claims 11-14, wherein the first denoising diffusion model is conditioned on the downsampled versions of the start image and the end image by concatenating the downsampled versions of the start image and the end image with the first noisy input on a channel axis before input to the first denoising diffusion model.

16. The computer-implemented method of claim 15, wherein minimum noise levels are set for the downsampled versions of the start image and the end image.

17. The computer-implemented method of any of claims 11-16, wherein the second denoising diffusion model is conditioned on the input versions of the start image and the end image by concatenating the input versions of the start image and the end image with the second noisy input on a channel axis before input to the second denoising diffusion model.

18. The computer-implemented method of claim 17, wherein minimum noise levels are set for the input versions of the start image and the end image.

19. The computer-implemented method of any of claims 11-18, wherein the operations further comprise performing classifier-free guidance on when conditioning one or both of the first denoising diffusion model and the second denoising diffusion model.

20. One or more non-transitory computer-readable media that collectively store the first denoising diffusion model and the second denoising diffusion model described in any preceding claim.

Citation Information

Patent Citations

  • Video generation method and device, model training method and device, equipment and medium

    CN116320216A