Image de-noising using guided updates

Through the stimulation accompanying guidance (SAG) method, multi-step denoising and stimulation accompanying numerical integration technology are used to solve the consistency and memory usage problems in the generation of diffusion model-guided images, and efficient and accurate image generation is achieved.

CN120410902APending Publication Date: 2025-08-01FACE CUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510115198.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-30
Filing Date
2025-01-24
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing diffusion model-guided image generation method does not require training, and the generation results are inconsistent with the one-step denoising approximation, resulting in inaccurate booting, and traditional backpropagation methods require a large amount of memory.

Method used

The final result is estimated by multi-step denoising, and the numerical integration method is used to perform gradient backpropagation to reduce memory usage and improve the accuracy of generated images.

Benefits of technology

More accurate image generation is achieved, reducing memory requirements, reducing computational costs, and improving the quality of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410902A_ABST
    Figure CN120410902A_ABST
Patent Text Reader

Abstract

The invention relates to image de-noising using boot updates. A computing system includes one or more processing devices configured to receive an image generation prompt and a reference image. The one or more processing devices compute a guide image over a plurality of denoising time steps by applying denoising updates to the generated image in a denoising diffusion model. At a subset of the de-noising time step, calculating the guidance image further includes applying a guidance update to the generated image based on the image generation hint, the reference image, and the generated image set. The one or more processing devices compute each guidance update by performing forward propagation and back propagation in the first integration time step and the second integration time step. The size of the generated image set and the number of the first integration time step and the second integration time step are each equal to a predefined integration time step count. The one or more processing devices output a final generated image calculated in the final de-noising time step.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Diffusion models are generative machine learning models used, for example, in image, video, and audio generation. In a diffusion model, noise is added to a data distribution to transform the data distribution into a simple noise distribution, such as a Gaussian distribution. Then, the diffusion model computes the inverse process of the noise addition process to generate new samples of the original data distribution. Thus, the diffusion model computes an output that matches the input distribution, such as an image, video, or sound. The input distribution can be computed based on an input that has a different modality compared to the output. For example, a diffusion model can compute an input distribution based on a text prompt and compute an image that is a sample of the input distribution for use in text-to-image synthesis. Summary of the Invention

[0002] According to one aspect of the present disclosure, there is provided a computing system including one or more processing devices configured to receive an image generation prompt and a reference image. The one or more processing devices are further configured to compute a guided image within a plurality of denoising time steps by at least partially iteratively applying a plurality of denoising updates to a generated image in a denoising diffusion model. At a subset of the plurality of denoising time steps, computing the guided image further includes: at least partially applying a corresponding guidance update to the generated image based on the image generation prompt, the reference image, and a set of generated images including the generated image at the current time step and the generated images at a plurality of previous time steps. The one or more processing devices are configured to compute each guidance update at least partially by: performing a forward pass on the set of generated images in a plurality of first integration time steps; and performing a backward pass on the set of generated images in a plurality of second integration time steps. The size of the set of generated images, the number of first integration time steps, and the number of second integration time steps are each equal to a predefined integration time step count. The one or more processing devices are further configured to output a final generated image computed at the last denoising time step of the plurality of denoising time steps as the guided image.

[0003] The Summary of the Invention is provided to introduce in a simplified form some concepts that will be further described in the Detailed Description below. The Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to embodiments that solve any or all of the disadvantages noted in any part of the present disclosure. Brief Description of the Drawings

[0004] Figure 1 Schematically shows a computing system according to an example, which includes one or more processing devices configured to generate a guided image.

[0005] Figure 2 Schematically shows a guided update calculation according to Figure 1 an example of

[0006] Figure 3A Schematically shows a guided update calculation according to Figure 1 more details of a computing system when one or more processing devices perform forward propagation according to an example of

[0007] Figure 3B Schematically shows a guided update calculation according to Figure 1 more details of a computing system when one or more processing devices perform backpropagation according to an example of

[0008] Figure 4 Schematically shows a guided update calculation according to Figure 1 a plurality of initial denoising time steps, a plurality of intermediate denoising time steps, and a plurality of subsequent denoising time steps that can be performed on one or more processing devices when calculating a guided image according to an example of

[0009] Figure 5 Schematically shows a guided update calculation according to Figure 1 a computing system when performing a plurality of self-recurrence time steps at one or more denoising time steps according to an example of [[ID=)]]

[0010] Figure 6 Shows an example algorithm that can be executed on a computing system to perform guided image generation according to an example of Figure 1 ]>

[0011] Figure 7 Schematically shows a computing system in an example of performing guided video generation according to an example of Figure 1

[0012] Figure 8 . Schematically shows a computing system in an example of using a trained feedback model when generating a guided image according to an example of Figure 1

[0013] Figure 9A Shows a flowchart of a method for use with a computing system to calculate a guided image according to an example of Figure 1

[0014] Figures 9B to 9D Shows additional steps of a method that can be performed in some examples of Figure 9A

[0015] Figure 10A Shows example images generated in a style-guided sampling experiment according to an example of Figure 1

[0016] ​​​​​​Figure 10B Shows an example image generated in an aesthetic-guided sampling experiment according to Figure 1 .

[0017] Figure 10C Shows an example image generated in an object-guided personalized sampling experiment according to Figure 1 .

[0018] Figure 10D Shows an example image generated in a face ID-guided personalized sampling experiment according to Figure 1 .

[0019] Figure 10E Shows an example guidance frame generated in a style-guided video editing experiment according to Figure 1 .

[0020] Figure 11A Shows an example guidance image generated at different values of a predefined integration time step count according to Figure 1 .

[0021] Figure 11B Shows a loss curve graph for different values of a predefined integration time step count according to Figure 1 .

[0022] Figure 11C Shows an example image generated at different values of a guidance strength hyperparameter according to Figure 1 .

[0023] Figure 12 Shows a schematic view of an example computing environment in which a Figure 1 computing system can be instantiated. DETAILED DESCRIPTION

[0024] In some previous methods, guidance sampling has been used with diffusion models to provide additional control over the output generation process. Guidance sampling has been used to control the outputs by conditioning the outputs of generative models according to various types of signals such as descriptive text, class labels, and images.

[0025] In some previous diffusion model guidance methods, guided sampling is performed by task-specific training of the diffusion model on paired data including the target output paired with the condition. For example, classifier guidance combines the score estimates computed in the diffusion model with the gradients computed in an image classifier to guide the generation process. Thus, classifier guidance can train the diffusion model to produce images corresponding to a specific class. Alternatively, classifier-free guidance directly trains a score estimator with the condition and uses a linear combination of the conditional and unconditional score estimators during sampling. Although training-based methods can effectively guide the diffusion model to generate data satisfying specified properties, the training-based diffusion model guidance methods have low flexibility due to the costs associated with training and the potential difficulty of collecting paired data.

[0026] Guidance methods without training are also used to perform guided sampling. In guidance sampling without training, at a certain sampling step \(t\), a guidance function is usually constructed based on the gradient of the loss function of the pre-trained diffusion model. More specifically, the guidance gradient is computed based on a one-step approximation of the denoised image of the noisy sample at the sampling step \(t\). Then the gradient is added to the corresponding sampling step as guidance for the generation process. Guidance methods without training allow the diffusion model to adapt to a wide range of guidance, thus providing greater flexibility. However, at some time steps of performing guidance, the generated results often do not match their one-step denoising approximation, resulting in inaccurate guidance. This inconsistency is obvious in the early steps of the generation process because the noisy samples are far from the final result. For example, in face ID-guided generation, when passing the blurred final approximation to a pre-trained face detection model, the pre-trained face detection model usually does not output accurate feature recognition. The inconsistency between the generated results and the one-step denoising approximation leads to inaccurate guidance for the specified input face.

[0027] To address the drawbacks of previous guided image generation methods, this paper introduces a diffusion model guidance method called Symplectic Adjoint Guidance (SAG). SAG is a guidance method without training. Compared with previous guidance methods without training, SAG estimates the final result through \(n\) steps of denoising. Multi-step sampling generates more accurate samples. However, multi-step sampling brings an additional challenge of backward propagating gradients from the output to each intermediate sampling step. If traditional backward propagation steps are performed, the backward propagation steps would require storing all intermediate states for \(n\) iterations, which would consume a large amount of memory. To reduce the memory used during backward propagation, SAG uses the symplectic adjoint numerical integration method when computing the guidance update, as discussed in further detail below. Thus, SAG achieves accurate gradient backward propagation while improving memory efficiency.

[0028] Figure 1 FIG. 1 schematically shows a computing system 10 that includes one or more processing devices 12 configured to generate a guidance image 54. The one or more processing devices 12 can include, for example, one or more central processing units (CPUs), graphics processing units (GPUs), tensor units, application specific integrated circuits (ASICs), and / or other types of processing devices 12.

[0029] The computing system 10 also includes one or more memory devices 14 coupled to the one or more processing devices 12. The one or more memory devices 14 can include volatile memory and non-volatile storage devices. In some examples, the computing system 10 is distributed across multiple physical computing devices, such as server computing devices located in one or more data centers. In other examples, the one or more processing devices 12 and the one or more memory devices 14 are included in a single physical computing device.

[0030] Figure 1 The example computing system 10 depicted in FIG. 1 also includes one or more input devices 16 and one or more display devices 18. The computing system 10 can also include one or more other types of output devices. In some examples, the one or more input devices 16 and the one or more display devices 18 can be included in a client computing device.

[0031] As Figure 1 depicted in the example of FIG. 1, the one or more processing devices 12 are configured to receive image generation conditions 20. These image generation conditions 20 include an image generation prompt 22, which can be a text prompt that describes the target semantic content and / or style features of the guidance image 54.

[0032] The image generation conditions 20 also include a reference image 24. In some examples, as Figure 1 depicted in FIG. 1, the reference image 24 is a style transfer reference image 24A. The style transfer reference image 24A received in such an example indicates the target style features of the guidance image 54. For example, the style features may be low-level features such as a color scheme or a texture pattern. In an example where the reference image 24 is the style transfer reference image 24A, the image generation prompt 22 specifies the semantic content of the guidance image 54.

[0033] In other examples, the reference image 24 can be the subject personalized reference image 24B. The subject personalized reference image 24B designates the target person or object indicated as being included in the guidance image 54. For example, the user can input their own photo as the subject personalized reference image 24B and can specify the target action (e.g., riding a bicycle) in the corresponding image generation prompt 22. In such an example, one or more processing devices 12 are configured to generate a guidance image 54 depicting these users performing the specified action. Thus, the subject personalized reference image 24B includes additional semantic data for description in the guidance image 54.

[0034] In other examples, other types of image data can be included in the reference image 24. Additionally or alternatively, data types other than the image generation prompt 22 and the reference image 24 can also be used as the image generation condition 20, as discussed in further detail below.

[0035] One or more processing devices 12 are also configured to compute the guidance image 54 within a plurality of denoising time steps 52 in the denoising diffusion model 30. In Figure 1 the example, the denoising diffusion model 30 is provided as a pre-trained noise prediction network ∈ θ . The noise prediction network ∈ θ includes a plurality of noise prediction network parameters θ, which can be included in multiple layers of a deep neural network.

[0036] One or more processing devices 12 are configured to process in a plurality of denoising time steps 52 starting from t = T,..., 0. At the plurality of denoising time steps 52, one or more processing devices 12 are configured to iteratively apply a plurality of denoising updates 34 to the generated image 32. One or more processing devices 12 can be configured to execute a scheduler to perform the denoising update 34 to compute the generated image x at the current time step t-1 .

[0037] When applying the denoising update 34, one or more processing devices 12 are configured to perform a noise process and a denoising process during each denoising time step 52. The forward noise process and the reverse denoising process can be performed by numerically solving the corresponding differential equation systems, as discussed in further detail below. The differential equation systems can be a set of stochastic differential equations (SDEs) or a set of ordinary differential equations (ODEs). The following discussion considers the denoising diffusion model 30 using a set of ODEs because the ODE-based diffusion model can be efficiently sampled in a deterministic manner.

[0038] The following discusses a denoising diffusion implicit model (DDIM) sampling method that can be used in the SAG method. At the DDIM sampler, one or more processing devices 12 are configured to solve the following ODE to perform discrete deterministic sampling: In the above equation, x t-1 is the generated image of the current time step calculated at the current denoising time step t, and x t is the generated image of the previous time step computed at the previous denoising time step t+1.

[0039] α t-1 is the value of the noise scheduling hyperparameter for the current denoising time step t. The one or more processing devices 12 are configured to modify the noise scheduling hyperparameter over multiple denoising time steps to adjust the amount of noise added to the generated image at different denoising time steps.

[0040] Noise prediction network ∈ θ is configured to reverse the noise process. At each denoising time step 52, the noise prediction network ∈ θ is configured to receive the generated image x at the previous time step t and the number of denoising time steps t as input.

[0041] is the estimated clean image, which is computed as an approximation of the fully denoised output image. It can be calculated according to the following equation:

[0042] Equation 1 can be used Parameterized, because σ t is monotonic in t. With this parameterization, we can also define the quantity When σ t-1 -σ t →0, we get the following ODE: In the above equation, By using the above ODE form of Equation 1, numerical methods can be used to accelerate sampling when performing denoising. As discussed in further detail below, the SAG method presented herein modifies the SAG in a way that takes into account the guidance. Calculation.

[0043] Back to Figure 1 For example, to perform the denoising update 34 during the denoising time step 52, the scheduler is configured to receive the image generation condition 20 and the generated image x of the previous time stept 。The scheduler is also configured to compute the generated image X at the current time step at least in part based on the output of the noise prediction network given these inputs ∈ θ to compute the generated image X at the current time step t-1 。At each denoising time step 52, the scheduler is correspondingly configured to estimate the noise in the generated image X at the previous time step t and remove that noise. The scheduler can for example be the DDIM scheduler discussed above. Other types of denoising diffusion schedulers can be used in other examples.

[0044] The computation of the guidance update 50 according to the SAG method is discussed below. At the corresponding denoising time step 52, one or more processing devices 12 are configured to apply these guidance updates 50 to the generated image 32. Each guidance update 50 is computed at least in part based on the image generation prompt 22, the reference image 24, the generated image x at the current time step t-1 and the generated images x at a plurality of previous time steps t and is computed. After the denoising update 34, the guidance update 50 can be applied to the generated image 32.

[0045] Figure 2 Schematically shows the computation of the guidance update 50 according to an example of Figure 1 According to an example of Figure 2 when computing the guidance update 50, one or more processing devices 12 are configured to perform a forward pass 40 over a plurality of first integration time steps 56. The number of first integration time steps 56 performed during the forward pass 40 is equal to a predefined integration time step count n. The predefined integration time step count n is a hyperparameter of the SAG process and is much smaller than the total number T of denoising time steps. For example, n can be equal to 4 or 5. At the first integration time step 56, one or more processing devices 12 are configured to compute estimates of n subsequent denoised images starting from the generated image x at the previous time step t up to the estimated clean image In an example of Figure 2 the intermediate estimate of the generated image x at the previous time step within the guidance time step 50 t is denoted as x t ′, and the estimated clean image is equal to x0′.

[0046] One or more processing devices 12 are also configured to compute a guidance loss 42 after performing the forward pass 40. The guidance loss 42 is at least in part based on the estimated clean image calculated with reference image 24. For example, the guidance loss 42 can be calculated as the L2 norm between the Gram matrix of the reference image 24 and the Gram matrix of the estimated clean image of.

[0047] When performing guided image generation at the denoising diffusion model 30, a guidance function can be added to the diffusion ODE of Equation 3. The resulting guided diffusion ODE can be expressed as follows: In the above equation, is the guidance strength hyperparameter, c is a set of image generation conditions 20, and is the guidance function. Gradient. For example, when the reference image 24 is the style transfer reference image 24A, this loss can be the style loss between the style transfer reference image 24A. During the calculation of the guidance update 50, one or more processing devices 12 are configured to generate images x for a set of previous time steps t as well as the generated image x at the current time step t-1 Perform forward propagation 40. As the output of forward propagation 40, one or more processing devices 12 are configured to obtain the value of the loss gradient. One or more processing devices 12 are also configured to solve the guided diffusion ODE in backpropagation 44 to calculate the guidance update 50.

[0049] Since the pre-trained denoising diffusion model 30 is trained using training images that do not contain noise, if it is used to directly obtain the loss value of the noisy input The loss gradient value obtained by the pre-trained denoising diffusion model 30 may be inaccurate. Instead, one or more processing devices 12 can be configured to use to approximate the loss where, is the estimated clean image discussed above with reference to Equation 2. Therefore, one or more processing devices 12 are configured to calculate the guidance loss 42.

[0050] Back to Figure 2 example, one or more processing devices 12 are also configured to perform backpropagation 44 in a plurality of second integration time steps 58. During backpropagation 44, one or more processing devices 12 are configured to calculate a plurality of adjoint states which represents the gradient of the guidance loss 42 with respect to the intermediate state x obtained as the generated image at the previous time step t ′. The pair of the intermediate state and the corresponding adjoint state (x t ′, a t) can be used as an augmented state during backpropagation 44. As shown in the example of Figure 2 , by integrating the augmented state backward in time, one or more processing devices 12 can be configured to compute a corresponding gradient estimate at a second integration time step 58.

[0051] In previous guided image generation methods, the following reverse ODE was used to obtain the gradient with respect to the intermediate state : After obtaining the gradient by solving the above differential equation , the gradient can be computed using the definition as However, when numerically solving the above differential equation, the previous backpropagation method using the adjoint state often results in large errors. Reducing these errors via the traditional method of reducing the step size leads to a significant increase in computational cost. To avoid these errors and the increase in computational cost, the SAG method uses the symplectic adjoint method to solve the differential equation, as discussed in further detail below.

[0052] Figure 3A More details of the computing system 10 are schematically shown when one or more processing devices 12 perform the forward propagation 40. The forward propagation 40 is performed at least in part based on generating an image set 60 that includes the generated image x at the current time step t-1 and the generated images x at multiple previous time steps t . The generated image x at the current time step t-1 and the generated images x at the previous time steps t are used to compute the estimated clean image

[0053] As discussed above, in previous image-guided methods, when only the generated image x at one previous time step t is used to compute the estimated clean image , the estimated clean image is often inconsistent with the finally output image, resulting in image artifacts. This inconsistency is particularly severe in the early stages of denoising iterations, when the noise samples are far from the final output. However, during the forward propagation, using multiple generated images x at previous time steps t-1 in addition to the generated image x at the current time step t can avoid the inconsistency between the estimated clean image and the guided image 54.

[0054] Since explicitly utilizing all the generated images 32 included in the generated image set 60 in each instance of the forward propagation 40 would consume a large amount of memory, one or more processing devices 12 are configured to iteratively compute the estimated clean image by numerically solving the first ordinary differential equation (ODE) 62 within a plurality of first integration time steps 56. When solving the first ODE 62, the noise prediction network ∈ θ , the noise schedule hyperparameter α t , the generated image x at the current time step t-1 , the generated image x from the previous time step of the immediately preceding denoising time step 52 t , and the current value of the estimated clean image are used as inputs.

[0055] The first ODE 62 can be expressed as follows: (Equation 6) In the first ODE 62 above, τ = n,..., 1 is the current first integration time step 56. x′ τ is an intermediate state of the process of predicting the estimated clean image . At the start of the clean image estimation process, one or more processing devices are configured to set x′ n = x t . Additionally, at the end of the clean image estimation process, Equation 6 is the discrete form of Equation 3.

[0056] As Figure 3A depicted in the example of, one or more processing devices 12 are configured to execute a numerical solver 64 to solve the first ODE 62. The numerical solver 64 is configured to use a symplectic solving method. For example, the numerical solver 64 can be a symplectic Euler solver 64A configured to execute the symplectic Euler method, or a symplectic Runge - Kutta solver 64B configured to execute the symplectic Runge - Kutta method. Equation 6 provides the update rule of the numerical solver 64 in the example where the numerical solver 64 is the symplectic Euler solver 64A. In the numerical solver 64, one or more processing devices 12 are configured to iteratively incorporate the data of the generated image x t from the previous time step into the calculation of the estimated clean image without storing the generated images X t of multiple previous time steps in volatile memory simultaneously.

[0057] Figure 3B Schematically shows more details of the computing system 10 when calculating the gradient 72 of the computational guidance loss 42 using the symplectic adjoint method during the backpropagation 44. InFigure 3B The gradient 72 calculated in the example of As Figure 3B Depicted, performing backpropagation 44 includes numerically solving the second ODE 70 within a plurality of second integration time steps 58. The number of second integration time steps 58 is also equal to a predefined integration time step count n.

[0058] One or more processing devices 12 are configured to perform backpropagation 44 on the generated image set 60. The guidance loss 42, the noise prediction network ∈ θ And the noise schedule hyperparameter α t Are also used as inputs to the second ODE 70. The inputs to the second ODE 70 also include the discrete step size One or more processing devices 12 are configured to iteratively calculate the gradient 72 within a plurality of second integration time steps 58 and use the calculated value of the gradient 72 as an input for subsequent second integration time steps 58.

[0059] In an example where the numerical solver 64 is the symplectic Euler solver 64A, the update rule given by the following equation can be used to solve the second ODE 70 given by Equation 5: Equations 7 and 8 are the discrete forms of Equation 5. In the above Equations 7 and 8, Is calculated at the corresponding second integration time step 58 Estimated value of. One or more processing devices 12 are configured to iteratively calculate the gradient Where τ = 0, 1,..., n - 1. After obtaining One or more processing devices 12 are also configured to calculate

[0060] Compared with using To update And The previous adjoint guidance method, the SAG method uses To update And The value used in backpropagation 44 is recovered from the value calculated in forward propagation 40. [[ID=4?]]

[0061] For the gradient calculated as the analytical solution of the continuous ODE of Equation 5 And the gradient calculated using the symplectic Euler solver 64A of Equation 8 Under a set of regular conditions, These regular conditions specify that the quantity s(δ, λ) = λ T δ is time-invariant, where And

[0062] After computing the gradients 72 in the reverse pass 44, one or more processing devices 12 are also configured to compute each guidance update 50 as the product of a guidance strength hyperparameter ρ t and the gradient 72 of the guidance loss 42. In some examples, the guidance strength hyperparameter ρ t varies over the course of multiple denoising time steps 52, while in other examples, the guidance strength hyperparameter ρ t remains constant. The guidance strength hyperparameter ρ t indicates the amount by which one or more processing devices 12 are configured to scale the gradient 72 when applying the guidance update 50 to the generated image x at the current time step t-1 .

[0063] Returning to Figure 1 the example of, one or more processing devices 12 are configured to output the final generated image 32 computed at the last denoising time step 52 of the multiple denoising time steps 52 as the guidance image 54. The guidance image 54 can be output to the display device 18 included in the computing system 10. In some examples, one or more processing devices 12 can be configured to transmit the guidance image 54 to a client computing device over a network.

[0064] In some examples, as Figure 4 shown, one or more processing devices 12 can be configured to compute the guidance image 54 in multiple initial denoising time steps 52A, multiple intermediate denoising time steps 52B, and multiple subsequent denoising time steps 52C. One or more processing devices 12 can be configured to apply the guidance update 50 at the multiple intermediate denoising time steps 52B. In such examples, the corresponding guidance update 50 is not performed at the initial denoising time step 52A or the subsequent denoising time step 52C. Early in the denoising diffusion process, the generated image 32 contains less information about the final output image than in the later denoising time steps 52. Late in the denoising diffusion process, the generated image 32 changes little between denoising time steps 52. By performing guidance during the multiple intermediate denoising time steps 52B, one or more processing devices 12 accordingly time the guidance to increase its impact compared to the early and late stages of the multiple denoising time steps 52. For example, guidance can be performed at 30% to 70% or from 20% to 60% of the total number T of denoising time steps 52. In other examples, other ranges of denoising time steps 52 can be used as the intermediate denoising time steps 52B.

[0065] In some examples, as Figure 5As schematically shown, one or more processing devices 12 may be configured to perform a plurality of self-loop time steps 80 at one or more denoising time steps 52. In each self-loop time step 80 (also referred to as a time travel step), one or more processing devices 12 may also be configured to repeatedly denoise and add noise to the generated image 32. Thus, one or more processing devices 12 may be configured to perform a plurality of denoising updates 34 and corresponding noise addition updates 82 during the denoising time step 52. Each noise addition update 82 includes adding noise to the generated image Xt of the previous time step. By performing a plurality of self-loop time steps during the denoising time step 52, one or more processing devices 12 may be configured to reduce visual artifacts that may otherwise occur due to adding the guidance update 50 to the generated image 32.

[0066] One or more processing devices 12 may be configured to calculate and apply the guidance update 50 at one or more self-loop time steps 80. The one or more self-loop time steps 80 at which the guidance update 50 is performed may occur during the intermediate denoising time step 52B. In contrast, one or more processing devices 12 may be configured not to perform the guidance update 50 during self-loop time steps 80 that do not occur during the initial denoising time step 52A and the subsequent denoising time step 52C.

[0067] Figure 6 An example algorithm 90 is shown that can be executed on the computing system 10 to perform guidance image generation. The algorithm 90 receives a pre-trained noise prediction network ∈ θ 、image generation condition c, loss function L, sampling scheduler guidance strength hyperparameter ρ t 、noise schedule hyperparameter α t 、guidance indicator array [g T ,...,g1] and a plurality of self-loop repetition counts (r T ,...,r1) as inputs. The guidance indicator array [g T ,...,g1] specifies the denoising time step 52 at which the guidance update 50 is performed. Thus, the guidance indicator array [g T ,...,g1] may specify a plurality of intermediate denoising time steps 52B. The self-loop repetition counts (r T ,...,r1) are the number of self-loop time steps 80 performed during the corresponding denoising time step 52.

[0068] In the algorithm 90, an initial generated image x is sampled from a normal distribution T ,where I is the identity matrix. For x TAfter initialization, the algorithm includes a first loop for denoising time steps \(t = T,\cdots,1\).

[0069] Each iteration of the first loop of Algorithm 90 includes one or more iterations of a second loop for self-loop time steps \(i = r\) t ,\(\cdots,1\). At each self-loop time step \(i\), Algorithm 90 includes updating the generated image \(x\) of the current time step by sampling at the scheduler such that t-1 Therefore, a denoising update 34 is performed on the generated image \(x\) of the previous time step t t t

[0070] After performing the denoising update 34 during the self-loop time step \(i\), Algorithm 90 also includes checking whether the guidance indicator \(g\) of the current denoising time step \(t\) is set to true t If \(g\) t is true, the algorithm also includes calculating the estimated clean image by solving Equation 6 within \(n\) integration time steps in the forward propagation 40 where \(n\) is a predefined integration time step count.

[0071] Algorithm 90 also includes calculating the loss gradient by solving Equation 7 and Equation 8 The loss gradient is also calculated within \(n\) integration time steps. After calculating the loss gradient, Algorithm 90 also includes updating the generated image \(x\) of the current time step by subtracting the guidance update 50 given by t-1 from the generated image \(x\) of the current time step t-1 .

[0072] During the self-loop time step \(i\), Algorithm 90 also includes a noise addition update 82. During the noise addition update 82, Algorithm 90 includes updating the generated image \(x\) of the previous time step according to the following equation t : (Equation 9) In the above equation, \(\in'\) is sampled from a normal distribution Therefore, the self-loop time step \(i\) includes preparing the generated image \(x\) of the previous time step for subsequent self-loop time steps t .

[0073] The values of the generated image \(x\) of the previous time step calculated in the last self-loop time step \(i\) of the denoising time step \(t\) t and the generated image \(x\) of the current time step t-1 are the \(x\) that can be used in subsequent denoising time steps \(t\)​t and x t-1 value. The value of x calculated in the last denoising time step t-1 is output as the guidance image 54.

[0074] Figure 7 Schematically shows the computing system 10 in an example of performing guidance video generation. In Figure 7 the example, one or more processing devices 12 are also configured to receive an input video 100 including a plurality of frames 102. One or more processing devices 12 are also configured to calculate corresponding depth maps 104 for the frames 102 of the input video 100.

[0075] In Figure 7 the example, one or more processing devices 12 are also configured to calculate a guidance video 110 at least in part based on the depth maps 104. The depth maps 104 are included in the image generation conditions 20, and one or more processing devices 12 are configured to use these conditions to adjust the denoising process. In Figure 7 the example, the image generation conditions 20 further include an image generation prompt 22 and a reference image 24. In Figure 7 the example, the calculated guidance video 110 includes a plurality of guidance frames 112, which can be calculated using the guidance image generation techniques discussed above. For example, one or more processing devices 12 can be configured to perform style transfer or object-guided personalization according to the reference image 24 and apply it to the guidance frames 112 of the guidance video 110. One or more processing devices 12 can also match the guidance video 110 with a text description provided as the image generation prompt 22.

[0076] One or more processing devices 12 are also configured to output the guidance video 110. For example, the guidance video 110 can be presented for display on a display device 18.

[0077] In some examples, as Figure 8 depicted, one or more processing devices 12 are configured to use a trained feedback model 120 to generate the guidance image 54. The trained feedback model 120 can be a prediction model including a plurality of feedback model parameters 122. For example, the trained feedback model 120 can be an aesthetic prediction model 120A trained to predict the aesthetic evaluations of human raters.

[0078] At a subset of the multiple denoising time steps 52 for performing guidance, one or more processing devices 12 can also be configured to calculate a feedback model reward value 124 at least in part based on the generated image x at the current time step t-1 to calculate the feedback model reward value 124. It can be at least in part by using the generated image x at the current time step t-1It is input into the trained feedback model 120 to calculate the feedback model reward value 124. Then, the feedback model reward value 124 can be included in the image generation condition 20. Thus, the feedback model reward value 124 can be utilized when calculating the denoising update 34 and the guidance loss 42. In an example where the feedback model 120 is the aesthetic prediction model 120A, the feedback model reward value 124 can be used to guide the calculation of the guidance image 54 towards an image predicted to have a high aesthetic value (as indicated by its feedback model reward value 124).

[0079] Figure 9A A flowchart of a method 200 for use with a computing system to calculate a guidance image according to the SAG method is shown. In step 202, method 200 includes receiving an image generation prompt. The image generation prompt can be a text prompt specifying the image generation conditions for the guidance image. In step 204, method 200 further includes receiving a reference image as an additional image generation condition. For example, the reference image can be a style transfer reference image or a subject personalization reference image.

[0080] In step 206, method 200 further includes calculating the guidance image within a plurality of denoising time steps. Performing these denoising time steps includes: in step 208, in the denoising diffusion model, iteratively applying a plurality of denoising updates to the generated image. The denoising diffusion model is a pre-trained noise prediction network that has been trained to reverse the noise process. The denoising updates can be calculated in a scheduler that performs a sampling process according to the image generation prompt and the reference image.

[0081] In step 210, step 206 further includes applying corresponding guidance updates to the generated image at a subset of the plurality of denoising time steps. These guidance updates are applied at least in part based on the image generation prompt, the reference image, and a set of generated images including the generated image at the current time step and the generated images at a plurality of previous time steps. Each guidance update is calculated as a modification to the generated image at the current time step. For example, each guidance update can be calculated as the product of a guidance strength hyperparameter and the gradient of the guidance loss.

[0082] Step 210 includes: in step 212, performing a forward pass on the set of generated images in a plurality of first integration time steps. The guidance loss can be calculated after performing the forward pass. Additionally, in step 214, step 210 further includes performing a backward pass on the set of generated images in a plurality of second integration time steps. The backward pass can output the gradient of the guidance loss. The size of the set of generated images, the number of first integration time steps, and the number of second integration time steps each equal a predefined integration time step count. The predefined integration time step count is a hyperparameter of the SAG method and can be equal to 4 or 5, for example.

[0083] In step 216, method 200 further includes outputting the final generated image calculated at the last denoising time step among the plurality of denoising time steps as a guidance image. The guidance image can be output to a display device. In some examples, the guidance image can be calculated at a first physical computing device (e.g., a server computing device) and output for display at a second computing device (e.g., a client computing device).

[0084] Figure 9B Illustrates additional steps of method 200 that can be performed when calculating the guidance update in step 208. As Figure 9B depicted, step 218 can be performed during the forward pass of step 212. In step 218, step 208 can further include calculating an estimated clean image at least in part based on the set of generated images. The estimated clean image is an approximation of the finally generated image. Step 218 can include: in step 220, numerically solving a first ordinary differential equation (ODE) within a plurality of first integration time steps. For example, the symplectic Euler method or the symplectic Runge-Kutta method can be used to solve the first ODE. In step 222, step 220 can include solving the first ODE at least in part based on noise schedule hyperparameters that vary within the plurality of denoising time steps. The noise schedule hyperparameters can be varied to adjust the magnitude of the guidance update during the guidance image generation process.

[0085] In step 224, step 208 can further include calculating a guidance loss at least in part based on the estimated clean image and the reference image. For example, the guidance loss can be calculated as the L2 norm between the Gram matrix of the reference image and the Gram matrix of the estimated clean image.

[0086] Step 226 can be performed during the backward pass of step 214. In step 226, method 200 can further include numerically solving a second ODE within a plurality of second integration time steps. The second ODE can also be solved using the symplectic Euler method or the symplectic Runge-Kutta method. In step 228, step 226 can include solving the second ODE at least in part based on the noise schedule hyperparameters. Thus, when using the noise schedule hyperparameters to calculate the guidance update, the values of the noise schedule hyperparameters for the corresponding denoising time steps are utilized during both the forward pass and the backward pass.

[0087] Figure 9CShows additional steps of method 200 performed in some examples when performing multiple denoising time steps in step 206. In step 230, method 200 may include performing a plurality of initial denoising time steps. During the initial time steps, no guidance updates are applied to the generated image. In step 232, method 200 may further include performing a plurality of intermediate denoising time steps at which guidance updates are applied. In step 234, method 200 may further include performing a plurality of subsequent denoising time steps at which no guidance updates are applied. Thus, the guidance updates may be performed during an intermediate stage of image generation, during which they may have a more significant impact on the finally generated image compared to the initial and subsequent stages. For example, the intermediate stage may include denoising time steps performed from 30% to 70% or from 20% to 60% of the plurality of denoising time steps.

[0088] Figure 9D Shows steps of method 200 that may be performed in an example of generating a guidance video. In step 236, method 200 may further include receiving an input video including a plurality of frames. In step 238, method 200 may further include calculating corresponding depth maps for the frames of the input video. The plurality of depth maps may be included in the image generation conditions together with a reference image and an image generation prompt.

[0089] In step 240, at least partially based on the depth maps, method 200 may further include calculating a guidance video including a plurality of guidance frames. For example, when the reference image is a style transfer reference image, the guidance frames may be guidance images to which the style of the reference image is transferred in a manner that preserves the depth relationships between the regions of the frames of the input video. In step 240, method 200 further includes outputting the guidance video.

[0090] The experimental results of the SAG method are discussed below. In the first experiment, style-guided sampling was performed using a style transfer reference image. To perform style-guided sampling, the features in the third layer of a pre-trained CLIP image encoder were used as feature vectors. The loss function was the L2 norm between the Gram matrix of the style transfer reference image and the Gram matrix of the estimated clean image, which was calculated using the CLIP feature vectors. StableDiffusion was used as the denoising diffusion model, and the pre-trained integral time step count was set to n = 4. The number of denoising time steps was set to T = 100, where guidance was applied from step t = 70 to t = 31. The number of self-loop iterations was set to r t = 1 from denoising time step 70 to 61, and set to r t = 2 from denoising time step 60 to 31.

[0091] The style-guided images generated using SAG were compared with the results obtained using the Free-Form Diffusion Model (FreeDoM) method and the results obtained using the Universal Guidance (UG) method. To obtain quantitative results for these three guidance image generation techniques, five style images and four prompts were randomly selected. For each technique, five images were generated for each style and each prompt. Figure 10A Example images 300 generated for two different styles of images when the image generation prompt was "a cat wearing glasses" are shown.

[0092] The following table shows the quantitative results obtained in the style-guided sampling experiment: Method Style loss (↓) CLIP (↑) FreeDoM 482.7 22.37 UG 805 23.02 SAG 386.6 23.51 As shown in the above table, SAG achieved the highest performance among these three techniques in terms of style loss and CLIP score.

[0093] In the second experiment, SAG was tested in the aesthetic-guided sampling task. SAG was tested using LAION, PickScore, and HPSv2 as aesthetic prediction models. The LAION aesthetic predictor is a linear head pre-trained on top of the CLIP visual embedding to predict a value ranging from 1 to 10, which represents the predicted aesthetic quality of the image. PickScore and HPSv2 are two reward functions trained on human preference data. In the aesthetic-guided sampling experiment, StableDiffusion was used as the denoising diffusion model, and the pre-trained integral time step count was set to n = 4. The feedback model reward value was calculated as the weighted sum of the scores output by LAION, PickScore, and HPSv2, with weights of 10, 2, and 0.5 respectively. The number of denoising time steps was set to T = 100, where guidance was applied from step t = 70 to t = 31. The number of self-loop iterations was set to r t = 2 from denoising time step 70 to 41, and set to r t = 1 from denoising time step 40 to 31.

[0094] In the aesthetic-guided sampling experiment, ten prompts were randomly selected from the following four prompt categories: animation, concept art, painting, and photo. One image was generated for each prompt. The resulting weighted aesthetic scores of all the generated images were compared with the baseline StableDiffusion (SD) v1.5, DOODL, and FreeDoM. Figure 10B Example images 310 generated in the aesthetic-guided sampling experiment are shown.

[0095] The following table shows the quantitative results obtained in the aesthetic-guided sampling experiment: Method Aesthetic loss (↓) SD v1.5 9.71 FreeDoM 9.18 DOODL 9.78 SAG 8.17 As shown in the above table, among the tested methods, SAG had the lowest aesthetic loss.

[0096] Personalized experiments were also conducted, including object-guided sampling experiments and face ID-guided sampling experiments. When calculating the guidance loss in the object-guided sampling experiment, the spherical distance loss was used to calculate the distance between the image features of the generated image and the reference image obtained from the ViT-H-14 CLIP model. StableDiffusion was adopted as the denoising diffusion model, and the pre-trained integral time step count was set to n = 4. The number of denoising time steps was set to T = 100, where guidance was applied from step t = 100 to t = 31. From denoising time step 100 to 31, the self-loop iteration count was set to r t = 2.

[0097] In the object-guided sampling experiment, the results of SAG were compared with DOODL, FreeDoM, and DreamBooth. The denoising diffusion model was fine-tuned for 400 steps on one training sample using the DreamBooth fine-tuning method, with a learning rate of 1×10 -6 . The cosine similarity between the CLIP embeddings of the generated image and the reference image (denoted as CLIP-I) and the cosine similarity between the CLIP embeddings of the generated image and the given text prompt (denoted as CLIP-T) were used to measure the model performance. An image of a dog was randomly selected as the reference, along with four prompts: "A dog (on the Acropolis / swimming / in a barrel / wearing sunglasses)." Four images were generated for each image for each prompt. Figure 10C Example image 320 generated in the object-guided sampling experiment is shown.

[0098] The following table shows the quantitative results of the object-guided sampling experiment: Method CLIP-I (↑) CLIP-T (↑) DreamBooth 0.724 0.277 FreeDoM 0.681 0.281 DOODL 0.743 0.277 SAG 0.774 0.270 As shown in the above table, the image generated by SAG has the highest CLIP image similarity with the reference image. Although the text similarity scores of these four methods are close to each other, FreeDoM has the highest similarity with the text prompt.

[0099] In another personalized experiment, face ID-guided sampling was conducted. In the face ID-guided sampling experiment, ArcFace was used to extract the target features of the reference face to represent the face ID. Additionally, ArcFace was also used to extract the features of the guidance image. The loss function is the l2 Euclidean distance between the face ID features extracted from the reference image and the guidance image. The pre-trained integration time step count was set to n = 5. Five face IDs were randomly selected, and 200 faces were generated for each ID. The face ID-guided generation results calculated using SAG were compared with those calculated using FreeDoM, represented by loss and Fréchet inception distance (FID). Figure 10D An example image 330 calculated in the face ID-guided generation experiment is shown.

[0100] The following table shows the quantitative results of the face ID-guided sampling experiment: Method ID loss (↓) FID (↓) FreeDoM 0.602 65.24 SAG 0.574 64.25 As shown in the above table, SAG is superior to FreeDoM in terms of both ID loss and FID.

[0101] A style-guided video editing experiment was also conducted. In the style-guided video editing experiment, MagicEdit was used as the denoising diffusion model. Given an input video, the depth map was extracted, and MagicEdit was used to generate a video with motion matching that of the input video based on the depth map and text prompts. Similar to the style-guided sampling experiment, the L2 loss between the Gram matrices of the generated image and the reference image was calculated. Since the depth map and text prompts provide a large amount of information about the final output video, the number of denoising time steps was set to T = 25. A 16-frame video with a size of 256×256 pixels per frame was rendered using MagicEdit. SAG guidance was applied at the denoising time steps t ∈ [20, 10].

[0102] Figure 10E An example guidance frame 340 generated in the style-guided video editing experiment is shown. As Figure 10E shown, SAG enables MagicEdit to generate videos in a specific style (e.g., a cat in the Chinese paper-cutting style). In contrast, without using SAG, the basic editing model can hardly synthesize a video with colors and textures consistent with the reference image.

[0103] Experiments were also conducted to test the effects of different choices of hyperparameter values. One such experiment tested different values of the predefined integral time step count n. This experiment was conducted using an image stylization task, where T = 100, and guidance was performed from step 70 to 31. For each tested value of n, 20 stylized images were generated using the prompts "a cat wearing glasses", "a butterfly", and "a photo of the Eiffel Tower". Example result images 400 are as Figure 11A shown. Additionally, Figure 11B a loss curve graph 410 for different values of n is shown. As Figure 11A shown, when n = 1, the generated images suffer from content distortion and the stylization effect is not very obvious. As n increases, the quality of the generated images also improves, and the reduction in the loss between the generated images and the style images becomes more obvious. However, when n increases above 4, the loss no longer decreases significantly.

[0104] Another experiment tested different values of the guidance strength hyperparameter ρ t . Again, the image stylization task was used as an example task, where the value of n was equal to 1 and 3. The guidance strength hyperparameter ρ t was gradually increased from 0.3 to 0.5. Figure 11C Example images 420 generated at different values of the guidance strength hyperparameter ρ t are shown. As Figure 11C depicted, as the guidance strength hyperparameter ρ t increases, the stylization becomes more obvious, but at higher values of ρ t , severe artifacts appear in the generated images.

[0105] Experiments were also conducted to test different ranges of the denoising time steps at which guidance is performed, and different numbers of self-loop time steps. As discussed above, the diffusion sampling process generally consists of three stages: a chaotic stage where x t has a lot of noise, a semantic stage where image semantic features appear, and a refinement stage where the generated results change minimally. Increasing the number of self-loop time steps r t extends the diffusion sampling process, helping to explore results that can both achieve guidance and ensure image quality. Therefore, in tasks such as stylization and aesthetic guidance that preserve the semantic content of the input image, a lower value of r t (e.g., r t = 2) can be used in the semantic stage. In contrast, for tasks such as object guidance or personalization such as face ID guidance, guidance can also be performed in the chaotic stage, and a larger value of r t (e.g., r t = 3) can be used.

[0106] As discussed above, the SAG method allows for guided image generation on denoising diffusion models in a way that does not require additional diffusion model training. SAG also avoids inconsistencies between the generated image and the estimated clean image, thus achieving higher image quality and reducing image artifacts. SAG also allows for memory-saving generation of guided images, which combines information from the generated images of previous time steps without having to store sequences of generated images of multiple previous time steps in volatile memory simultaneously. Therefore, SAG is capable of efficiently computing high-quality guided images.

[0107] In some embodiments, the methods and processes described herein can be related to the computing systems of one or more computing devices. Specifically, such methods and processes can be implemented as a computer application or service, an application programming interface (API), a library, and / or other computer program products.

[0108] Figure 12 A non-limiting embodiment of a computing system 500 that can implement one or more of the above methods and processes is schematically shown. The computing system 500 is shown in a simplified form. The computing system 500 can embody the computing system 10 described above and illustrated in Figure 1 The computing system 500 can take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, video game devices, mobile computing devices, mobile communication devices (e.g., smartphones), and / or other computing devices, as well as wearable computing devices (such as smartwatches and head-mounted augmented reality devices).

[0109] The computing system 500 includes a logical processor 502, a volatile memory 504, and a non-volatile storage device 506. The computing system 500 can optionally include a display subsystem 508, an input subsystem 510, a communication subsystem 512, and / or Figure 12 other components not shown in

[0110] The logical processor 502 includes one or more physical devices configured to execute instructions. For example, the logical processor can be configured to execute instructions as part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions can be implemented to perform a certain task, implement a certain data type, transform the state of one or more components, achieve a certain technical effect, or otherwise achieve an expected result.

[0111] A logical processor may include one or more physical processors configured to execute software instructions. Additionally or alternatively, a logical processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processors of the logical processor 502 may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and / or distributed processing. Optionally, the various components of the logical processor may be distributed among two or more separate devices, which may be remotely located and / or configured to coordinate processing. The various aspects of the logical processor may be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration. It will be understood that in such cases, these virtualized aspects run on different physical logical processors of various different machines.

[0112] The non-volatile storage device 506 includes one or more physical devices configured to store instructions executable by the logical processor to implement the methods and processes described herein. When implementing such methods and processes, the state of the non-volatile storage device 506 may change, for example to store different data.

[0113] The non-volatile storage device 506 may include removable and / or built-in physical devices. The non-volatile storage device 506 may include optical memory, semiconductor memory, and / or magnetic memory, or include other mass storage device technologies. The non-volatile storage device 506 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable devices. The non-volatile storage device 506 is configured to store instructions even when the power to the non-volatile storage device 506 is cut off.

[0114] The volatile memory 504 may include a physical device that includes random access memory. The volatile memory 504 is typically used by the logical processor 502 to temporarily store information during the processing of software instructions. It should be understood that when the power to the volatile memory 504 is cut off, the volatile memory 504 generally does not continue to store instructions.

[0115] Aspects of the logical processor 502, the volatile memory 504, and the non-volatile storage device 506 may be integrated together into one or more hardware logic components. Such hardware logic components may include field programmable gate arrays (FPGAs), program and application specific integrated circuits (PASIC / ASICs), program and application specific standard products (PSSP / ASSPs), system on a chip (SOCs), and complex programmable logic devices (CPLDs), among others.

[0116] The terms "module", "program", and "engine" can be used to describe an aspect of computing system 500 that is typically implemented in software by a processor to perform a specific function using portions of volatile memory, the function involving transformational processing for specifically configuring the processor to perform the function. Thus, a module, program, or engine can be instantiated by executing instructions stored in non-volatile storage device 506 using portions of volatile memory 504 via logical processor 502. It should be understood that different modules, programs, and / or engines can be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Also, the same module, program, and / or engine can be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module", "program", and "engine" can encompass single or multiple sets of executable files, data files, libraries, drivers, scripts, database records, etc.

[0117] Display subsystem 508 (when included) can be used to present a visual representation of data stored in non-volatile storage device 506. The visual representation can take the form of a graphical user interface (GUI). Since the methods and processes described herein change the data stored in the non-volatile storage device and thus transform the state of the non-volatile storage device, the state of display subsystem 508 can likewise be transformed to visually represent the changes in the underlying data. Display subsystem 508 can include one or more display devices utilizing almost any type of technology. Such display devices can be combined with logical processor 502, volatile memory 504, and / or non-volatile storage device 506 in a shared enclosure, or such display devices can be peripheral display devices.

[0118] Input subsystem 510 (when included) can include one or more user input devices, such as a keyboard, mouse, touch screen, or game controller, or interface with these user input devices. In some embodiments, the input subsystem can include selected natural user input (NUI) components, or interface with these components. These components can be integrated or peripheral components, and the conversion and / or processing of input actions can be performed on-board or off-board. Example NUI components can include a microphone for voice and / or speech recognition; infrared, color, stereo, and / or depth cameras for machine vision and / or gesture recognition; a head tracker, eye tracker, accelerometer, and / or gyroscope for motion detection and / or intent recognition; and an electric field sensing component for evaluating brain activity; and / or any other suitable sensor.

[0119] The communication subsystem 512 (when included) may be configured to communicatively couple the various computing devices described herein to each other and to other devices. The communication subsystem 512 may include wired and / or wireless communication devices compatible with one or more different communication protocols. As a non-limiting example, the communication subsystem may be configured to communicate via a wireless telephone network, or a wired or wireless local or wide area network. In some embodiments, the communication subsystem may allow the computing system 500 to send messages to and / or receive messages from other devices via a network such as the Internet.

[0120] The following paragraphs provide additional description of the subject matter of the present disclosure. According to one aspect of the present disclosure, there is provided a computing system including one or more processing devices configured to receive an image generation prompt and a reference image. The one or more processing devices are further configured to compute a guided image within a plurality of denoising timesteps by at least partially iteratively applying a plurality of denoising updates to a generated image in a denoising diffusion model. At a subset of the plurality of denoising timesteps, computing the guided image further includes applying a respective guidance update to the generated image based at least in part on the image generation prompt, the reference image, and a set of generated images including the generated image at the current timestep and the generated images at a plurality of previous timesteps. The one or more processing devices are configured to compute each guidance update at least in part by: performing a forward pass on the set of generated images over a plurality of first integration timesteps; and performing a backward pass on the set of generated images over a plurality of second integration timesteps. The size of the set of generated images, the number of first integration timesteps, and the number of second integration timesteps each equal a predefined integration timestep count. The one or more processing devices are further configured to output a final generated image computed at a last denoising timestep of the plurality of denoising timesteps as the guided image. The above features may have the following technical effects: performing guided image generation without additional training of the denoising diffusion model. These technical effects may further include improved memory efficiency and reduced image artifacts.

[0121] According to this aspect, the one or more processing devices may be configured to compute each guidance update as a product of a guidance strength hyperparameter and a gradient of a guidance loss. The above features may have the following technical effects: generating a guided image with an adjustable guidance strength.

[0122] According to this aspect, performing the forward pass may include computing an estimated clean image based at least in part on the set of generated images. The guidance loss may be computed based at least in part on the estimated clean image and the reference image. The above features may have the following technical effects: guiding image denoising using a loss function that depends on the reference image.

[0123] According to this aspect, performing forward propagation may include numerically solving a first ordinary differential equation (ODE) within a plurality of first integration time steps. Performing backpropagation may include numerically solving a second ODE within a plurality of second integration time steps. The above features may have the following technical effects: calculating the estimated clean image and the gradient of the guidance loss.

[0124] According to this aspect, one or more processing devices may be configured to solve the first ODE and the second ODE using the symplectic Euler method or the symplectic Runge-Kutta method. The above features may have the following technical effects: calculating the estimated clean image and the gradient in a memory-saving manner.

[0125] According to this aspect, one or more processing devices may be configured to calculate a guidance update based at least in part on a noise schedule hyperparameter that varies within a plurality of denoising time steps. The above features may have the following technical effects: changing the amount of noise applied to the generated image at different denoising time steps.

[0126] According to this aspect, one or more processing devices may be configured to apply the guidance update at a plurality of intermediate denoising time steps that are preceded by a plurality of initial denoising time steps and followed by a plurality of subsequent denoising time steps. The above features may have the following technical effects: timing the guidance of the image generation process such that the guidance can affect the semantic content of the guidance image.

[0127] According to this aspect, in one or more of the denoising time steps, one or more processing devices may be configured to repeatedly denoise and add noise to the generated image at each of a plurality of self-loop time steps. The above features may have the following technical effects: reducing visual artifacts that may otherwise occur due to adding the guidance update to the generated image.

[0128] According to this aspect, one or more processing devices are configured to calculate and apply the guidance update at one or more of the self-loop time steps. The above features may have the following technical effects: reducing artifacts in the guidance image by performing self-loop and guidance simultaneously at the same denoising time step.

[0129] According to this aspect, the reference image may be a style transfer reference image or a subject personalization reference image. The above features may have the following technical effects: performing style transfer on the generated image or inserting a user-selected subject into the generated image.

[0130] According to this aspect, one or more processing devices may also be configured to receive an input video including a plurality of frames and compute corresponding depth maps for the frames of the input video. One or more processing devices may also be configured to compute a guidance video including a plurality of guidance frames based at least in part on the depth maps. One or more processing devices may also be configured to output the guidance video. The above features may have the following technical effects: performing guidance video generation.

[0131] According to another aspect of the present disclosure, there is provided a method for use with a computing system, the method including receiving an image generation prompt and receiving a reference image. The method further includes computing a guidance image within a plurality of denoising timesteps by at least partially iteratively applying a plurality of denoising updates to a generated image in a denoising diffusion model. At a subset of the plurality of denoising timesteps, the method further includes applying a corresponding guidance update to the generated image based at least in part on the image generation prompt, the reference image, and a set of generated images including the generated image at the current timestep and the generated images at a plurality of previous timesteps. Computing each guidance update includes performing a forward pass on the set of generated images over a plurality of first integration timesteps; and performing a backward pass on the set of generated images over a plurality of second integration timesteps. The size of the set of generated images, the number of first integration timesteps, and the number of second integration timesteps are each equal to a predefined integration timestep count. The method further includes outputting a final generated image computed at a last denoising timestep of the plurality of denoising timesteps as the guidance image. The above features may have the following technical effects: performing guidance image generation without additional training of the denoising diffusion model. These technical effects may also include improving memory efficiency and reducing image artifacts.

[0132] According to this aspect, each guidance update may be computed as the product of a guidance strength hyperparameter and the gradient of a guidance loss. The above features may have the following technical effects: generating a guidance image with an adjustable guidance strength.

[0133] According to this aspect, performing the forward pass may include computing an estimated clean image based at least in part on the set of generated images. The guidance loss may be computed based at least in part on the estimated clean image and the reference image. The above features may have the following technical effects: guiding image denoising using a loss function that depends on the reference image.

[0134] According to this aspect, performing the forward pass may include numerically solving a first ordinary differential equation (ODE) over a plurality of first integration timesteps. Performing the backward pass may include numerically solving a second ODE over a plurality of second integration timesteps. The above features may have the following technical effects: computing the gradient of the estimated clean image and the guidance loss.

[0135] According to this aspect, the guidance update can be calculated based at least in part on noise schedule hyperparameters that vary within multiple denoising time steps. The above features can have the following technical effects: changing the amount of noise applied to the generated image at different denoising time steps.

[0136] According to this aspect, the guidance update can be applied at multiple intermediate denoising time steps, preceded by multiple initial denoising time steps and followed by multiple subsequent denoising time steps. The above features can have the following technical effects: timing the guidance of the image generation process such that the guidance can affect the semantic content of the guidance image.

[0137] According to this aspect, the reference image can be a style transfer reference image or a subject personalization reference image. The above features can have the following technical effects: performing style transfer on the generated image or inserting a user-selected subject into the generated image.

[0138] According to this aspect, the method can further include receiving an input video including multiple frames and calculating corresponding depth maps for the frames of the input video. The method can further include calculating a guidance video including multiple guidance frames based at least in part on the depth maps. The method can further include outputting the guidance video. The above features can have the following technical effects: performing guidance video generation.

[0139] According to another aspect of the present disclosure, a computing system is provided. The computing system includes one or more processing devices configured to receive an image generation prompt. The one or more processing devices are further configured to compute a guided image within a plurality of denoising timesteps by at least partially iteratively applying a plurality of denoising updates to a generated image in a denoising diffusion model. At a subset of the plurality of denoising timesteps, computing the guided image further includes: computing a feedback model reward value in a trained feedback model based at least in part on the generated image at the current timestep. At a subset of the plurality of denoising timesteps, computing the guided image further includes: applying a corresponding guided update to the generated image based at least in part on the image generation prompt, the feedback model reward value, and a set of generated images including the generated image at the current timestep and a set of generated images at a set of previous timesteps. The one or more processing devices are configured to compute each guided update by at least partially: performing a forward pass on the set of generated images over a plurality of first integration timesteps; and performing a backward pass on the set of generated images over a plurality of second integration timesteps. The size of the set of generated images, the number of first integration timesteps, and the number of second integration timesteps each equal a predefined integration timestep count. The one or more processing devices are further configured to output a final generated image computed at a last denoising timestep of the plurality of denoising timesteps as the guided image. The above features may have the following technical effects: Guided image generation can be performed without additional training of the denoising diffusion model. These technical effects may also include improved memory efficiency and reduced image artifacts.

[0140] As used herein, "and / or" is defined as inclusive or, ∨, as specified in the following truth table: A B A ∨ B True True True True False True False True True False False False

[0141] It should be understood that the configurations and / or methods described herein are exemplary in nature, and these specific embodiments or examples should not be considered restrictive as there may be various variations. The specific routines or methods described herein may represent one or more of any number of processing strategies. Accordingly, the various acts illustrated and / or described may be performed in the order illustrated and / or described, in other orders, in parallel, or omitted. Similarly, the order of the above processes may be changed.

[0142] The subject matter of the present disclosure includes all novel and non-obvious combinations and subcombinations of the various processes, systems, and configurations, as well as other features, functions, acts, and / or characteristics disclosed herein, and any and all equivalents thereof.

Claims

1. A computing system, comprising: One or more processing devices configured to: Receive an image generation prompt; Receive a reference image; Within a plurality of denoising time steps, calculate a guidance image at least in part by: At a denoising diffusion model, iteratively applying a plurality of denoising updates to a generated image; and And At a subset of the plurality of denoising time steps, at least in part based on the image generation prompt, the reference image, and a set of generated images, apply a corresponding guidance update to the generated image, the set of generated images including the generated image at the current time step and the generated images at a plurality of previous time steps, wherein the one or more processing devices are configured to calculate each guidance update at least in part by: Performing forward propagation on the set of generated images over a plurality of first integration time steps; and performing backward propagation on the set of generated images over a plurality of second integration time steps, wherein the size of the set of generated images, the number of first integration time steps, and the number of second integration time steps each equal a predefined integration time step count; and Output a final generated image calculated at the last denoising time step among the plurality of denoising time steps as the guidance image.

2. The computing system according to claim 1, wherein The one or more processing devices are configured to calculate each guidance update as a product of a guidance strength hyperparameter and a gradient of a guidance loss.

3. The computing system according to claim 2, wherein: Performing the forward propagation includes: calculating an estimated clean image at least in part based on the set of generated images; The guidance loss is calculated at least in part based on the estimated clean image and the reference image.

4. The computing system according to claim 1, wherein: Performing the forward propagation includes: numerically solving a first ordinary differential equation (ODE) over the plurality of first integration time steps; and Performing the backward propagation includes: numerically solving a second ODE over the plurality of second integration time steps.

5. The computing system according to claim 4, wherein, The one or more processing devices are configured to solve the first ODE and the second ODE using a symplectic Euler method or a symplectic Runge-Kutta method.

6. The computing system according to claim 1, wherein, The one or more processing devices are configured to calculate the guidance update at least in part based on a noise schedule hyperparameter that varies over the plurality of denoising time steps.

7. The computing system according to claim 1, wherein, The one or more processing devices are configured to apply the guidance update at a plurality of intermediate denoising time steps, preceded by a plurality of initial denoising time steps and followed by a plurality of subsequent denoising time steps.

8. The computing system according to claim 1, wherein, At one or more of the denoising time steps, the one or more processing devices are configured to repeat denoising and noise addition to the generated image at each of a plurality of self-loop time steps.

9. The computing system according to claim 8, wherein, The one or more processing devices are configured to calculate and apply the guidance update at one or more of the self-loop time steps.

10. The computing system according to claim 1, wherein, The reference image is a style transfer reference image or a subject personalization reference image.

11. The computing system according to claim 1, wherein, The one or more processing devices are further configured to: Receive an input video including a plurality of frames; Calculate the corresponding depth map of the frames of the input video; Calculate a guidance video including a plurality of guidance frames, at least partially based on the depth map; And Output the guidance video.

12. A method for use with a computing system, the method comprising: Receiving an image generation prompt; Receiving a reference image; Calculating a guidance image within a plurality of denoising time steps, at least partially by: Iteratively applying a plurality of denoising updates to the generated image at a denoising diffusion model; And At a subset of the plurality of denoising time steps, applying a corresponding guidance update to the generated image, at least partially based on the image generation prompt, the reference image, and a set of generated images, the set of generated images including the generated image at the current time step and the generated images at a plurality of previous time steps, wherein calculating each guidance update includes: Performing forward propagation on the set of generated images over a plurality of first integration time steps; and Performing backward propagation on the set of generated images over a plurality of second integration time steps, wherein the size of the set of generated images, the number of first integration time steps, and the number of second integration time steps each equal a predefined integration time step count; and Outputting the final generated image calculated at the last denoising time step among the plurality of denoising time steps as the guidance image.

13. The method according to claim 12, wherein, Each guidance update is calculated as the product of a guidance strength hyperparameter and the gradient of a guidance loss.

14. The method according to claim 13, wherein: Performing the forward propagation includes: calculating an estimated clean image, at least partially based on the set of generated images; The guidance loss is calculated at least partially based on the estimated clean image and the reference image.

15. The method according to claim 12, wherein: Performing the forward propagation includes numerically solving a first ordinary differential equation (ODE) over the plurality of first integration time steps; and Performing the backward propagation includes numerically solving a second ODE over the plurality of second integration time steps.

16. The method according to claim 12, wherein, The guidance update is calculated at least partially based on a noise schedule hyperparameter that varies over the plurality of denoising time steps.

17. The method according to claim 12, wherein, The guidance update is applied at a plurality of intermediate denoising time steps, preceded by a plurality of initial denoising time steps and followed by a plurality of subsequent denoising time steps.

18. The method according to claim 12, wherein, The reference image is a style transfer reference image or a subject personalization reference image.

19. The method according to claim 12, further comprising: Receiving an input video including a plurality of frames; Calculating the corresponding depth map of the frames of the input video; Calculating a guidance video including a plurality of guidance frames, at least partially based on the depth map; And Outputting the guidance video.

20. A computing system, comprising: One or more processing devices configured to: Receive an image generation prompt; Calculate a guidance image within a plurality of denoising time steps, at least partially by: Iteratively applying a plurality of denoising updates to the generated image in a denoising diffusion model; And At a subset of the plurality of denoising time steps: At the trained feedback model, calculate a feedback model reward value based at least in part on the generated image of the current time step; and Apply a corresponding guidance update to the generated image based at least in part on the image generation prompt, the feedback model reward value, and the set of generated images, the set of generated images including the generated image of the current time step and a set of generated images of a set of previous time steps, wherein the one or more processing devices are configured to calculate each guidance update at least in part by: Performing forward propagation on the set of generated images over a plurality of first integration time steps; and Performing backward propagation on the set of generated images over a plurality of second integration time steps, wherein the size of the set of generated images, the number of first integration time steps, and the number of second integration time steps each equal a predefined integration time step count; and Output the final generated image calculated in the last denoising time step of the plurality of denoising time steps as the guidance image.