Diffusion model noise inversion calculation method based on Euler discrete sampling

By using the noise inversion calculation method of the diffusion model through Euler discrete sampling, the noise is optimized to solve the problem of unsatisfactory generation results of the diffusion model, achieving efficient generation and wide applicability, and reducing the need for model fine-tuning.

CN121746530APending Publication Date: 2026-03-27NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing diffusion models require significant computational resources for fine-tuning when generating images and videos, and struggle to provide targeted initial noise based on given conditions, resulting in unsatisfactory generation results.

Method used

A noise inversion calculation method based on Eulerian discrete sampling is adopted. By combining Eulerian discrete sampling inference and inversion with a classifier-free guided method, the noise is optimized to improve the generation quality without the need for fine-tuning of the diffusion model.

Benefits of technology

It improves the quality and efficiency of the generated results, has strong generalization ability, is applicable to a variety of diffusion models, and reduces the consumption of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746530A_ABST
    Figure CN121746530A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of diffusion models, and particularly discloses a diffusion model noise inversion calculation method based on Euler discrete sampling. Optimized noise is collected, so that a model generation result is improved on the premise that the model is not finely adjusted. The method specifically comprises the steps of 1, performing multi-step reasoning on initial noise by using an Euler discrete sampling reasoning method, 2, performing multi-step inversion on the reasoned noise by using an Euler discrete sampling inversion method, and 3, screening the collected initial noise-optimized noise pair. The method is used in the fields of image generation, video generation and new view angle synthesis, and has a good market prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of noise inversion calculation technology, and in particular to a noise inversion calculation method based on a diffusion model using Euler discrete sampling. Background Technology

[0002] From artistic creation to medical diagnosis, from design and manufacturing to the entertainment industry, image and video generation technologies are profoundly changing the way we work and produce across all sectors with their high efficiency and innovative characteristics. Due to their powerful generative capabilities, diffusion models are gradually becoming the mainstream method for solving image generation, video generation, and novel perspective synthesis problems. Most methods improve the generated results by adjusting the diffusion model structure. However, this method requires a significant amount of computational resources.

[0003] Currently, some methods propose that selecting specific initial noise can improve the image and video generation performance of diffusion models. For example, the Sv3d model, a novel perspective synthesis model based on diffusion models, suffers from problems such as unreasonable generated results and a lack of local detail. MV-adapter, another novel perspective synthesis model based on diffusion models, also suffers from a mismatch between the generated result and the actual appearance contour size. BANSA proposed an active noise selection framework based on a principled Bayesian formula for attention uncertainty. DLBS-LA proposed a diffusion latent space bundle search method with a prospective estimator, which maximizes the given alignment reward by selecting a better diffusion latent space during inference. However, both of these methods search multiple noise sources and cannot provide targeted initial noise based on given conditions. How to provide targeted initial noise based on given conditions remains a crucial issue. Summary of the Invention

[0004] To address the problem that existing technologies require fine-tuning of diffusion models, which consumes significant computational resources and time, this invention aims to provide a diffusion model noise inversion calculation method based on Euler discrete sampling. This method eliminates the need for fine-tuning of the diffusion model, thereby greatly improving efficiency. By limiting sampling conditions, it optimizes noise, improves the quality of generated results, and has wide applicability, applicable to various diffusion models that use Euler discrete sampling inference and classifier-free guided methods.

[0005] To address the problems in the existing technology, the technical solution adopted by this invention is as follows:

[0006] A method for noise inversion calculation based on a diffusion model using Euler discrete sampling includes the following steps:

[0007] Step 1, given time step t is initially set to T, and random Gaussian noise is generated. The initial value of the number of inferences x is 1, and the noise is calculated. The mean and standard deviation of are denoted as . , ,right Perform Euler discrete sampling scaling and calculate the scaled initial noise. mean with standard deviation ,Will By feeding the given conditions and other conditions into the diffusion model network, the prediction results under the given conditions are obtained; The empty condition, along with other conditions, is fed into the diffusion model network to obtain the prediction results under the empty condition. The mean and standard deviation of the two prediction results are calculated respectively, denoted as . , and , Using a classifier-free guided method, the prediction results under two different conditions are linearly combined, and the result is... Using the Euler discrete sampling inference method, through and Obtain noise Increment the number of inferences x by 1;

[0008] Step 2: Determine if the number of inferences x is less than or equal to m, and m is less than T. If so, repeat step 1; otherwise, obtain... ,Will Considered Proceed to step 3;

[0009] Step 3, given time step The initial value of t is T-m+1, and the initial value of the given inversion number y is 1. For noise... Perform Euler discrete sampling scaling to obtain Then use and right Scaling by mean-standard deviation yields the noise. ,Will By feeding the given conditions and other conditions into the diffusion model network, the prediction results under the given conditions are obtained. The empty condition and other conditions are fed into the diffusion model network to obtain the prediction results under the empty condition, which are then used respectively. , and , The two prediction results are scaled by mean-standard deviation, and a classifier-free guided method is used to linearly combine the prediction results under the two different conditions. The result is... Using the Euler discrete sampling inversion method, through and Obtain noise ,use , right Scaling by mean-standard deviation yields the noise. Invert the number y by 1;

[0010] Step 4: Determine if the inversion number y is less than or equal to m. If it is, repeat step 3; otherwise, obtain... Proceed to step 5;

[0011] Step 5, the noise obtained from steps 1-4 With optimized noise This involves initial noise - optimized noise pairs, filtering the collected noise pairs, and calculating the initial random noise for each pair. and optimize noise The generated results are compared with the corresponding real images using evaluation metrics scores, and the initial noise-optimized noise pairs with better scores are retained.

[0012] Preferably, in step 1, the given conditions are the input conditions specified by the diffusion model, including images or text.

[0013] Preferably, in step 1, the specific formula for the classifier-free guided method is:

[0014]

[0015] In the formula, Indicates a time step. Indicates the scaling factor. This indicates the given conditions for the input. This represents an empty condition corresponding to the given condition. This represents all other input conditions besides the given conditions. In the text-based diffusion model, this term does not exist; in the new perspective synthesis model, this term is a camera pose condition. and These correspond to the predicted noise output by the diffusion model under different cues.

[0016] Preferably, in step 1, the specific formula for the Euler discrete sampling inference method is as follows:

[0017]

[0018] In the formula, and These are preset parameters. This represents the noise predicted by the diffusion model. Indicates a time step.

[0019] Preferably, in step 3, the specific formula for the Euler discrete sampling inversion method is as follows:

[0020]

[0021] In the formula, and These are preset parameters. This represents the noise predicted by the diffusion model. Indicates a time step.

[0022] Preferably, in step 3, the specific process of scaling based on the mean and standard deviation is as follows:

[0023] Step 3.1, calculate the current noise. The mean and standard deviation of are denoted as . , ;

[0024] Step 3.2, according to , and corresponding , The noise is scaled using the following formula:

[0025] .

[0026] Preferably, in step 3, the scaling factor of the classifier-guided method is not used. It is always less than the scaling factor of the classifier-free guided method in step 1.

[0027] Preferably, in step 4, the evaluation metric score is a learnable perceptual image patch similarity score. The diffusion model can solve various tasks, such as generating images and videos from given text, and generating images from a new perspective from given reference images. The learnable perceptual image patch similarity score is an evaluation metric suitable for new perspective synthesis tasks. In image generation tasks, the image-text comparison pre-training score can be selected as the evaluation metric, and the initial noise-optimized noise pair with the better score can be retained.

[0028] Beneficial effects:

[0029] Compared with existing technologies, the noise inversion calculation method of the diffusion model based on Euler discrete sampling of the present invention has the following advantages:

[0030] The optimized noise obtained through Euler discrete sampling technique can add more details to the generated results of the diffusion model, thereby improving the quality of the generated results;

[0031] This method eliminates the need for fine-tuning the diffusion model, significantly improving efficiency.

[0032] With strong generalization ability, this method can be applied to a variety of diffusion models that use Euler discrete sampling inference and classifier-free guided methods. Attached Figure Description

[0033] Figure 1 This is a flowchart of a noise inversion calculation method for a diffusion model based on Euler discrete sampling according to the present invention;

[0034] Figure 2 The following is a comparison of the noise effect after processing by the method of the present invention: (a) is the initial noise, and (b) is the noise after processing. Detailed Implementation

[0035] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0036] This invention proposes a noise inversion calculation method for diffusion models based on Euler discrete sampling. An Euler discrete sampling inversion method is designed, in which the difference between the inference and inversion scaling coefficients (without classifier guidance) is used to inject corresponding information into the initial random noise, transforming the initial random noise into optimized noise. This method improves the generation quality of the diffusion model without fine-tuning and can be applied to various diffusion models using Euler discrete sampling inference and classifier-free guidance methods, demonstrating strong generalization ability.

[0037] Tested on the novel perspective synthetic diffusion model Sv3d. This model uses the Euler discrete sampling inference method to predict velocity, and obtains images from 21 different viewpoints through 25 inference steps. The inference process also uses a classifier-free guidance technique. The model requires a single PNG image with a 3D subject and a transparent background as input, in RGBA format, with selectable size, and the optimal size is 576×576 pixels (in this embodiment, a sofa image of the optimal size is selected). The model also requires specifying the azimuth and pitch angles of the camera pose corresponding to the 21 generated results (in this embodiment, the pitch angle is set to 0 for all, and the azimuth angle is set to [15.0, 30.0, 45.0, 60.0, 75.0, 90.0, 108.0, 126.0, 144.0, 162.0, 180.0, 198.0, 216.0, 234.0, 252.0, 270.0, 285.0, 300.0, 315.0, 337.5, 0.0]). In this embodiment, the hyperparameters of the Sv3d model not mentioned in the text are set to default. The specific process is as follows:

[0038] Step 1: Given a time step The initial value of t is T, and in this embodiment, T is 25. The initial random Gaussian noise is generated using a random number seed of 23. The shape is (1, 21, 4, 72, 72), the initial value of the number of inferences x is 1, and the noise is calculated. The mean and standard deviation of are denoted as . , ,right Perform Euler discrete sampling scaling and calculate the scaled initial noise. mean with standard deviation .Will Given the reference image conditions and camera pose conditions, a list of three elements (each element of shape (1, 21) is generated. (Besides the variational autoencoder encoding result (shape (1, 4, 72, 72)), the CLIP model encoding result of the reference image is also given (shape (1, 1, 1024)). This list is fed into the diffusion model network to obtain the prediction result under the given reference image conditions, which has the shape (1, 21, 4, 72, 72). The empty condition and the camera pose condition (the empty condition is two encoded results whose shape is consistent with the given reference image condition, but whose values ​​are all 0) are fed into the diffusion model network to obtain the prediction result under the empty condition, with a shape of (1, 21, 4, 72, 72). The mean and standard deviation of the above two prediction results are calculated and denoted as . , , , Using a classifier-free guided method, the prediction results under two different conditions are linearly combined, and the result is... The shape is (1, 21, 4, 72, 72), and the formula used in the classifier-free guided method is as follows:

[0039]

[0040] In the formula, Indicates a time step. Indicates the scaling factor. This indicates the given conditions for the input. This represents an empty condition corresponding to the given condition. Indicates camera attitude conditions. and These correspond to the predicted noise output by the diffusion model under different cues.

[0041] Then, using the Euler discrete sampling inference method, through and Obtain noise The shape is (1, 21, 4, 72, 72).

[0042]

[0043] In the formula, and These are preset parameters. This represents the noise predicted by the diffusion model. Indicates the time step. Increment the inference count by 1 (x).

[0044] Step 2: Determine if the number of inferences x is less than or equal to m (where m is less than T). If yes, repeat step 1; otherwise, obtain the result. ,Will Considered This embodiment Set it to 16.

[0045] Step 3, given time step The initial value of t is T-m+1, and the initial value of the given inversion number y is 1. For noise... The shape (1, 21, 4, 72, 72) is obtained by Euler discrete sampling scaling. Then use and right Perform mean-standard deviation scaling again to obtain the noise. ,Will Given the reference image conditions and camera pose conditions, a list of three elements (each element of shape (1, 21) is generated. (Besides the variational autoencoder encoding result (shape (1, 4, 72, 72)), the CLIP model encoding result of the reference image is also given (shape (1, 1, 1024)). This list is fed into the diffusion model network to obtain the prediction result under the given reference image conditions, which has the shape (1, 21, 4, 72, 72). The empty condition (where the shape matches the two encoded results of the given reference image condition, but both values ​​are 0) and the camera pose condition are fed into the diffusion model network to obtain the prediction result under the empty condition, with a shape of (1, 21, 4, 72, 72). These are then used to... , , , The two prediction results are then scaled using mean-standard deviation. Using the classifier-free guided method, the same as in step 1, the prediction results under the two different conditions are linearly combined, and the result is... The shape is (1, 21, 4, 72, 72). Using the Euler discrete sampling inversion method, the formula is as follows:

[0046]

[0047] In the formula, and These are preset parameters. This represents the noise predicted by the diffusion model. Indicates a time step;

[0048] pass and Obtain noise ,use , right Scaling by mean-standard deviation yields the noise. The shape is (1, 21, 4, 72, 72), and the inversion number y is increased by 1;

[0049] Specifically, the scaling steps are as follows:

[0050] Step 3.1, calculate the current noise. The mean and standard deviation of are denoted as . , ;

[0051] Step 3.2, according to , and corresponding , The noise is scaled using the following formula:

[0052] .

[0053] It is important to emphasize that in step 1, the scaling factor of the classifier-guided method is not specified. The scaling factor is always greater than the scaling factor of the classifier-free guided method in step 3. In this example, the scaling factor of the classifier-free guided method in step 1 is set to 6.0 to 2.5, and the scaling factor of the classifier-free guided method in step 3 is set to 0.0.

[0054] Step 4: Determine if the inversion number y is less than or equal to m. If it is, repeat step 3; otherwise, obtain... This embodiment Set it to 16.

[0055] Step 5, the noise obtained from steps 1-4 With optimized noise That is, the initial noise-optimized noise pair is selected, and the collected noise pairs are filtered. The learnable perceptual image patch similarity scores between the results generated by the initial random noise and the noise cue and the corresponding real images are calculated respectively. The initial noise-optimized noise pair with the better score is retained.

[0056] The optimized noise and the initial noise were fed into the Sv3d model, and the final partial results are as follows: Figure 2As shown in the figure, the image obtained by optimizing the noise is clearer and has richer details compared to the image obtained with the initial noise.

Claims

1. A method for noise inversion calculation based on a diffusion model using Euler discrete sampling, characterized in that, Includes the following steps: Step 1, given time step t is initially set to T, and random Gaussian noise is generated. The initial value of the number of inferences x is 1, and the noise is calculated. The mean and standard deviation of are denoted as . , ,right Perform Euler discrete sampling scaling and calculate the scaled initial noise. mean with standard deviation ,Will By feeding the given conditions and other conditions into the diffusion model network, the prediction results under the given conditions are obtained; The empty condition, along with other conditions, is fed into the diffusion model network to obtain the prediction results under the empty condition. The mean and standard deviation of the two prediction results are calculated respectively, denoted as . , and , Using a classifier-free guided method, the prediction results under two different conditions are linearly combined, and the result is... Using the Euler discrete sampling inference method, through and Obtain noise Increment the number of inferences x by 1; Step 2: Determine if the number of inferences x is less than or equal to the set number of inferences m, and m is less than T. If so, repeat step 1; otherwise, obtain the result. ,Will Considered Proceed to step 3; Step 3, given time step The initial value of t is T-m+1, and the initial value of the given inversion number y is 1. For noise... Perform Euler discrete sampling scaling to obtain Then use and right Scaling by mean-standard deviation yields the noise. ,Will By feeding the given conditions and other conditions into the diffusion model network, the prediction results under the given conditions are obtained. The empty condition and other conditions are fed into the diffusion model network to obtain the prediction results under the empty condition, which are then used respectively. , and , The two prediction results are scaled by mean-standard deviation, and a classifier-free guided method is used to linearly combine the prediction results under two different conditions, resulting in: Using the Euler discrete sampling inversion method, through and Obtain noise ,use , right Scaling by mean-standard deviation yields the noise. Invert the number y by 1; Step 4: Determine if the inversion number y is less than or equal to m. If it is, repeat step 3; otherwise, obtain... Proceed to step 5; Step 5, the noise obtained from steps 1-4 With optimized noise This involves initial noise - optimized noise pairs, filtering the collected noise pairs, and calculating the initial random noise for each pair. and optimize noise The generated results are compared with the corresponding real images using evaluation metrics scores, and the initial noise-optimized noise pairs with better scores are retained.

2. The method for noise inversion calculation based on Euler discrete sampling in a diffusion model according to claim 1, characterized in that, In step 1, the given conditions are the input conditions specified by the diffusion model, including images or text.

3. The noise inversion calculation method based on Euler discrete sampling for a diffusion model according to claim 1, characterized in that, In step 1, the specific formula for the classifier-free guided method is: , In the formula, Indicates a time step. Indicates the scaling factor. This indicates the given conditions for the input. This represents an empty condition corresponding to the given condition. This represents all other conditions input besides the given condition. and These correspond to the predicted noise output by the diffusion model under different cues.

4. The method for noise inversion calculation based on Euler discrete sampling in a diffusion model according to claim 1, characterized in that, In step 1, the specific formula for the Euler discrete sampling inference method is as follows: , In the formula, and These are preset parameters. This represents the noise predicted by the diffusion model. Indicates a time step.

5. The method for noise inversion calculation based on a diffusion model using Euler discrete sampling according to claim 1, characterized in that, In step 3, the specific formula for the Euler discrete sampling inversion method is as follows: , In the formula, and These are preset parameters. This represents the noise predicted by the diffusion model. Indicates a time step.

6. The method for noise inversion calculation based on Euler discrete sampling in a diffusion model according to claim 1, characterized in that, In step 3, the specific process of scaling based on the mean and standard deviation is as follows: Step 3.1, calculate the current noise. The mean and standard deviation of are denoted as . , ; Step 3.2, according to , and corresponding , The noise is scaled using the following formula: 。 7. The method for noise inversion calculation based on a diffusion model using Euler discrete sampling according to claim 1, characterized in that, In step 3, the scaling factor of the classifier-free guided method. It is always less than the scaling factor of the classifier-free guided method in step 1.

8. The method for noise inversion calculation based on Euler discrete sampling in a diffusion model according to claim 1, characterized in that, In step 4, the evaluation index score is the learnable perceptual image patch similarity score.