Diffusion model updating method and device, computer equipment and storage medium

By determining the inference step count, calculating global errors and updating model parameters in the diffusion model, the problem that the diffusion model cannot maximize its generation ability within limited computing resources and inference time in the prior art is solved, and more efficient image generation in the online image generation business is achieved.

CN120106217APending Publication Date: 2025-06-06SHUXING TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510166826.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

Smart Images

  • Figure CN120106217A_ABST
    Figure CN120106217A_ABST
Patent Text Reader

Abstract

The invention relates to a diffusion model updating method and device, computer equipment and a storage medium. The method comprises the following steps: determining the inference step number of a diffusion model according to the inference time of an image generation service allowable diffusion model; determining the actual output of the diffusion model after the step number reasoning; determining a global error according to the actual output and the ideal output; the model parameters of the diffusion model are updated with the minimum global error as the target, so that an updated diffusion model is obtained, and the updated diffusion model is used for generating an image based on the input data of the image generation service. The diffusion model optimized by the method can exert the generation capability to the greatest extent under the fixed reasoning step number.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a diffusion model updating method, apparatus, computer equipment, storage medium and computer program product. Background Art

[0002] Image generation technology based on diffusion models has been widely used in many fields and business scenarios. However, existing diffusion model acceleration methods only optimize the number of floating-point operations per second (Floating-point Operations Per Second, Flops) required for inference to achieve the same generation effect. Due to the limitation of computing resources, diffusion models often find it difficult to give full play to their maximum generation capabilities. For example, online image generation services and products usually have strict requirements on timeliness, and the allowed inference time of diffusion models is very limited. For a given inference time, based on the computing resources used, we can roughly estimate the available Flops. For this type of online business, the current diffusion model acceleration methods cannot give full play to the generation capabilities of the diffusion model. Summary of the invention

[0003] Based on this, it is necessary to provide a diffusion model updating method, apparatus, computer equipment, computer readable storage medium and computer program product to address the above technical issues.

[0004] In a first aspect, the present application provides a diffusion model updating method, comprising:

[0005] Determine the number of inference steps of the diffusion model according to the inference time of the diffusion model allowed by the image generation business;

[0006] Determining the actual output of the diffusion model at the last iteration point after the number of inference steps;

[0007] Determining a global error based on the actual output and the ideal output;

[0008] Taking the global error minimum as a goal, the model parameters of the diffusion model are updated to obtain an updated diffusion model, and the updated diffusion model is used to generate an image based on the input data of the image generation service.

[0009] In one embodiment, determining the actual output of the diffusion model at the last iteration point after the number of inference steps includes:

[0010] Based on the sampler of the diffusion model, determining the iterative relationship between the outputs of two adjacent iteration points in the diffusion model according to the time function of the diffusion model and the neural network;

[0011] According to the iterative relationship, the actual output of the diffusion model at the last iterative point after the number of inference steps is determined.

[0012] In one embodiment, the updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0013] Taking the global error minimum as the goal, the weights of the neural network are fixed, and the time function is updated to obtain the updated diffusion model.

[0014] In one embodiment, the updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0015] Taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network to obtain the updated diffusion model.

[0016] In one embodiment, the updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0017] Alternately performing step 1 and step 2 to obtain the updated diffusion model;

[0018] The step 1 comprises: fixing the weight of the neural network and updating the time function with the goal of minimizing the global error;

[0019] The step 2 includes: taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network.

[0020] In one embodiment, determining the number of inference steps of the diffusion model according to the inference time allowed by the diffusion model for the image generation service includes:

[0021] Determine the number of floating point operations per second based on the inference time allowed for the diffusion model of the image generation business and the available computing resources;

[0022] A number of inference steps of the diffusion model corresponding to the number of floating-point operations per second is determined.

[0023] In a second aspect, the present application also provides a diffusion model updating device, comprising:

[0024] An inference step determination module, used to determine the inference step number of the diffusion model according to the inference time allowed by the diffusion model for the image generation service;

[0025] A global error determination module is used to determine the actual output of the diffusion model at the last iteration point after the number of inference steps; and determine the global error based on the actual output and the ideal output;

[0026] The model updating module is used to update the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model, wherein the updated diffusion model is used to generate an image based on the input data of the image generation service.

[0027] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method as described in the first aspect or any one of its embodiments is implemented.

[0028] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method as described in the first aspect or any one of its embodiments is implemented.

[0029] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in the first aspect or any one of its embodiments.

[0030] The above diffusion model updating method, device, computer equipment, storage medium and computer program product can determine the number of inference steps of the diffusion model according to the inference time allowed by the image generation business for the diffusion model; and determine the actual output of the diffusion model at the last iteration point after the number of inference steps; then determine the global error according to the actual output and the ideal output; finally, update the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model, and the updated diffusion model is used to generate images based on the input data of the image generation business. Through this scheme, according to the requirements of the image generation business for timeliness, the number of inference steps of the corresponding diffusion model is determined, and based on the fixed number of inference steps, the global error is determined, and the diffusion model is optimized with the goal of minimizing the global error, so that the optimized diffusion model can maximize the generation capability under the fixed number of inference steps, and when the updated diffusion model is used to generate images based on the input data of the image generation business, the generated image effect is better. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0032] Figure 1 A flowchart of a diffusion model updating method;

[0033] Figure 2 A sampling comparison diagram is shown;

[0034] Figure 3 A schematic diagram of the process of image generation using the updated diffusion model;

[0035] Figure 4 A structural block diagram of a diffusion model updating device;

[0036] Figure 5 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0038] The diffusion model updating method provided in the embodiment of the present application can be applied to a diffusion model updating device or a computer device. The diffusion model updating device can be a functional module or functional entity in the computer device for implementing the diffusion model updating method. The computer device can be a server, a computer, etc.

[0039] At present, image generation technology based on diffusion models has been widely used in many fields and business scenarios. However, in actual use, due to limited computing resources, diffusion models often cannot exert their maximum generation capabilities. For example, businesses and products such as online image generation have strict timeliness requirements, so the allowed inference time of the diffusion model is limited. The existing diffusion model acceleration method only optimizes the number of floating-point operations per second (Flops) required to achieve the same generation effect, but does not consider how to maximize the generation capacity of the model under a given Flops budget. In addition, the current diffusion model acceleration method requires relevant personnel to conduct a large number of manual parameter adjustment experiments for manual optimization.

[0040] Flops is an important indicator to measure computer performance. Conceptually, it stands for the number of floating-point operations performed per second. Floating-point operations are a type of operation involving numbers with decimal parts, such as addition, subtraction, multiplication, and division. In computer science and engineering, many complex computing tasks, especially those involving scientific computing, graphics processing, artificial intelligence, etc., require a large number of floating-point operations.

[0041] In the image generation business, Flops reflects the computing power of the system in processing the diffusion model reasoning process. Higher Flops means that the system can perform more floating-point operations per unit time, which can speed up the reasoning of the diffusion model and thus generate images faster. At the same time, Flops is also limited by available computing resources, including processor performance, memory capacity, graphics processing unit (GPU) capabilities and other factors. If the image generation business has high timeliness requirements, then a sufficiently high Flops is required to ensure that the diffusion model reasoning is completed within the specified time and an image that meets business needs is generated.

[0042] In an embodiment of the present application, the available Flops within the inference time allowed by the diffusion model can be used as a constraint condition, and a constrained optimization problem is constructed on this basis as the goal of fine-tuning the diffusion model, so that the generation capacity of the model can be maximized under given computing resources.

[0043] In an exemplary embodiment, Figure 1 As shown, a flow chart of a diffusion model updating method is provided, and the method may include but is not limited to the following steps 101 to 104.

[0044] 101. Determine the number of inference steps of the diffusion model according to the inference time of the diffusion model allowed by the image generation business.

[0045] The image generation service may be any service that requires image generation. In some embodiments, the image generation service may be an online image generation service, which has high requirements for timeliness and allows a diffusion model to have a shorter reasoning time.

[0046] In some embodiments, the available floating point operations may be determined based on the inference time allowed by the diffusion model for the image generation service and the available computing resources, and then the number of inference steps of the diffusion model corresponding to the floating point operations may be determined.

[0047] Among them, this is an important constraint on the diffusion model reasoning time allowed by the business. If the business has high timeliness requirements, the allowed reasoning time is shorter. The available computing resources determine the total amount of computing power that can be invested in diffusion model reasoning. By considering these two factors comprehensively, we can derive the floating-point operations that can be used in a specific business scenario.

[0048] Then, determine the number of inference steps of the diffusion model corresponding to the floating-point operand. The number of inference steps of the diffusion model directly affects the quality and time consumption of image generation. Different floating-point operands correspond to different inference steps to maximize the quality of image generation within the time required by the business. For example, more floating-point operands may allow more inference steps, thereby generating more detailed images, but at the same time, it is also necessary to consider that the inference time cannot exceed the business allowed.

[0049] 102. Determine the actual output of the diffusion model after the number of inference steps.

[0050] The actual output of the diffusion model after the inference steps is the actual output of the diffusion model at the last iteration point after the inference steps.

[0051] The above-mentioned determination of the actual output of the diffusion model after the number of inference steps may include but is not limited to: based on the sampler of the diffusion model, determining the iterative relationship between the outputs of two adjacent iteration points in the diffusion model according to the time function and neural network of the diffusion model; and determining the actual output of the last iteration point of the diffusion model after the number of inference steps according to the iterative relationship.

[0052] The sampler of the diffusion model can be any of the following:

[0053] First-order differential equation solver Euler algorithm, probability-based sampler, second-order or higher-order differential equation solver.

[0054] In some embodiments, when the first-order differential equation solver Euler algorithm is used as a sampler of the diffusion model, first of all, its algorithm principle and calculation process are relatively simple, making it very convenient to implement and deploy. Developers can quickly build a diffusion model sampling system based on the Euler algorithm without complex code logic and a large amount of computing resources. Secondly, it performs well in terms of computational efficiency. The Euler algorithm only needs to perform a simple linear approximation calculation at each time step, and can complete a large number of iterative operations in a very short time. For online image generation services with extremely high timeliness requirements, such as real-time image preview, fast image synthesis and other scenarios, the Euler algorithm can quickly generate images that meet basic requirements within the specified reasoning time, greatly improving the response speed of the service and user experience. In addition, simple calculation methods also mean lower computing costs, reducing the consumption of hardware resources, and reducing the operating costs of the system.

[0055] In some embodiments, when the probability-based sampler is a diffusion model sampler, images with extremely high diversity can be generated. By introducing randomness or adjusting the noise injection method, it can deeply explore the potential image distribution space and generate images with different styles and rich content. In the fields of art creation, advertising design, etc., this diversity can meet the needs of different users for personalization and innovation, and provide more inspiration and choices for designers and creators. At the same time, some probability-based samplers, such as the Denoising Diffusion Implicit Models (DDIM) sampler, significantly improve the sampling speed while ensuring a certain image quality. This enables the probability-based sampler to generate images quickly and with high quality in some business scenarios that have certain requirements for image quality but cannot tolerate too long inference time, such as online image editing tools, real-time special effects generation, etc., and find a balance between efficiency and quality, which can effectively improve the competitiveness of the business and user satisfaction.

[0056] In some embodiments, when a second-order or higher-order differential equation solver is used as a sampler of a diffusion model, it is characterized by high-precision computing power, and through more complex multi-order approximation methods, it can more accurately simulate the state changes in the diffusion process. Compared with the above-mentioned Euler algorithm, under the same number of reasoning steps, the second-order or higher-order differential equation solver can obtain results closer to the true solution, thereby generating images with higher quality and richer details. In fields with strict requirements on image quality, such as medical image generation, high-precision image rendering, etc., this high-precision sampler can provide more reliable and clearer images, providing strong support for professional research and applications. In addition, due to its fast convergence speed, fewer iteration steps are often required to achieve the same error requirements. This not only improves computing efficiency and reduces reasoning time, but also reduces the consumption of computing resources to a certain extent, which has important practical significance for scenes with limited computing resources but pursuing high-quality images.

[0057] In the following description, the Euler algorithm, a first-order differential equation solver, is mainly used as an example for illustrating a sampler of a diffusion model.

[0058] Exemplarily, the above determination of the actual output of the diffusion model after the number of inference steps can use the first-order differential equation solver Euler algorithm as a sampler of the diffusion model, and determine the iterative relationship between the outputs of two adjacent iterative points in the diffusion model based on the time function and neural network of the diffusion model; based on the iterative relationship, determine the actual output of the last iterative point of the diffusion model after the number of inference steps.

[0059] Considering the first-order differential equation solver Euler algorithm as the sampler of the diffusion model, the iteration relationship can be expressed by an iteration formula, which is shown in the following formula (1):

[0060]

[0061] Among them, x i is the output of the diffusion model at the current iteration point, x i+1 is the output of the diffusion model at the next iteration point, σ i is a function of time, D θ (x i ,σ i ) is the output of the neural network at the current iteration point. Since the time function is strictly decreasing, this iteration formula can be rewritten as the following formula (2):

[0062]

[0063] It can be seen from the above formula (2) that in the generation process of the diffusion model, the next iteration point is the current iteration point and the neural network D θ(x i ,σ i ) outputs a convex combination.

[0064] For example, taking the above reasoning step number T as an example, considering a diffusion generation process of T steps, recursively using the above formula (2), we can express the output of the last iteration point as a linear combination of the outputs of each iteration point before the last iteration point in the generation process. Therefore, after the diffusion generation process of T steps, the output of the diffusion model at the last iteration point can be expressed in the form of the following formula (3):

[0065]

[0066] in, is the weighting coefficient, x 0 is the initial noise of the generation process. Since the time function is strictly decreasing, this formula shows that the output x of the last iteration point generated by the diffusion model T , which is essentially a neural network D on the sampling trajectory θ (x i ,σ i ) outputs a convex combination.

[0067] 103. Determine the global error based on the actual output and the ideal output.

[0068] The essence of the generation process of the diffusion model is to solve a complex neural differential equation. Each forward propagation of the neural network in the diffusion model outputs an estimate of the iteration direction at the current iteration point. When the number of forward propagation calculations tends to infinity, the error of the generation process is smaller. Therefore, we can use the error generated by numerical integration as a starting point to analyze the global error of the diffusion model under a given number of inference steps.

[0069] Based on the above formula (3), we can define the global error of the diffusion model generation process as shown in the following formula (4):

[0070]

[0071] in, is the ideal true solution obtained by the integration, that is, the ideal output for image generation.

[0072] Above ∈ represents random noise that satisfies Gaussian distribution, which can be combined with the above formula (4) to obtain the following formula (5):

[0073]

[0074] Based on the above formula (5) combined with the triangle inequality, the following formula (6) can be obtained:

[0075]

[0076] Since ∈ 0 With ∈ T Independent of each other, according to the law of large numbers, |∈ 0 -∈ T |=2d, where d represents the dimension of random noise, i.e. the number of pixels in the generated image. Ignore the fixed constant σ T |∈ 0 -∈ T |, due to the weighting factor λ i It is a convex combination. Combining the above formula (6) and formula (4), using Jensen's inequality, we can get that this global error has an upper bound, which can be expressed as the following formula (7):

[0077]

[0078] It should be noted that the above process of determining the global error is an exemplary description, and any other method can also be used to determine the global error.

[0079] 104. With the goal of minimizing the global error, update the model parameters of the diffusion model to obtain an updated diffusion model.

[0080] The updated diffusion model is used to generate an image based on input data of the image generation service.

[0081] Among them, combined with the above formula (7), it can be seen that in order to make the global error Minimum, then the upper bound of the global error is Minimum.

[0082] Assume that the pre-trained diffusion model approximates the learned denoising function D(x i ,σ i ), the error in the generation process mainly comes from the discretization method when numerically solving neural differential equations, that is, the time function σ i .

[0083] In some embodiments, the weight of the neural network can be fixed and the time function can be updated with the goal of minimizing the global error to obtain an updated diffusion model. θ (x i ,σ i ) weight θ remains unchanged, and the optimization time function σ i It is the optimal discretization method for the current diffusion model to obtain the updated diffusion model.

[0084] In some embodiments, the weights of the neural network can be updated with the global error minimum as the goal and the time function fixed to obtain an updated diffusion model. iKeep unchanged and optimize the above D θ (x i ,σ i ) to obtain the updated diffusion model so that it is in the current time function σ i The loss of the above prediction is the smallest.

[0085] In some embodiments, updating the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0086] Alternately perform step 1 and step 2 to obtain an updated diffusion model;

[0087] Step 1 includes: fixing the weights of the neural network and updating the time function with the goal of minimizing the global error;

[0088] Step 2 includes: taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network.

[0089] Exemplarily, combined with the upper bound of the global error, a two-stage alternating optimization algorithm is constructed: in the first stage, the diffusion model network D can be fixed θ (x i ,σ i ) weight, optimizing the time function σ i is the optimal discretization method for the current diffusion model; in the second stage, the time function σ is fixed i , optimized diffusion model D θ (x i ,σ i ) so that it is at the current time function σ i The loss of the above prediction is the smallest.

[0090] The above diffusion model updating method can maximize the effect of the pre-trained diffusion model for the operation scenario of the online image generation business.

[0091] In the embodiment of the present application, after executing the above diffusion model updating method, the updated diffusion model may be used to generate an image.

[0092] Through mathematical analysis, we explicitly include the computing resource constraints into the optimization objective function, thereby transforming a difficult-to-solve constrained optimization problem into an unconstrained optimization problem, and provide a two-stage alternating algorithm for solving this unconstrained optimization problem, which can be optimized into a fixed-step generation model based on the pre-trained diffusion model. Experimental results show that the diffusion model update method in this application is applicable to any pre-trained diffusion model, and the sampling steps are improved from 5 to 20 steps. In addition, the fine-tuned diffusion model and its time function are still applicable to other accelerated sampling algorithms, and the quality of generated samples can be further improved by combining accelerated sampling algorithms. Among them, the comparison of FID effects under different model types is shown in Table 1 below:

[0093] Table 1

[0094]

[0095] Among them, FID stands for Fréchet Inception Distance. FID is mainly used to evaluate the quality of images generated by generative models (such as generative adversarial networks, diffusion models, etc.) and their similarity to the distribution of real images. Specifically, FID measures the distance between the feature distribution of generated images and the feature distribution of real images. A lower FID value means that the feature distribution of the generated image is closer to the feature distribution of the real image, which means that the generated image is of higher quality, more realistic, and more similar to the statistical characteristics of the real data. Conversely, a higher FID value indicates that the difference between the generated image and the real image is large, and the generated quality is relatively low.

[0096] The above-mentioned accelerated sampling algorithm may include but is not limited to: a denoising probabilistic model solver (Denosing Probabilistic Models SOLVER, DPM SOLVER), or a pluggable neural diffusion model (Pluggable Neural Diffusion Models, PNDM).

[0097] For example, Figure 2 As shown, Figure 2 A sampling comparison diagram. Figure 2 The dashed arrow on the left side of the figure shows the sampling trajectory of the current diffusion model in the high-dimensional manifold space. Since the training goal of the diffusion model is to fit the iteration direction at the current point, that is, the gradient in the probability space, the trained diffusion model will move along the sample manifold according to the gradient direction, and its sampling path is Figure 2The half of the deterministic sampling path shown on the right. After the diffusion model updating method provided in the embodiment of the present application, the iterative trajectory after global error optimization can allow a large deviation from the gradient direction locally, but the local error directions at different time steps are opposite, thereby reducing the global error. The sampling trajectory is as follows Figure 2 As shown by the arrow on the left side of the figure, the sampling path is as follows Figure 2 The optimized sampling path is shown on the right.

[0098] In an exemplary embodiment, Figure 3 As shown, a schematic diagram of a process for generating an image using an updated diffusion model is provided. The method may include but is not limited to the following steps:

[0099] 301. Obtain input data of the image generation service.

[0100] The image generation service may be an online image generation service. The input data may be image data or text data, or may be comprehensive data composed of image data and text data.

[0101] The above text data may be some description information for the image to be generated.

[0102] 302. Input the input data to the updated diffusion model.

[0103] 303. Obtain a generated image output by the updated diffusion model.

[0104] In the embodiment of the present application, after the input data is input into the updated diffusion model, the updated diffusion model can generate an image based on the input data.

[0105] Since the updated diffusion model is optimized based on the number of inference steps of the diffusion model allowed by the image generation service, when the updated diffusion model is used to generate images for the image generation service, the capability of the diffusion model can be maximized within a fixed number of inference steps, thereby meeting the timeliness requirements of the image generation service.

[0106] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0107] Based on the same inventive concept, the embodiment of the present application also provides a diffusion model updating device for implementing the diffusion model updating method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more diffusion model updating device embodiments provided below can refer to the limitations on the diffusion model updating method above, and will not be repeated here.

[0108] In an exemplary embodiment, Figure 4 As shown, a diffusion model updating device is provided, comprising:

[0109] The inference step number determination module 401 is used to determine the inference step number of the diffusion model according to the inference time allowed by the diffusion model for the image generation service;

[0110] A global error determination module 402 is used to determine the actual output of the diffusion model after the number of inference steps; and determine the global error based on the actual output and the ideal output;

[0111] The model updating module 403 is used to update the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model, and the updated diffusion model is used to generate an image based on the input data of the image generation service.

[0112] In some embodiments, the global error determination module 402 is specifically used to: use the first-order differential equation solver Euler algorithm as a sampler of the diffusion model, and determine the iterative relationship between the outputs of two adjacent iteration points in the diffusion model according to the time function and the neural network of the diffusion model;

[0113] According to the iterative relationship, the actual output of the diffusion model at the last iterative point after the number of inference steps is determined.

[0114] In some embodiments, the model updating module 403 is specifically configured to:

[0115] Taking the global error minimum as the goal, the weights of the neural network are fixed, and the time function is updated to obtain the updated diffusion model.

[0116] In some embodiments, the model updating module 403 is specifically configured to:

[0117] Taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network to obtain the updated diffusion model.

[0118] In some embodiments, the model updating module 403 is specifically configured to:

[0119] Alternately performing step 1 and step 2 to obtain the updated diffusion model;

[0120] The step 1 comprises: fixing the weight of the neural network and updating the time function with the goal of minimizing the global error;

[0121] The step 2 includes: taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network.

[0122] In some embodiments, the inference step number determination module 401 is specifically used to:

[0123] Determine the number of floating point operations per second based on the inference time allowed for the diffusion model of the image generation business and the available computing resources;

[0124] A number of inference steps of the diffusion model corresponding to the number of floating-point operations per second is determined.

[0125] Each module in the above diffusion model updating device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0126] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a diffusion model updating method is implemented.

[0127] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0128] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0129] Determine the number of inference steps of the diffusion model according to the inference time of the diffusion model allowed by the image generation business;

[0130] determining an actual output of the diffusion model after the number of inference steps;

[0131] Determining a global error based on the actual output and the ideal output;

[0132] Taking the global error minimum as a goal, the model parameters of the diffusion model are updated to obtain an updated diffusion model, and the updated diffusion model is used to generate an image based on the input data of the image generation service.

[0133] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0134] Determining the actual output of the diffusion model after the number of inference steps includes:

[0135] Using a first-order differential equation solver Euler algorithm as a sampler of a diffusion model, and determining an iterative relationship between outputs of two adjacent iteration points in the diffusion model according to a time function and a neural network of the diffusion model;

[0136] According to the iterative relationship, the actual output of the diffusion model at the last iterative point after the number of inference steps is determined.

[0137] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0138] The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0139] Taking the global error minimum as the goal, the weights of the neural network are fixed, and the time function is updated to obtain the updated diffusion model.

[0140] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0141] The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0142] Taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network to obtain the updated diffusion model.

[0143] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0144] The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0145] Alternately performing step 1 and step 2 to obtain the updated diffusion model;

[0146] The step 1 comprises: fixing the weight of the neural network and updating the time function with the goal of minimizing the global error;

[0147] The step 2 includes: taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network.

[0148] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0149] The step of determining the number of inference steps of the diffusion model according to the inference time allowed by the diffusion model for the image generation service includes:

[0150] Determine the number of floating point operations per second based on the inference time allowed for the diffusion model of the image generation business and the available computing resources;

[0151] A number of inference steps of the diffusion model corresponding to the number of floating-point operations per second is determined.

[0152] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0153] Determine the number of inference steps of the diffusion model according to the inference time of the diffusion model allowed by the image generation business;

[0154] determining an actual output of the diffusion model after the number of inference steps;

[0155] Determining a global error based on the actual output and the ideal output;

[0156] Taking the global error minimum as a goal, the model parameters of the diffusion model are updated to obtain an updated diffusion model, and the updated diffusion model is used to generate an image based on the input data of the image generation service.

[0157] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0158] Determining the actual output of the diffusion model after the number of inference steps includes:

[0159] Using a first-order differential equation solver Euler algorithm as a sampler of a diffusion model, and determining an iterative relationship between outputs of two adjacent iteration points in the diffusion model according to a time function and a neural network of the diffusion model;

[0160] According to the iterative relationship, the actual output of the diffusion model at the last iterative point after the number of inference steps is determined.

[0161] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0162] The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0163] Taking the global error minimum as the goal, the weights of the neural network are fixed, and the time function is updated to obtain the updated diffusion model.

[0164] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0165] The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0166] Taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network to obtain the updated diffusion model.

[0167] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0168] The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0169] Alternately performing step 1 and step 2 to obtain the updated diffusion model;

[0170] The step 1 comprises: fixing the weight of the neural network and updating the time function with the goal of minimizing the global error;

[0171] The step 2 includes: taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network.

[0172] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0173] The step of determining the number of inference steps of the diffusion model according to the inference time allowed by the diffusion model for the image generation service includes:

[0174] Determine the number of floating point operations per second based on the inference time allowed for the diffusion model of the image generation business and the available computing resources;

[0175] A number of inference steps of the diffusion model corresponding to the number of floating-point operations per second is determined.

[0176] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps:

[0177] Determine the number of inference steps of the diffusion model according to the inference time of the diffusion model allowed by the image generation business;

[0178] determining an actual output of the diffusion model after the number of inference steps;

[0179] Determining a global error based on the actual output and the ideal output;

[0180] Taking the global error minimum as a goal, the model parameters of the diffusion model are updated to obtain an updated diffusion model, and the updated diffusion model is used to generate an image based on the input data of the image generation service.

[0181] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0182] Determining the actual output of the diffusion model after the number of inference steps includes:

[0183] Based on the sampler of the diffusion model, determining the iterative relationship between the outputs of two adjacent iteration points in the diffusion model according to the time function of the diffusion model and the neural network;

[0184] According to the iterative relationship, the actual output of the diffusion model at the last iterative point after the number of inference steps is determined.

[0185] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0186] The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0187] Taking the global error minimum as the goal, the weights of the neural network are fixed, and the time function is updated to obtain the updated diffusion model.

[0188] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0189] The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0190] Taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network to obtain the updated diffusion model.

[0191] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0192] The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes:

[0193] Alternately performing step 1 and step 2 to obtain the updated diffusion model;

[0194] The step 1 comprises: fixing the weight of the neural network and updating the time function with the goal of minimizing the global error;

[0195] The step 2 includes: taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network.

[0196] In some embodiments, when the processor executes the computer program, the processor further implements the following steps:

[0197] The step of determining the number of inference steps of the diffusion model according to the inference time allowed by the diffusion model for the image generation service includes:

[0198] Determine the number of floating point operations per second based on the inference time allowed for the diffusion model of the image generation business and the available computing resources;

[0199] A number of inference steps of the diffusion model corresponding to the number of floating-point operations per second is determined.

[0200] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0201] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0202] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0203] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A diffusion model updating method, characterized in that: The method comprises: Determine the number of inference steps of the diffusion model according to the inference time of the diffusion model allowed by the image generation business; determining an actual output of the diffusion model after the number of inference steps; Determining a global error based on the actual output and the ideal output; Taking the global error minimum as a goal, the model parameters of the diffusion model are updated to obtain an updated diffusion model, and the updated diffusion model is used to generate an image based on the input data of the image generation service.

2. The method according to claim 1, characterized in that Determining the actual output of the diffusion model after the number of inference steps includes: Based on the sampler of the diffusion model, determining the iterative relationship between the outputs of two adjacent iteration points in the diffusion model according to the time function of the diffusion model and the neural network; According to the iterative relationship, the actual output of the diffusion model at the last iterative point after the number of inference steps is determined.

3. The method according to claim 2, characterized in that The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes: Taking the global error minimum as the goal, the weights of the neural network are fixed, and the time function is updated to obtain the updated diffusion model.

4. The method according to claim 2, characterized in that: The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes: Taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network to obtain the updated diffusion model.

5. The method according to claim 2, characterized in that: The updating of the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model includes: Alternately performing step 1 and step 2 to obtain the updated diffusion model; The step 1 comprises: fixing the weight of the neural network and updating the time function with the goal of minimizing the global error; The step 2 includes: taking the global error minimum as the goal, fixing the time function, and updating the weights of the neural network.

6. The method according to claim 1, characterized in that The step of determining the number of inference steps of the diffusion model according to the inference time allowed by the diffusion model for the image generation service includes: Determine the number of floating point operations per second based on the inference time allowed for the diffusion model of the image generation business and the available computing resources; A number of inference steps of the diffusion model corresponding to the number of floating-point operations per second is determined.

7. A diffusion model updating device, characterized in that: The device comprises: An inference step determination module, used to determine the inference step number of the diffusion model according to the inference time allowed by the diffusion model for the image generation service; A global error determination module is used to determine the actual output of the diffusion model at the last iteration point after the number of inference steps; and determine the global error based on the actual output and the ideal output; The model updating module is used to update the model parameters of the diffusion model with the goal of minimizing the global error to obtain an updated diffusion model, wherein the updated diffusion model is used to generate an image based on the input data of the image generation service.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.