Diffusion model mixing precision quantification method based on dynamic sensitivity
By dynamically adjusting the weight and activation bit width of the diffusion model, the problem of quantitative error accumulation is solved, and more efficient resource utilization and generation performance is achieved.
Patent Information
- Application Number
- CN202510334094.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-25
AI Technical Summary
The existing hybrid precision quantization method fails to fully consider the dynamic characteristics of the diffusion model in the denoising time step, resulting in the accumulation of quantization errors and reduce the generation quality.
By optimizing the weight of each layer of the diffusion model and the activated candidate bit width set, using quantization parameters and sensitivity metrics, dynamically adjusting the bit width allocation, and using an adaptive bit width configuration strategy to minimize the quantization sensitivity, obtaining a hybrid precision quantization model.
While maintaining the generation quality, reduce hardware resource consumption and improve the generation performance and efficiency of diffusion model quantization.
Smart Images

Figure CN120373369A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of diffusion model mixed-precision quantization methods based on dynamic sensitivity, and in particular to a diffusion model mixed-precision quantization method, device, and electronic device based on dynamic sensitivity. Background Art
[0002] Diffusion models have achieved remarkable results in image generation tasks in recent years. However, since diffusion models typically contain large neural network architectures, they require extremely high computational and storage resources, making it extremely difficult to deploy them in low-latency and memory-constrained environments. In this context, quantization techniques can significantly reduce the computational and storage requirements of diffusion models by mapping high-precision floating-point numerical representations to low-precision fixed-point numerical formats. Considering the huge cost of training diffusion models from scratch, post-training quantization has received extensive attention in related research due to its low overhead and ease of use. The main advantage of post-training quantization is that it does not require large-scale modification or retraining of the pre-trained network, and only a small calibration dataset is needed to complete the quantization process. Existing post-training quantization methods for diffusion models can be roughly divided into two categories: fixed-precision quantization and mixed-precision quantization. Fixed-precision quantization methods use the same bit width for all layers of the model. Prior art one proposed a fixed-precision quantization method to quantize the diffusion model at the block granularity and layer granularity. However, such methods fail to effectively distinguish the layers in the network that are more sensitive to quantization operations, resulting in a decline in the performance of the quantized model. Mixed-precision quantization methods evaluate the importance of each layer of the model and assign different bit widths to layers with different importance levels to achieve a better balance between performance and efficiency. Prior art two uses the signal-to-noise ratio as a metric for quantization sensitivity and selects the optimal bit width for activations at each denoising time step, thus improving the generation quality of the quantized model.
[0003] However, existing mixed-precision quantization methods are usually based on static strategies, that is, a fixed bit width is assigned to each layer of the model, and they do not fully consider the dynamic characteristics of the diffusion model at the denoising time steps. The generation process of the diffusion model is a dynamic process of gradual denoising, and the quantization sensitivity of its layers may change significantly at different time steps, resulting in difficulty for static mixed-precision quantization strategies to adapt to the sensitivity requirements of the model during the denoising process. For example, such a strategy may assign inappropriate bit widths to key layers at certain time steps, leading to the accumulation of quantization errors and a reduction in generation quality. Summary of the Invention
[0004] Embodiments of the present application provide a diffusion model mixed-precision quantization method, device, and electronic device based on dynamic sensitivity, aiming to solve the problem that existing quantization strategies may assign inappropriate bit widths to key layers at certain time steps, resulting in the accumulation of quantization errors and a reduction in generation quality.
[0005] To achieve the above object, an embodiment of the present application further provides a method for mixed-precision quantization of a diffusion model based on dynamic sensitivity, including: for the candidate bit-width sets of weights and activations of each layer of the diffusion model to be quantized, optimizing the quantization parameters under all candidate bit-widths through a reconstruction-based quantization method, where the quantization parameters include a scaling factor term and a zero-point value; calculating the weighted sum of the scaling factors of weights and activations as a quantization sensitivity metric based on the optimized quantization parameters; under given resource constraints, minimizing the quantization sensitivity of each layer at all time steps using an adaptive bit-width configuration strategy to obtain the final activation quantization bit-width configuration and weight quantization bit-width configuration, and quantizing the diffusion model according to the final activation quantization bit-width configuration and weight quantization bit-width configuration to obtain the diffusion model after mixed-precision quantization.
[0006] Optionally, the reconstruction-based quantization method for optimizing the quantization parameters under all candidate bit-widths includes: fixing the quantization parameters of the diffusion model to be quantized; setting a quantization objective function, where the objective function is:
[0007]
[0008] In the formula, min||·|| represents minimizing the expression, represents the output of the full-precision network, and w and x t respectively represent the full-precision weight and the full-precision activation at the denoising time step t, represents the output of the quantized network, b represents the bit-width of the current network, and respectively represent the scaling factor and zero-point value in the weight quantization process when the bit-width is b, and respectively represent the scaling factor and zero-point value in the activation quantization process at time step t when the bit-width is b, and respectively represent the quantized weight after quantization and the quantized activation at the denoising time step t; solving the objective function to obtain the quantization parameters under the candidate bit-widths, and inputting the quantization parameters under the candidate bit-widths into a preset quantization reconstruction formula to screen out the final quantization parameters that approximate the original precision.
[0009] Optionally, the solving the objective function to obtain the quantization parameters under the candidate bit-widths includes: in the weight quantization process, calculating the Frobenius norm loss of the diffusion model to be quantized block by block, and updating the weight quantization parameters of the diffusion model to be quantized with b-bit width through backpropagation; in the activation quantization process, calculating the mean square error loss between the full-precision network and the quantized network of the diffusion model to be quantized by grouping time steps to obtain the activation quantization parameters.
[0010] Optionally, the preset quantization reconstruction formula is:
[0011]
[0012] Among them, x represents the full-precision input value to be quantized, s represents mapping x to the quantized integer range, z represents the zero-point position for adjusting the quantized value, b represents the bit width, round(·) represents rounding the quantization result to convert the floating-point value to the nearest integer, and clamp(·) represents restricting the quantization result within a given range. represents the floating-point value restored from the integer obtained by the quantization operation through the dequantization operation.
[0013] Optionally, there are at least two types of weight bit widths and activation bit widths for each layer of the diffusion model to be quantized, and the expression of the quantization sensitivity metric is:
[0014]
[0015] Among them, λ is a hyperparameter used to balance the influence of weights and activations on quantization sensitivity, and respectively represents the weight scaling factor under the bit width b i below, represents the activation scaling factor under the bit width b j,t below, b i represents the quantized weight bit width, and b j,t represents the quantized activation bit width at time step t, and S is the quantization sensitivity metric of layer l of the quantized diffusion model at time step t.
[0016] Optionally, under the given resource constraints, an adaptive bit-width configuration strategy is used to minimize the quantization sensitivity of each layer at all time steps to obtain the final activation quantization bit-width configuration and weight quantization bit-width configuration, including: according to the noise accumulation characteristics of the denoising process of the diffusion model to be quantized, decomposing the global bit-width budget into each time step according to the exponential decay function; establishing an integer linear programming model with the goal of minimizing the sum of the quantization sensitivities of each layer, solving the integer linear programming model to obtain the final activation quantization bit-width configuration and weight quantization bit-width configuration, where, during the process of solving the integer linear programming model, the activation quantization bit-width configuration is dynamically adjusted at each time step while keeping the weight quantization bit-width fixed.
[0017] Optionally, the expression for decomposing the global bit-width budget into each time step according to the exponential decay function is:
[0018]
[0019] Among them, α is a hyperparameter, T is the total number of time steps for denoising, t represents the time step, and k is an intermediate variable.
[0020] Optionally, an integer linear programming model is established with the goal of minimizing the sum of quantization sensitivities of each layer, including: for each time step t, an optimization objective function is established, and the expression of the optimization objective function is:
[0021]
[0022] where S l (b i ,b j ,t) represents the quantization sensitivity of layer l under the bitwidth configuration (b i ,b j ) at time step t;
[0023] The constraint conditions of the expression of the optimization objective function are:
[0024]
[0025] where BitOps = #MACs × b i ×b j , representing the computational overhead of each layer under a given bitwidth, and #MACs represents the number of multiply-accumulate operations.
[0026] To achieve the above object, the present application also provides a diffusion model mixed-precision quantization method, device, and electronic device based on dynamic sensitivity, including: a quantization parameter optimization module for optimizing the quantization parameters under all candidate bitwidths for the weight and activation configuration candidate bitwidth sets of each layer of the diffusion model to be quantized through a reconstruction-based quantization method, where the quantization parameters include a scaling factor term and a zero point value; a metric calculation module for calculating the weighted sum of the scaling factors of the weights and activations as a quantization sensitivity metric based on the optimized quantization parameters; a quantization model determination module for minimizing the quantization sensitivities of each layer at all time steps using an adaptive bitwidth configuration strategy under given resource constraints to obtain the final activation quantization bitwidth configuration and weight quantization bitwidth configuration, and quantizing the diffusion model according to the final activation quantization bitwidth configuration and weight quantization bitwidth configuration to obtain a diffusion model after mixed-precision quantization.
[0027] To achieve the above object, the present application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the diffusion model mixed-precision quantization method based on dynamic sensitivity provided in the foregoing embodiments.
[0028] A method, apparatus, and electronic device for hybrid precision quantization of a diffusion model based on dynamic sensitivity proposed in an embodiment of the present application. By setting a bitwidth indicator and quantization parameters, the bitwidth indicator is used to represent the combination between the weight bitwidth and the activation bitwidth in each layer, and the quantization parameters include a scaling factor and a zero-point value. Determine the weight bitwidth and the activation bitwidth based on the bitwidth indicator, and optimize the quantization parameters at multiple bitwidths based on minimizing the output difference between the full-precision network and the quantization networks at each preset bitwidth. Obtain the quantization sensitivity of any layer at each time step based on the weighted sum of the scaling factor in the weight quantization process and the scaling factor in the activation quantization process. Under the given resource constraints, minimize the sum of the sensitivities of all layers at all time steps to obtain an adaptive bitwidth configuration strategy, and quantize the diffusion model based on the bitwidth configuration strategy to obtain a hybrid-precision quantization diffusion model. The present application can dynamically adjust the bitwidth allocation of network layers according to the quantization sensitivity of the diffusion model at the denoising time steps, so as to better adapt to the sensitivity changes during the denoising process. This strategy can reduce the consumption of hardware resources while maintaining the generation quality. The present application can reduce the storage and computational requirements of the diffusion model and significantly improve the generation performance and efficiency after quantization of the diffusion model. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a flowchart of a method for hybrid precision quantization of a diffusion model based on dynamic sensitivity provided in an embodiment of the present application;
[0030] Figure 2 is a structural block diagram of an apparatus for hybrid precision quantization of a diffusion model based on dynamic sensitivity provided in another embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are proposed for the readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present application can be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation to the specific implementation manner of the present application. Each embodiment can be combined and cross-referenced with each other without conflict.
[0032] An embodiment of the present application proposes a method for mixed-precision quantization of a diffusion model based on dynamic sensitivity, which is applied to an electronic device. The electronic device can be a terminal or a server. In this embodiment and the following embodiments, the server is taken as an example for illustration. The implementation details of the method for mixed-precision quantization of the diffusion model based on dynamic sensitivity proposed in this embodiment are specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution.
[0033] The specific process of the method for mixed-precision quantization of the diffusion model based on dynamic sensitivity proposed in this embodiment can be as Figure 1 shown, including:
[0034] S101. For the candidate bit-width sets of the weights and activation configurations of each layer of the diffusion model to be quantized, optimize the quantization parameters under all candidate bit-widths through a quantization method based on reconstruction. The quantization parameters include a scaling factor term and a zero-point value.
[0035] It should be noted that activation is the output of a neuron, which is generated by a non-linear transformation of the weighted input by an activation function. The weight is a parameter of the model, which is essentially a value learned from training data and is used to adjust the linear combination relationship of input features.
[0036] In an embodiment of the present application, step S101 may include the following execution process:
[0037] S1011. Fix the quantization parameters of the pre-trained model;
[0038] S1012. Set the objective function of quantization. The objective function is:
[0039]
[0040] In the formula, min||·|| represents minimizing the expression, represents the output of the full-precision network, and w and x t respectively represent the full-precision weight and the full-precision activation at the denoising time step t, represents the output of the quantization network, b represents the bit-width of the current network, and respectively represent the scaling factor and the zero-point value in the weight quantization process when the bit-width is b, and respectively represent the scaling factor and the zero-point value in the activation quantization process when the time step is t and the bit-width is b, and respectively represent the restored weight and the quantized activation at the denoising time step t after quantization;
[0041] S1013. Solve the objective function to obtain the quantization parameters under the candidate bitwidth, and input the quantization parameters under the candidate bitwidth into the preset quantization reconstruction formula to screen out the final quantization parameters that approximate the original accuracy.
[0042] In this step, first, the present application configures a candidate bitwidth set B = {b0, …, b n-1} and a bitwidth indicator (b i , b j ) for each layer of weights and activations of the diffusion model to be quantized, which are used to perform the quantization process. Among them, the candidate bitwidth set contains multiple possible bitwidth options, aiming to provide a flexible adjustment space for subsequent bitwidth allocation, while the bitwidth indicator is used to clearly specify the weight bitwidth and the combination of bitwidths of activations used in the current layer.
[0043] In an embodiment of the present application, step S1013 may include the following execution process:
[0044] S10131. In the weight quantization process, calculate the Frobenius norm loss of the diffusion model to be quantized block by block, and update the weight quantization parameters of the diffusion model to be quantized under the b-bitwidth through backpropagation;
[0045] S10132. In the activation quantization process, calculate the mean square error loss between the full-precision network and the quantized network of the diffusion model to be quantized by grouping time steps to obtain the activation quantization parameters.
[0046] Specifically, for the weight quantization process, the model is regarded as consisting of multiple blocks, and the present application uses the following loss function to optimize block by block:
[0047]
[0048] Among them, represents the Frobenius norm, R k (·) and respectively represent the output of the k-th full-precision block and the output of the quantized block in the diffusion model. For the activation quantization process, to improve the quantization efficiency, the present application divides T time steps into G groups and optimizes the activation quantization parameters within each group, which is applied to adjacent T / G time steps. The optimization objective for each time step is:
[0049]
[0050] Among them, ‖·‖ 2 represents the mean square error, x t represents the input at time step t, ∈ θ (x t , t) and respectively represent the outputs of the full-precision network and the quantized network. By optimizing the activation quantization parameters within each time step group in this application, the optimal results of the quantized diffusion model at multiple time steps can be obtained.
[0051] In one embodiment of this application, the preset quantization reconstruction formula is:
[0052]
[0053] where x represents the full-precision input value to be quantized, s represents mapping x to the integer range after quantization, z represents the zero point position for adjusting the quantized value, b represents the bit width, round(·) represents rounding the quantization result to convert the floating-point value to the nearest integer, clamp(·) represents restricting the quantization result within a given range, represents the floating-point value restored from the integer obtained by the quantization operation through the inverse quantization operation.
[0054] In this application, the uniform quantization method is adopted to convert the full-precision floating-point value into an integer value. The scaling factor is used to map the floating-point input value to a suitable integer range, while the zero point value is used to adjust the reference position of the integer value after quantization. In the quantization process, first, the full-precision input value is divided by the scaling factor to convert it into a floating-point number approximately close to the target integer range, and then a rounding operation is performed on it to obtain an integer representation, so that the difference between each quantization level remains consistent, thereby ensuring the uniform distribution of data. Then, this application restricts the rounded result within the target integer range through the clamping function, thereby ensuring that the data does not exceed the representable bit width range. By participating in the quantization process with appropriate scaling factors and zero point values, this application can effectively compress the floating-point representation into an integer, reduce the computing and storage requirements, and at the same time maintain the accuracy of the original data as much as possible.
[0055] S102, based on the optimized quantization parameters, calculate the weighted sum of the scaling factors of the weights and activations as the quantization sensitivity metric.
[0056] In one embodiment of this application, this application sequentially selects a bit width b from the candidate bit width set B and switches the bit width indicator to (b, b), that is, the weight and activation bit width indicators of all network layers are uniformly set to b. In this way, this application can ensure that the quantization parameters of each layer can be optimized under all possible bit widths, avoiding missing any configurations. Exemplarily, in the optimization process of each bit width, this application fixes the parameters of the full-precision pre-trained model and updates the quantization parameters under the b-bit width through backpropagation. The optimization goal of this application is to minimize the difference between the full-precision network and the output of the quantized network under each specified bit width.
[0057] Among them, there are at least two weight bit widths and activation bit widths for each layer of the diffusion model to be quantized. Based on the weighted sum of the scaling factor in the weight quantization process and the scaling factor in the activation quantization process, the expression for the quantization sensitivity of any layer at each time step is obtained as follows:
[0058]
[0059] Among them, λ is a hyperparameter used to balance the influence of weights and activations on quantization sensitivity. represents the weight scaling factor at bit width b i below. represents the activation scaling factor at time step t with bit width b j,t below, and b i represents the quantized weight bit width, and b j,t represents the quantized activation bit width. S is the quantization sensitivity metric for layer l of the quantized diffusion model at time step t.
[0060] S103. Under the given resource constraints, an adaptive bit width configuration strategy is used to minimize the quantization sensitivity of each layer at all time steps, obtaining the final activation quantization bit width configuration and weight quantization bit width configuration. According to the final activation quantization bit width configuration and weight quantization bit width configuration, the diffusion model is quantized to obtain a diffusion model after mixed-precision quantization.
[0061] The efficient quantization sensitivity calculation method designed in this application can avoid the computational cost brought by the huge bit width combination space caused by the number of network layers and time steps of the diffusion model. To simplify this process, this application uses the weighted sum of the scaling factors in the weight and activation quantization parameters as the quantization sensitivity metric, thus eliminating the need for inference-based evaluation. Each layer of weights and activations in the diffusion model of this application has n possible bit widths, and the bit width indicator (b i , b j ) represents using b i bit width to quantize weights and b j bit width to quantize activations. In actual operation, if the quantization sensitivity of a certain layer at a specific time step t is high, it means that this layer is more sensitive to quantization errors, and the quantization errors will significantly affect the generation performance of the model. Therefore, it is necessary to increase the quantization bit width of this layer. On the contrary, if the sensitivity is low, the corresponding quantization bit width can be reduced. In this way, this application can not only efficiently evaluate the quantization sensitivity of each layer at different time steps, but also avoid the bottleneck of a large amount of calculation required in the traditional inference evaluation process.
[0062] In an embodiment of this application, step S103 may include the following execution process:
[0063] S1031. Decompose the global bitwidth budget into each time step according to the noise accumulation characteristics of the denoising process of the diffusion model to be quantized by an exponential decay function;
[0064] S1032. Establish an integer linear programming model with the goal of minimizing the sum of quantization sensitivities of each layer, solve the integer linear programming model, and obtain the final activation quantization bitwidth configuration and weight quantization bitwidth configuration. During the process of solving the integer linear programming model, dynamically adjust the activation quantization bitwidth configuration at each time step while keeping the weight quantization bitwidth fixed.
[0065] This application proposes a dynamic bitwidth adjustment method based on quantization sensitivity. The goal of this method is to obtain the most suitable dynamic bitwidth allocation strategy for the weights and activations of each layer at different time steps according to the quantization sensitivities obtained in the foregoing steps without relying on additional inference calculations, so as to achieve efficient utilization of computing resources.
[0066] Specifically, first, according to the characteristics of noise accumulation in the denoising process of the diffusion model, this application should allocate different resource constraints to the quantization model in the early and late stages of denoising. In the initial stage of denoising, since the quantization sensitivity of the model is small, the resource allocation can be reduced. In the later stage of denoising, as the quantization sensitivity of the model increases, more resources need to be allocated to ensure the accuracy of the model. For this reason, this application decomposes the total resource constraint C into the resource constraint C at each time step t , and the specific decomposition method is exponential decay adjustment. Among them, the allocation method of resource constraints is carried out through the following resource decomposition formula:
[0067]
[0068] In the formula, α is a hyperparameter, T is the total number of time steps of denoising, and k is an intermediate variable. This application converts the search problem of the optimal allocation method into an integer linear programming problem, and the optimization goal is to minimize the sum of quantization sensitivities of each layer under the resource constraint C at each time step t . For each time step t, the optimization goal can be expressed as:
[0069]
[0070] In the formula, S l (b i , b j , t) represents the quantization sensitivity of layer l under the bitwidth configuration (b i , b j ) at time step t. To ensure that the computing resources at each time step do not exceed the constraint C t , the following constraint conditions also need to be satisfied:
[0071]
[0072] where BitOps = #MACs × b i × b j , representing the computational cost of each layer at a given bitwidth, and #MACs is the number of multiply-accumulate operations. Through integer linear programming, the present application can adaptively adjust the bitwidth configuration of the model according to the resource constraints at each time step. In the actual implementation process, the bitwidth of weight quantization remains fixed at all time steps, while the bitwidth of activation quantization is dynamically adjusted according to the change of time steps. This method can not only achieve optimized resource allocation, but also avoid the overhead of storing and loading multiple model state files, because the activation quantizer only needs to change the scaling factor and zero-point value, with almost no additional storage overhead.
[0073] After solving the optimal bitwidth allocation strategy, the present application can be directly applied to the image generation task instead of the original full-precision diffusion model.
[0074] In summary, the quantization sensitivity evaluation method based on quantization parameters effectively avoids additional inference processes and computational overhead, making the performance evaluation of the quantized model more efficient. At the same time, based on the dynamic change characteristics of quantization sensitivity during the denoising process, the present application can automatically adjust the bitwidth allocation of each layer according to different stages of the model, ensuring that the quantization bitwidth matches the actual requirements, thereby achieving more precise resource allocation and better model performance. After the model pre-trained using a large publicly available dataset is processed by the quantization method proposed in the present application, the generation performance remains almost unchanged, verifying the effectiveness and robustness of the method in practical applications.
[0075] It can be understood that the step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of the present application; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process are all within the protection scope of this application.
[0076] Refer to Figure 2, based on the above embodiments, the present application further provides a diffusion model mixed-precision quantization device based on dynamic sensitivity. The diffusion model mixed-precision quantization device 1000 includes a quantization parameter optimization module 1001, a metric calculation module 1002, and a quantization model determination module 1003. Among them, the quantization parameter optimization module 1001 is used to optimize the quantization parameters under all candidate bit widths for the weight and activation configuration candidate bit width sets of each layer of the diffusion model to be quantized through a reconstruction-based quantization method. The quantization parameters include a scaling factor term and a zero point value; the metric calculation module 1002 is used to calculate the weighted sum of the scaling factors of the weights and activations as a quantization sensitivity metric based on the optimized quantization parameters; the quantization model determination module 1003 is used to minimize the quantization sensitivity of each layer at all time steps using an adaptive bit width configuration strategy under given resource constraints to obtain the final activation quantization bit width configuration and weight quantization bit width configuration, and quantize the diffusion model according to the final activation quantization bit width configuration and weight quantization bit width configuration to obtain the diffusion model after mixed-precision quantization.
[0077] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details and technical effects mentioned in the above embodiments are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.
[0078] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of the present application, units that are not closely related to solving the technical problems proposed by the present application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0079] Another embodiment of the present application proposes an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the diffusion model mixed-precision quantization method based on dynamic sensitivity in the above method embodiments.
[0080] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present application.
Claims
1. A method for mixed-precision quantization of a diffusion model based on dynamic sensitivity, characterized in that Including: For the candidate bit-width sets of the weights and activation configurations of each layer of the diffusion model to be quantized, optimize the quantization parameters under all candidate bit-widths through a reconstruction-based quantization method, where the quantization parameters include a scaling factor term and a zero point value; Based on the optimized quantization parameters, calculate the weighted sum of the scaling factors of the weights and activations as a quantization sensitivity metric; Under given resource constraints, use an adaptive bit-width configuration strategy to minimize the quantization sensitivity of each layer at all time steps, obtain the final activation quantization bit-width configuration and weight quantization bit-width configuration, and quantize the diffusion model according to the final activation quantization bit-width configuration and weight quantization bit-width configuration to obtain a diffusion model with mixed-precision quantization.
2. The method for mixed-precision quantization of a diffusion model based on dynamic sensitivity according to claim 1, wherein The optimization of the quantization parameters under all candidate bit-widths by the reconstruction-based quantization method includes: Fix the quantization parameters of the diffusion model to be quantized; Set the quantization objective function, and the objective function is: where min||·|| represents minimizing the expression, represents the full-precision network output, and w and x t represent the full-precision weight and the full-precision activation at the denoising time step t, respectively, represents the output of the quantization network, b represents the bit width of the current network, and represent the scaling factor and zero-point value in the weight quantization process when the bit width is b, respectively, and represent the scaling factor and zero-point value in the activation quantization process when the bit width is b and the time step is t, respectively, and represent the restored weight after quantization and the quantized activation at the denoising time step t, respectively; Solve the objective function to obtain the quantization parameters under the candidate bit-widths, input the quantization parameters under the candidate bit-widths into the preset quantization reconstruction formula, and screen out the final quantization parameters that approximate the original precision.
3. The method for mixed-precision quantization of a diffusion model based on dynamic sensitivity according to claim 2, wherein The solving of the objective function to obtain the quantization parameters under the candidate bit-widths includes: During the weight quantization process, calculate the Frobenius norm loss of the diffusion model to be quantized block by block, and update the weight quantization parameters of the diffusion model to be quantized with b-bit width through backpropagation; During the activation quantization process, calculate the mean square error loss between the full-precision network and the quantized network of the diffusion model to be quantized by grouping time steps to obtain the activation quantization parameters.
4. The method for mixed-precision quantization of a diffusion model based on dynamic sensitivity according to claim 2, wherein, The preset quantization reconstruction formula is: Among them, x represents the full-precision input value to be quantized, s represents mapping x to the quantized integer range, z represents the zero-point position for adjusting the quantized value, b represents the bit width, round(·) represents rounding the quantization result to convert the floating-point value to the nearest integer, and clamp(·) represents restricting the quantization result within a given range. represents the floating-point value restored from the integer obtained by the quantization operation through the dequantization operation.
5. The method for mixed-precision quantization of a diffusion model based on dynamic sensitivity according to claim 1, wherein The expression of the quantization sensitivity metric is: Among them, λ is a hyperparameter used to balance the influence of weights and activations on quantization sensitivity. denotes the weight scaling factor at bit width b i below, denotes the activation scaling factor at bit width b j,t below, and b i denotes the quantization weight bit width, and b j,t denotes the quantization activation bit width at time step t. S is the quantization sensitivity metric of layer l in the quantization diffusion model at time step t.
6. The method for mixed-precision quantization of a diffusion model based on dynamic sensitivity according to claim 1, wherein, The minimizing the quantization sensitivity of each layer at all time steps by using an adaptive bit-width configuration strategy under given resource constraints to obtain the final activation quantization bit-width configuration and weight quantization bit-width configuration includes: According to the noise accumulation characteristics of the denoising process of the diffusion model to be quantized, decompose the global bit-width budget into each time step according to an exponential decay function; Establish an integer linear programming model with the goal of minimizing the sum of the quantization sensitivities of each layer, solve the integer linear programming model, and obtain the final activation quantization bit-width configuration and weight quantization bit-width configuration; Among them, during the process of solving the integer linear programming model, dynamically adjust the activation quantization bit-width configuration at each time step and keep the weight quantization bit-width unchanged.
7. The method for mixed-precision quantization of a diffusion model based on dynamic sensitivity according to claim 6, wherein The expression for decomposing the global bit-width budget into each time step according to an exponential decay function is: Among them, α is a hyperparameter. T is the total number of time steps for denoising, t represents the time step, and k is an intermediate variable.
8. The method for mixed-precision quantization of a diffusion model based on dynamic sensitivity according to claim 4, wherein Establishing an integer linear programming model with the goal of minimizing the sum of the quantization sensitivities of each layer includes: For each time step t, establish an optimization objective function, and the expression of the optimization objective function is: Among them, S l (b i ,b j ,t) represents the quantization sensitivity of layer l at time step t under the bit-width configuration (b i ,b j ). The constraint conditions of the expression of the optimization objective function are: where BitOps = #MACs × b i × b j , representing the computational cost of each layer at a given bit width, and #MACs represents the number of multiply-accumulate operations.
9. A diffusion model mixed-precision quantization device based on dynamic sensitivity, characterized in that, Including: A quantization parameter optimization module for optimizing the quantization parameters under all candidate bit-widths of the weights and activation configurations of each layer of the diffusion model to be quantized through a reconstruction-based quantization method, where the quantization parameters include a scaling factor term and a zero point value; A metric calculation module, configured to calculate the weighted sum of the scaling factors of the weights and activations as a quantization sensitivity metric based on the optimized quantization parameters; A quantization model determination module, configured to minimize the quantization sensitivity of each layer at all time steps by using an adaptive bit-width configuration strategy under given resource constraints, obtain the final activation quantization bit-width configuration and weight quantization bit-width configuration, and quantize the diffusion model according to the final activation quantization bit-width configuration and weight quantization bit-width configuration to obtain a mixed-precision quantized diffusion model.
10. An electronic device, characterized in that, Comprising: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the dynamic sensitivity-based diffusion model mixed-precision quantization method according to any one of claims 1 to 8.
Citation Information
Cited By
Risk perception neural network quantification method and device for aero-engine health management
CN121920453A
Diffusion model weight calibration method and content generation method based on time sequence sensitivity
CN122087775A