Adaptive model parameter quantification method and device based on Taylor expansion approximation
Through grouping time step and hybrid precision quantization, combined with Taylor Expanded Approximation and Lagrangian optimization, the quantization problem of the diffusion model on resource-constrained devices is solved, and efficient and accurate image generation is achieved, suitable for resource-constrained devices such as mobile phones and autonomous vehicles.
Patent Information
- Application Number
- CN202510445203.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
The existing diffusion model quantization methods have time step sensitivity, quantization distortion problems and lack of hybrid precision quantization on resource-constrained devices, resulting in reduced generated image quality and waste of resources.
By grouping time steps and quantizing with hybrid precision, combining Taylor expansion approximation and Lagrangian optimization, the bit allocation of the diffusion model is optimized, and the optimal bit allocation of each layer is found using MSE and SSIM as distortion metrics, and the model is efficiently deployed on resource-constrained devices.
It significantly reduces the computational complexity and memory usage of the model, improves the quality and compression ratio of generated images, maintains efficient inference efficiency and accuracy, and is suitable for resource-constrained devices such as mobile phones and autonomous vehicles.
Smart Images

Figure CN120339513A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of diffusion model optimization, and particularly to an adaptive model parameter quantization method and device based on Taylor expansion approximation. Background Art
[0002] Diffusion models perform excellently in image and video generation tasks, but their high computational complexity and memory requirements limit their application on resource-constrained devices (such as mobile phones, drones, smart watches, etc.).
[0003] The quantization method for diffusion models is Post-Training Quantization (PTQ). Existing PTQ techniques have made significant progress in fields such as Convolutional Neural Networks (CNNs), but there are still some deficiencies and defects in their application to diffusion models. One is: time step sensitivity; the output distributions of diffusion models vary greatly at different time steps. Existing PTQ methods usually adopt a unified quantization strategy and cannot effectively handle this distribution change, resulting in poor performance of the quantized model at some time steps and a decrease in the quality of the generated images. The second is: quantization distortion problem; existing PTQ methods are mainly based on simple distortion metrics such as Minimum Mean Square Error (MMSE) and cannot comprehensively capture the impact of quantization on the quality of the generated images. Especially in low-bit quantization (such as 4-bit, 5-bit), the fidelity and diversity of the generated images decrease significantly. The third is: lack of mixed-precision quantization; existing PTQ methods usually use a unified bit width for quantization and cannot flexibly allocate bits according to the sensitivity of different layers, resulting in resource waste or performance degradation. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above-mentioned defects in the prior art, and provide an adaptive model parameter quantization method and device based on Taylor expansion approximation. By optimizing the quantization technology of bit allocation in the diffusion model, the compression ratio is improved, while maintaining high inference efficiency and accuracy, so as to achieve efficient deployment on resource-constrained devices, reduce memory and computational overhead, and enhance real-time inference capabilities.
[0005] To achieve the above purpose, the present invention provides an adaptive model parameter quantization method based on Taylor expansion approximation, including the following steps;
[0006] Step S1: Adopt a Denoising Diffusion Probabilistic Model (DDPM) or a Diffusion Model (DDIM) as the base model; the base model is trained through a unified time step, and a noisy image is generated at each time step, and the final image is gradually denoised.
[0007] Step S2: Group the time steps. Divide the unified time steps into multiple groups, and after grouping, perform independent quantization configuration optimization for each group.
[0008] Step S3: For each group of time steps, model the quantization problem as a rate-distortion optimization problem. The goal is to minimize the output distortion caused by quantization under the given model size constraint.
[0009] Step S4: Mixed-precision quantization; adopt mixed-precision quantization for the weights and activation values of different layers.
[0010] Step S5: Taylor approximation; utilize the first-order Taylor expansion approximation to decompose the output distortion into the sum of quantization distortions of each layer, so as to find the optimal bit allocation for each layer within linear time complexity.
[0011] Step S6: Lagrangian optimization; adopt the Lagrange multiplier method to transform the rate-distortion optimization problem into the minimization problem of the Lagrangian cost function, ensuring the minimum output distortion under the model size constraint.
[0012] Preferably, in step S1, the base model is trained through 1000 time steps. Each time step generates a noisy image, and the final image is gradually denoised.
[0013] Preferably, in step S2, the 1000 time steps are divided into 10 groups, with 100 time steps in each group. The output distributions of different time steps vary greatly. After grouping, independent quantization configuration optimization is performed for each group.
[0014] Preferably, in step S3, the distortion is measured by a combination of the mean squared error (MSE) and the structural similarity index (SSIM); the mean squared error (MSE) and the structural similarity index (SSIM) are introduced as distortion metrics to comprehensively evaluate the impact of quantization on the quality of the generated image; by combining these two metrics, the fidelity and structural similarity of the image can be better maintained under low-bit quantization.
[0015] Preferably, in step S4, for the layers sensitive to quantization, a higher bit width is allocated, while for the insensitive layers, a lower bit width is allocated.
[0016] Preferably, in step S6, the optimal bit allocation for each layer is found by enumerating different slopes, ensuring the minimum output distortion under the model size constraint.
[0017] Preferably, step S6 further includes the following steps:
[0018] Step S61: Quantize the weights and activation values of the model to generate a rate-distortion curve; each curve corresponds to the distortion of the weights or activation values of one layer at different bit widths.
[0019] Step S62: Enumerate different slopes and find the points on each rate-distortion curve where the slope is equal to this value. These points correspond to the optimal bit allocation for each layer.
[0020] Step S63: Quantize the weights and activation values of the model according to the optimal bit allocation to generate a quantized model.
[0021] Step S64: Perform inference on the quantized model to generate an image, and evaluate the quality of the generated image through metrics such as IS (Inception Score) and FID (Fréchet Inception Distance).
[0022] The present invention also provides an adaptive model parameter quantization device based on Taylor expansion approximation, including the above-mentioned adaptive model parameter quantization method based on Taylor expansion approximation.
[0023] It further includes:
[0024] Input module: Receive an input image or a noise image as the initial input of the diffusion model.
[0025] Quantization module: The quantization module is connected to the input module and quantizes the weights and activation values of the model to reduce the computational complexity and memory occupancy.
[0026] Optimization module: The optimization module is connected to the quantization module and finds the optimal bit allocation for each layer through rate-distortion optimization and Taylor approximation to ensure that the output distortion of the quantized model is minimized under a given size constraint.
[0027] Inference module: The inference module is connected to the optimization module and performs inference on the quantized model to gradually denoise and generate the final image.
[0028] Output module: The output module is connected to the inference module, outputs the generated image, and evaluates the image quality.
[0029] The present invention also provides an electronic terminal device, including the above-mentioned adaptive model parameter quantization device based on Taylor expansion approximation. The electronic device is either connected to an information display device, and this information display device is used to display the denoised image with the display parameters, attributes set by the user or through an artificial intelligence model.
[0030] Preferably, the electronic terminal device is a resource-constrained terminal device; the computing resources and storage resources of the terminal device are relatively limited, and the terminal device is a mobile phone, a robot or an autonomous vehicle.
[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0032] By optimizing the quantization configuration of the time step, the present invention minimizes the quantization distortion, thereby significantly reducing the computational complexity and memory occupancy of the model while maintaining the quality of the generated images. Specifically, it includes: First, time step grouping optimization: The time steps are divided into different groups, and the quantization configuration of each group is optimized respectively to cope with the output distribution differences of different time steps; Second, distortion minimization: By introducing a new distortion metric (MSE+SSIM), the impact of quantization on the quality of the generated images is comprehensively evaluated to ensure high fidelity even under low-bit quantization; Third, mixed-precision quantization: According to the sensitivity of different layers, the bit width is flexibly allocated to achieve the optimal utilization of resources; Fourth, efficient optimization algorithm: Using Taylor series expansion and first-order approximation, an optimization algorithm with linear time complexity is proposed to quickly find the optimal bit allocation scheme; thereby improving the compression ratio, while maintaining high inference efficiency and accuracy, realizing efficient deployment on resource-constrained devices, reducing memory and computational overhead, and enhancing real-time inference capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0034] Figure 1 It is a schematic diagram of the steps of an adaptive model parameter quantization method based on Taylor expansion approximation provided by the present invention;
[0035] Figure 2 It is a schematic diagram of an adaptive model parameter quantization device based on Taylor expansion approximation provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are one embodiment of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0037] Example 1
[0038] Please refer to Figure 1 and Figure 2 , the present invention provides an adaptive model parameter quantization method based on Taylor expansion approximation.
[0039] It includes the following steps;
[0040] Step S1: Adopt the Denoising Diffusion Probabilistic Models (DDPM) or the Denoising Diffusion Implicit Models (DDIM) as the basic model; the basic model is trained through 1000 time steps, and a noisy image is generated at each time step, and the final image is gradually denoised to generate.
[0041] Step S2: Group the time steps. Divide 1000 time steps into 10 groups, with 100 time steps in each group. The output distributions of different time steps vary greatly. After grouping, independent quantization configuration optimization is performed for each group.
[0042] Step S3: For each time step group, model the quantization problem as a rate-distortion optimization problem, and the goal is to minimize the output distortion caused by quantization under the given model size constraint;
[0043] In the step S3, the distortion is measured by a combination of the mean squared error (MSE) and the structural similarity index (SSIM); the mean squared error (MSE) and the structural similarity index (SSIM) are introduced as distortion metrics to comprehensively evaluate the impact of quantization on the quality of the generated image; by combining these two metrics, the fidelity and structural similarity of the image can be better maintained under low-bit quantization.
[0044] Step S4: Mixed-precision quantization; adopt mixed-precision quantization for the weights and activation values of different layers;
[0045] In the step S4, for the layers sensitive to quantization, a higher bit width (such as 8 bits) is allocated, while for the insensitive layers, a lower bit width (such as 4 bits) is allocated.
[0046] Step S5: Taylor approximation; use the first-order Taylor expansion approximation to decompose the output distortion into the sum of the quantization distortions of each layer. In this way, the optimal bit allocation for each layer can be found within the linear time complexity.
[0047] Step S6: Lagrangian optimization; using the Lagrange multiplier method, the rate-distortion optimization problem is transformed into the minimization problem of the Lagrangian cost function to ensure the minimum output distortion under the model size constraint. In step S6, by enumerating different slopes, the optimal bit allocation for each layer is found to ensure the minimum output distortion under the model size constraint.
[0048] Step S6 further includes the following steps:
[0049] Step S61: Quantize the weights and activation values of the model to generate rate-distortion curves; each curve corresponds to the distortion of the weights or activation values of one layer under different bit widths;
[0050] Step S62: Enumerate different slopes and find the points on each rate-distortion curve where the slope is equal to this value. These points correspond to the optimal bit allocation for each layer;
[0051] Step S63: Quantize the weights and activation values of the model according to the optimal bit allocation to generate a quantized model;
[0052] Step S64: Perform inference on the quantized model to generate images, and evaluate the quality of the generated images through IS (Inception Score) and FID (Fréchet Inception Distance) metrics.
[0053] The present invention mainly performs approximation and solution based on Taylor expansion. Specifically, the present invention first defines the problem of neural network parameter quantization as a mathematical optimization problem of minimizing the output error, and the object of optimization is the number of bits by which each layer of parameters is quantized. Subsequently, by using Taylor expansion approximation, the defined mathematical optimization problem is simplified, and an efficient dynamic programming algorithm is used to solve this problem.
[0054] 1. Time step grouping optimization:
[0055] Divide the time steps of the diffusion model into different groups, and optimize the quantization configuration of each group respectively to cope with the output distribution differences of different time steps. Determine the best grouping strategy through experiments to ensure a relatively high distribution similarity within each group. For example, if the total number of time steps is 1000, it can be evenly divided into 10 groups, namely 1 - 100, 101 - 200, 203 - 300, …, 901 - 1000. Experimental data shows that the more groups are divided, the better the compression performance.
[0056] Meaning of time step: In the diffusion model, the time step refers to different time points or steps during the denoising process. Specifically: The diffusion model generates images by gradually denoising, starting from completely noisy data and gradually removing the noise until a clear image is generated. Each time step corresponds to a specific stage of the denoising process. For example, time step 100 represents an early stage of the denoising process, while time step 1000 may represent a stage close to the end of the denoising process. At different time steps, the distribution of the model's activation values and the performance of quantization distortion will vary, so specific quantization strategies need to be designed for each time step.
[0057] Grouping the time steps and optimizing the quantization configuration for each group separately effectively solves the problem of output distribution differences of the diffusion model at different time steps. Experiments show that this method significantly outperforms existing methods in terms of FID (Frechet Inception Distance) and IS (Inception Score) of the generated images under low-bit quantization (such as 5-bit, 6-bit).
[0058] As shown in the following table, when testing with DDPM (Denoising Diffusion Probabilistic Models) on the CIFAR-10 dataset, the FID score of dividing 1000 time steps into 10 groups (the right column, that is, every 100 time steps as a group) is lower than that of dividing them into 2 groups (the left column, that is, every 500 time steps as a group).
[0059]
[0060] Comparison of quantization results (FID) under different numbers of groups
[0061] Note: FID (Frechet Inception Distance) is an index used to evaluate the quality of images generated by generative models (such as GAN, diffusion models, etc.). It measures the performance of the generative model by comparing the distribution differences between the generated images and the real images in the feature space. It has the advantages of capturing high-level semantics, being sensitive to distributions, and high stability, and is one of the most commonly used evaluation indexes in the field of generative models. The lower the FID value, the higher the quality of the generated images.
[0062] 2. New distortion metric (MSE + SSIM)
[0063] Introduce MSE (Mean Squared Error) and SSIM (Structural Similarity) as distortion metrics to comprehensively evaluate the impact of quantization on the quality of generated images. By combining these two metrics, it is possible to better maintain the fidelity and structural similarity of images under low-bit quantization. The formula is as follows:
[0064]
[0065] 0 represents the output of the original model, represents the output of the quantized model
[0066] μ O and are the pixel means of O and respectively.
[0067] and are the pixel variances of O and respectively.
[0068] The pixel covariance of O and respectively.
[0069] c1 and c2: Constants used to stabilize the denominator and prevent division-by-zero errors.
[0070] As shown in the following table, when testing on the CIFAR-10 dataset using DDPM and using MSE+SSIM as the distortion metric, the FID after model quantization is less than that of the quantization model using MSE as the distortion metric when the bit width is 4-7 bits; at 8-bit width, the performance is similar. That is, in most cases, using MSE+SSIM as the distortion metric to optimize the quantization parameters of the diffusion model has more advantages.
[0071]
[0072] Experiments show that using MSE+SSIM as the distortion metric, the FID and IS of the generated images are significantly better than the method using only MSE under low-bit quantization.
[0073] 3. Mixed-precision quantization
[0074] According to the sensitivity of different layers of the neural network model, flexibly allocate the bit width to achieve the optimal utilization of resources. Determine the optimal bit width for each layer through experiments to ensure that while maintaining the quality of the generated images, the computational complexity and memory occupancy of the model are significantly reduced.
[0075] Theoretical derivation shows that in the process of quantizing the weight parameters and activation parameters of the diffusion model, the quantization optimal solution based on the rate-distortion optimization problem can be obtained only when the output error of each layer of the model is equal to the ratio of the quantized parameter size.
[0076] Based on this theory, the process of this part is as follows:
[0077] Quantize the weights and activation parameters of each layer with different bit widths (e.g., from 1 bit to 8 bits), calculate the output distortion (MSE + SSIM) caused by quantization, and generate the rate-distortion curves of the weights and activation degrees of each layer. Use a value, e.g., -1, to select the point on each curve where the slope is equal to this value. The points selected on all curves correspond to a set of bit allocation schemes. Traverse multiple values, e.g., (-0.5, -1, -1.5, -2), until a set of schemes that meet the size constraints of the quantization network is found. The abscissa of each point in this scheme is the optimal quantization bit width value corresponding to that layer. This set of schemes is also the quantization bit width scheme for this time step group.
[0078] The pseudo-code of the algorithm is as follows:
[0079] Generate the quantization rate-distortion curves of the weights and activation parameters
[0080]
[0081] Quantization parameter optimization algorithm based on the Lagrange formula
[0082]
[0083] Experiments show that mixed-precision quantization significantly reduces the memory footprint of the model while maintaining the quality of the generated images.
[0084] 4. Efficient optimization algorithm
[0085] Using Taylor series expansion and first-order approximation, an optimization algorithm with linear time complexity is proposed to quickly find the optimal bit allocation scheme. Experiments show that when optimizing the bit allocation, the computational complexity of this algorithm increases linearly and can be quickly deployed in practical applications.
[0086] Computational complexity optimization
[0087] Assume that N represents the number of rate-distortion curves of the diffusion model, that is, the number of layers of the model weight parameters plus the number of layers of the activation parameters; M represents the number of points on each curve. For example, if the bit width range is 4 - 8 bits, then M is 5; K is the total number of slopes to be evaluated. For example, if the slope values are taken as (-0.5, -1, -1.5, -2), then K is 4. The time complexity of solving the joint optimization problem is O(K * M * N). This shows that this scheme has linear time complexity and is independent of the number of parameters. Reduction of computational complexity: Through quantization, the weights and activation values of the model are reduced from 32-bit floating-point numbers to 8 bits or lower, significantly reducing the computational complexity and memory footprint. Image quality maintenance: Through rate-distortion optimization and mixed-precision quantization, the quantized model can still maintain high image generation quality at low bit widths, and the FID and IS metrics are close to those of the full-precision model.
[0088] Overall performance benefits
[0089] As shown in the following table, performance comparison was conducted on four commonly used image generation datasets (CIFAR-10, CelebA-HQ, LSUN-Bedroom, LSUN-Church) using the original unquantified diffusion model DDPM (FP represents floating point), four different quantization methods including Q-Diffusion (Quantized Diffusion), PTQ4DM (Post-Training Quantization for Diffusion Models), PTQD (Post-Training Quantization for Diffusion), APQ-DM (Adaptive Post-Training Quantization for Diffusion Models), and the quantization method our proposed in this technology.
[0090] Two metrics, IS (Inception Score) and FID (Fréchet Inception Distance), were used to evaluate the quality of the generated images. FID measures the distribution difference between the generated images and the real images in the feature space, and a lower value indicates higher quality of the generated images; IS measures the diversity and clarity of the generated images, and a higher value indicates higher quality of the generated images.
[0091] W8A8 represents that the weight parameters and activation parameters are quantized using 8-bit width, and W6A6 represents that the weight parameters and activation parameters are quantized using 6-bit width.
[0092] Experiments show that the performance of this technology on multiple datasets is better than other quantization methods.
[0093]
[0094] Example 2
[0095] Please refer to Figure 1 and Figure 2 , the present invention provides an adaptive model parameter quantization device based on Taylor expansion approximation. The device includes an adaptive model parameter quantization method based on Taylor expansion approximation described in Example 1.
[0096] The device further includes:
[0097] Input module: Receives an input image or a noise image as the initial input of the diffusion model.
[0098] Quantization module: The quantization module is connected to the input module to quantize the weights and activation values of the model, reducing the computational complexity and memory occupation.
[0099] Optimization module: The optimization module is connected to the quantization module. By rate-distortion optimization and Taylor approximation, it finds the optimal bit allocation for each layer to ensure that the output distortion of the quantized model is minimized under a given size constraint.
[0100] Inference module: The inference module is connected to the optimization module and performs inference on the quantized model to gradually denoise and generate the final image.
[0101] Output module: The output module is connected to the inference module, outputs the generated image, and evaluates the image quality.
[0102] Embodiment III
[0103] Please refer to Figure 1 and Figure 2 , the present invention provides an electronic terminal device, including an adaptive model parameter quantization device based on Taylor expansion approximation described in Embodiment II. This electronic device is either connected to an information display device, which is used to display the denoised image with display parameters, attributes set by the user, or through an artificial intelligence model.
[0104] The electronic terminal device is a resource-constrained terminal device; the computing resources and storage resources of the terminal device are relatively limited, and the terminal device is a mobile phone, a robot, or an autonomous driving vehicle.
[0105] By quantizing the model using the method of the present invention, the size of the model can be effectively reduced, the inference efficiency can be improved, and then the model can be deployed to resource-constrained terminal devices.
[0106] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An adaptive model parameter quantization method based on Taylor expansion approximation, characterized in that: Including the following steps; Step S1: Adopt a denoising diffusion probabilistic model (DDPM) or a diffusion model (DDIM) as the base model; the base model is trained with a unified time step, and a noisy image is generated at each time step, and the final image is gradually denoised; Step S2: Group the time steps, divide the unified time steps into multiple groups, and perform independent quantization configuration optimization for each group after grouping; Step S3: For each group of time steps, model the quantization problem as a rate-distortion optimization problem, with the goal of minimizing the output distortion caused by quantization under the given model size constraint; Step S4: Mixed-precision quantization; adopt mixed-precision quantization for the weights and activation values of different layers; Step S5: Taylor approximation; Using the first-order Taylor expansion approximation, decompose the output distortion into the sum of quantization distortions of each layer, so as to find the optimal bit allocation for each layer within the linear time complexity; Step S6: Lagrangian optimization; adopt the Lagrangian multiplier method to transform the rate-distortion optimization problem into the minimization problem of the Lagrangian cost function, ensuring the minimum output distortion under the model size constraint.
2. The adaptive model parameter quantization method based on Taylor expansion approximation according to claim 1, wherein: In the said step S1, the base model is trained with 1000 time steps, and a noisy image is generated at each time step, and the final image is gradually denoised.
3. The adaptive model parameter quantization method based on Taylor expansion approximation according to claim 2, wherein: In the said step S2, the 1000 time steps are divided into 10 groups, with 100 time steps in each group. The output distributions of different time steps vary greatly. After grouping, independent quantization configuration optimization is performed for each group.
4. An adaptive model parameter quantization method based on Taylor expansion approximation according to claim 1, characterized in that: In the said step S3, the distortion is measured by a combination of the mean squared error (MSE) and the structural similarity index (SSIM); the mean squared error (MSE) and the structural similarity index (SSIM) are introduced as distortion metrics to comprehensively evaluate the impact of quantization on the quality of the generated image; by combining these two metrics, the fidelity and structural similarity of the image can be better maintained under low-bit quantization.
5. An adaptive model parameter quantization method based on Taylor expansion approximation according to claim 1, characterized in that: In the said step S4, for the layers sensitive to quantization, a higher bit width is allocated, while for the insensitive layers, a lower bit width is allocated.
6. The adaptive model parameter quantization method based on Taylor expansion approximation according to claim 1, characterized in that: In the said step S6, by enumerating different slopes, the optimal bit allocation for each layer is found to ensure the minimum output distortion under the model size constraint.
7. An adaptive model parameter quantization method based on Taylor expansion approximation according to claim 6, characterized in that: The said step S6 further includes the following steps: Step S61: Quantize the weights and activation values of the model to generate a rate-distortion curve; each curve corresponds to the distortion situation of the weights or activation values of a layer at different bit widths; Step S62: Enumerate different slopes, and find the points on each rate-distortion curve where the slope is equal to this value. These points correspond to the optimal bit allocation for each layer; Step S63: Quantize the weights and activation values of the model according to the optimal bit allocation to generate a quantized model; Step S64: Perform inference on the quantized model to generate an image, and evaluate the quality of the generated image through metrics such as IS (Inception Score) and FID (Fréchet Inception Distance).
8. An adaptive model parameter quantization device based on Taylor expansion approximation, characterized in that: Including an adaptive model parameter quantization method based on Taylor expansion approximation according to any one of claims 1 to 7; Further including: Input module: Receives an input image or a noise image as the initial input to the diffusion model; Quantization module: The quantization module is connected to the input module to quantize the weights and activation values of the model, reducing the computational complexity and memory footprint; Optimization module: The optimization module is connected to the quantization module. By rate-distortion optimization and Taylor approximation, it finds the optimal bit allocation for each layer to ensure that the quantized model outputs the minimum distortion under a given size constraint; Inference module: The inference module is connected to the optimization module and performs inference on the quantized model to gradually denoise and generate the final image; Output module: The output module is connected to the inference module to output the generated image and evaluate the image quality.
9. An electronic terminal device, characterized in that, Comprises an adaptive model parameter quantization device based on Taylor expansion approximation as described in claim 8. The electronic device is either connected to an information display device, which is used to display the denoised image with display parameters, attributes set by the user or through an artificial intelligence model.
10. An electronic terminal device according to claim 9, characterized in that: The electronic terminal device is a resource-constrained terminal device; the computing resources and storage resources of the terminal device are relatively limited, and the terminal device is a mobile phone, a robot or an autonomous vehicle.