Post-training quantitative data calibration system and method for providing visual service for Internet of Things equipment
By selecting calibration data with smooth pixel distribution and no-return sampling strategy, the weight rounding parameters and activation value scale factor of the quantized model are optimized, which solves the problems of low model performance and large calibration data requirements on IoT devices, and achieves efficient visual services.
Patent Information
- Application Number
- CN202510388608.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
Existing quantitative technologies have problems with low model performance and high calibration data requirements on IoT devices, resulting in high resource consumption and slow response speed.
The calibration data selection and no return sampling strategy with smooth pixel distribution are adopted, combined with block-based quantization loss optimization, and the weight rounding parameters and activation value scale factors of the quantization model are optimized through forward propagation and backpropagation to reduce rounding and reconstruction losses.
Without increasing resource consumption, the performance of the quantitative model is significantly improved, the use of calibration data is reduced, and the efficient operation and high-quality visual services are ensured on the IoT device side.
Smart Images

Figure CN120339759A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of post-training quantization of neural network models, and particularly relates to a post-training quantization data calibration system and method for providing visual services for Internet of Things devices. Background Art
[0002] According to the prediction of International Data Corporation, there will be 41.6 billion Internet of Things devices connected to the Internet by 2025. However, the computing and storage resources of Internet of Things devices are limited, and at the same time, users have high requirements for the response speed of applications. Quantization technology maps high-performance full-precision neural network models trained on the cloud to low bits such as INT4, and deploys these low-bit models to ubiquitous Internet of Things devices, so as to provide high-quality visual services (such as image classification and object detection) with low resource requirements and fast response has become the mainstream, as Figure 2 and Figure 3 shown.
[0003] However, due to the lack of detailed analysis of calibration data, most of the existing related methods have two limitations: (i) The performance of the low-bit model is low. (ii) A large amount of calibration data is required, resulting in a large amount of resource consumption. For example, when the QDrop method quantizes ResNet-18 from FP32 to INT8, the model size can be reduced by 75%, and the performance loss on the ImageNet dataset is controlled within 2%. When the quantization bit width is W2A4, the performance drops by up to 5.8%. In addition, the existing post-training quantization methods use more calibration data to improve performance, but it prolongs the quantization period and increases resource consumption. In addition, its effect shows a trend of diminishing marginal utility, that is, as the number of calibration data increases, the amplitude of model performance improvement gradually decreases, which is mentioned in the studies of QDrop and MRECG. Therefore, in order to ensure the high-performance operation and deployment of the quantization model on the Internet of Things device side, it is necessary to improve its performance and reduce the number of calibration data used, especially in today's era of Internet of Everything.
[0004] Existing research shows that the selection of calibration data affects the performance of the downstream services of the model. And using one-time random sampling as calibration data in PTQ is not an optimal strategy. A dynamic clustering method based on activation statistics is proposed to more accurately select calibration data, but it will significantly increase the time consumption. In the PTQ method, the calibration data is input into the model for forward propagation to calculate the quantization loss. Then, based on this loss, a backpropagation operation is performed to optimize the quantization parameters of each block, thereby reducing the accumulation of quantization errors. It can be seen that the calibration data runs through the entire quantization process, directly affecting the rounding loss and reconstruction loss, and has a greater impact on the performance of the quantization model. Therefore, the present invention focuses on the selection of calibration data. By selecting appropriate calibration data and studying the pixel distribution and quantity of the calibration data, the quantization performance loss and the quantity of the selected calibration data are reduced, providing a more efficient solution for actual deployment. Summary of the Invention
[0005] Aiming at the problems of the existing technology, the present invention provides a post-training quantization data calibration system and method for providing visual services for Internet of Things devices, so as to reduce the performance loss of the quantization model and the usage amount of calibration data, ensure the high-performance operation and deployment of the quantization model at the Internet of Things device side, and provide high-quality visual services with low resource requirements and fast response, such as Figure 5 shown.
[0006] To solve the problems of the existing technology, the present invention adopts the following technical solutions:
[0007] A post-training quantization data calibration system for providing visual services for Internet of Things devices, the data calibration system includes a deep neural network model pre-training module, a first data calibration module, a first quantization module, a second data calibration module and a second quantization module; wherein:
[0008] The first data calibration module collects data samples with smooth pixel distribution from the training dataset as calibration data;
[0009] The first quantization module inputs the calibration data into the pre-training module for forward propagation, and obtains the quantized weight W Q and activation value X Q , and calculates the output of the quantization model;
[0010] The second data calibration module adopts a sampling strategy without replacement to extract batch-size samples from the calibration data, and removes the selected data after each sampling, and iteratively calibrates the quantization model obtained by the first quantization module;
[0011] The second quantization module takes blocks as basic units, and calculates the loss Loss i, and optimize the weight rounding parameter V and activation value scale factor scale of each block of the model based on the batch-size calibration samples x .
[0012] Furthermore, the first quantization module inputs the calibration data into the pre-trained model for forward propagation, and obtains the quantized weights W Q and activation values X Q , including:
[0013] Extract 16 samples from the calibration data and input them into the pre-trained module for forward propagation to optimize four quantization variables, namely scale w , zero_point w , scale x and zero_point x ; where scale w and zero_point w are the quantization parameter scale factor and zero point of the weights, and scale x and zero_point x are the quantization parameter scale factor and zero point of the activation values;
[0014] Calculate the quantization parameter scale factor scale w and zero point zero_point w of the weights according to the following formula:
[0015]
[0016] where: w max is the maximum value of the weights, and w min is the minimum value of the weights; n w is the quantization bit width, which determines the maximum value and minimum value
[0017] Calculate the quantization parameter scale factor scale x and zero point zero_point x of the activation values according to the following formula:
[0018]
[0019] where: x max is the maximum value of the activation values, and x min is the minimum value of the activation values; n x represents the quantization bit width, which determines the maximum value and minimum value
[0020] Based on the weight quantization parameter scale obtained above w and zero_point w , calculate the quantized weight W of the neural network model according to the following formula Q :
[0021]
[0022] Where: W is the weight of the full-precision model; Round() represents the rounding function; Clamp() represents the truncation function, which restricts the quantized weight W Q within the quantization domain; n w represents the quantization bit width of the weight;
[0023] Based on the activation value quantization parameter scale obtained above x and zero_point x , calculate the quantized activation value X of the neural network model according to the following formula Q :
[0024]
[0025] Where: X is the activation value of the full-precision model, Round() represents the rounding function; Clamp() represents the truncation function, which restricts the quantized activation value X Q within the quantization domain; n x represents the quantization bit width of the activation value;
[0026] Based on the quantized weight W Q and activation value X Q , calculate the output Output of the quantized model Q :
[0027] Output Q = W Q · X Q .
[0028] Furthermore, the second quantization module takes blocks as the basic unit and calculates the loss Loss based on the outputs before and after quantization i , and optimizes the weight rounding parameter V and activation value scale factor scale of each block of the model based on batch-size calibration samples x ; including:
[0029] (1) Taking blocks as the basic unit, calculate the quantization loss Loss based on the outputs Output i before and after quantization and : i :
[0030]
[0031] Where: h(V i ) is the modified Sigmoid function, aiming to relax the Round() function in into a continuous form; V i is the rounding parameter of the weights; the parameter λ controls the weight of the rounding loss; τ is the temperature parameter;
[0032] Minimize the quantization loss Loss i as the objective function, that is, minimize the rounding loss of the i-th block and the reconstruction loss
[0033] (2) Input batch-size calibration data samples into the model, and based on the rounding loss of V i gradient, optimize the rounding parameter V of the weights through the following formula i to minimize the rounding loss:
[0034]
[0035] Where: represents the learning rate of V i ;
[0036] Based on the reconstruction loss of gradient, optimize the scale factor of the activation value through the following formula parameter to minimize the reconstruction loss:
[0037]
[0038] Where: represents the learning rate of ;
[0039] Based on the above steps, use gradient optimization to update V i and to minimize the rounding loss and the reconstruction loss, thereby minimizing the quantization loss;
[0040] Perform the same operations as above on the remaining blocks of the model obtained by the second quantization module to obtain a high-precision quantization model.
[0041] The present invention can also adopt the following technical solution, including the following steps:
[0042] S1. Collect data samples with smooth pixel distribution from the training dataset as calibration data;
[0043] S2. Input the calibration data into the pre-trained model for forward propagation, and obtain the quantized weights W according to the pre-trained weights and quantization bit widthQ and the activation value X Q , and calculate the quantized model output, including:
[0044] 201. Extract 16 samples from the calibration data and perform forward propagation on the pre-trained model to optimize four quantization variables, namely scale w , zero_point w , scale x and zero_point x ; among them, scale w and zero_point w are the quantization parameter scale factor and zero point of the weights, and scale x and zero_point x are the quantization parameter scale factor and zero point of the activation value;
[0045] 202. Calculate the quantization parameter scale factor scale w and zero point zero_point w of the weights according to the following formula:
[0046]
[0047] where: w max is the maximum value of the weights, w min is the minimum value of the weights; n w is the quantization bit width, which determines the maximum value and minimum value
[0048] 203. Calculate the quantization parameter scale factor scale x and zero point zero_point x of the activation value according to the following formula:
[0049]
[0050] where: x max is the maximum value of the activation value, x min is the minimum value of the activation value; n x represents the quantization bit width, which determines the maximum value and minimum value
[0051] 204. Based on the above-obtained weight quantization parameters scale w and zero_point w , calculate the quantized weight W Q of the model according to the following formula:
[0052]
[0053] Where: W is the weight of the full-precision model; Round() represents the rounding function; Clamp() represents the truncation function, which restricts the quantized weight W Q within the quantization domain; n w represents the quantization bit width of the weight;
[0054] 205. Based on the activation value quantization parameters scale x and zero_point x obtained above, calculate the activated value X of the quantized model according to the following formula Q :
[0055]
[0056] Where: X is the activation value of the full-precision model, Round() represents the rounding function; Clamp() represents the truncation function, which restricts the quantized activation value X Q within the quantization domain; n x represents the quantization bit width of the activation value;
[0057] 206. Based on the quantized weight W Q and the activation value X Q , calculate the output Output of the quantized model Q :
[0058] Output Q =W Q ·X Q ;
[0059] S3. Adopt a sampling strategy without replacement to draw batch-size samples from the calibration data multiple times, and remove the selected data after each sampling, and iteratively calibrate the quantized model obtained by the first quantization module;
[0060] S4. Taking blocks as the basic unit, calculate the loss Loss based on the outputs before and after quantization i , and optimize the weight rounding parameter V and the activation value scale factor scale of each block of the model based on batch-size calibration samples x ; including:
[0061] 401. Taking blocks as the basic unit, calculate the quantization loss Loss based on the outputs Output i and before and after quantization i :
[0062]
[0063] Where: h(V i ) is the modified Sigmoid function, aiming to relax the Roud() function in into a continuous form; V i is the rounding parameter of the weights; the parameter λ controls the weight of the rounding loss; τ is the temperature parameter;
[0064] 402. Minimize the quantization loss Loss i as the objective function, that is, minimize the rounding loss of the i-th block and the reconstruction loss Input batch-size calibration data samples into the model. Based on the gradient of the rounding loss with respect to V i , optimize the rounding parameter V of the weights through the following formula i to minimize the rounding loss:
[0065]
[0066] Where: represents the learning rate of V i ;
[0067] Based on the gradient of the reconstruction loss with respect to , optimize the scale factor of the activation value through the following formula of the parameter to minimize the reconstruction loss:
[0068]
[0069] Where: represents the learning rate of ;
[0070] Based on the above steps, use gradient optimization to update V i and to minimize the rounding loss and the reconstruction loss, thereby minimizing the quantization loss;
[0071] 403. Perform the same operations as above on the remaining blocks of the model obtained by the second quantization module to obtain a high-precision quantization model.
[0072] Beneficial effects
[0073] Compared with the traditional technical solution, the beneficial effects brought by the present invention are:
[0074] (1) Improve the accuracy of the model under multiple quantization bit widths: The present invention selects calibration data with smooth pixel distribution during quantization to reduce the rounding loss and the reconstruction loss, thereby improving the performance of the quantization model on resource-constrained IoT devices, such as Figure 7As shown. Under the W4A4 quantization configuration, based on the ResNet-101 model, the accuracy of QDrop is 72.98%, and the present invention reaches 76.33%, achieving a 2.88% improvement. Under the W2A4 quantization configuration, based on the MNasx2 model, the accuracy of QDrop is 62.36%, and the accuracy of the present invention is 64.79%, an increase of 2.43%. Under the W3A3 quantization configuration, based on the MobileNetV2 model, the accuracy of QDrop is 54.27%, and the accuracy of the present invention is 57.76%, an increase of 3.49%. This performance improvement is mainly due to the present invention's selection of samples with smooth pixel distributions as calibration data. Without increasing any model complexity, the present invention significantly improves the accuracy of existing quantization methods for models at different quantization bit widths.
[0075] (2) Demonstrates a significant accuracy improvement at low-bit quantization bit widths such as W2A2: Based on the ResNet-18 model, the accuracy of BRECQ is 42.54%, and the accuracy of the present invention is 42.68%, an increase of 0.14%. While the accuracy of QDrop is 51.14%, and the accuracy of the present invention reaches 55.26%, an increase of 4.12%. Based on the ResNet-101 model, the accuracy of QDrop is 59.68%, and the accuracy of the present invention is 64.32%, an increase of 4.64%. Based on the MobileNetV2 model, the accuracy of QDrop is 8.46%, and the accuracy of the present invention is 14.47%, an increase of 6.01%. The significant improvement in the accuracy of low-bit quantization models is mainly due to effectively alleviating the quantization error accumulation effect. Specifically, the present invention reduces the rounding loss and reconstruction loss of the model by selecting appropriate calibration data and adopts a sampling strategy without replacement to avoid calibrating the quantization model with the same samples, thereby improving the performance of low-bit quantization models and promoting the efficient deployment of low-bit models on Internet of Things device terminals.
[0076] (3) Reduces the number of calibration data required for the quantization model: To solve the problem of the large number of calibration data used when calibrating the quantization model, the present invention adopts a random sampling strategy without replacement for calibration data to ensure that the calibration data sampled each time is not repeated, avoiding quantization losses concentrating on specific samples, thereby maintaining the performance of the quantization model without decline while reducing the number of calibration data, as Figure 6 shown. In addition, calibration data refers to a small data subset extracted from the original training data, and reducing its quantity better meets the requirements of actual application scenarios. Description of the Drawings
[0077] Figure 1 is a flowchart of the post-training quantization data calibration method for the present invention to provide visual services for Internet of Things devices.
[0078] Figure 2It is an overview diagram of the post-training quantization data calibration system for providing visual services to Internet of Things (IoT) devices in the present invention. In the cloud server, a neural network model is first trained and quantized to low bits, and then these low-bit models are sent to the IoT device side for deployment, so as to provide high-quality visual services, such as image classification, object detection, and target tracking, with low resource consumption and fast response speed.
[0079] Figure 3 It is an example diagram of the post-training quantization data calibration method for providing visual services to IoT devices in the present invention to provide visual services such as object detection on the IoT device side.
[0080] Figure 4 It is a framework diagram of the post-training quantization data calibration system for providing visual services to IoT devices in the present invention, that is, a data calibration strategy, which specifically includes steps such as pre-training of the neural network model, calibration data selection, quantization, calibration data sampling strategy, and calibration, and finally obtains a quantized model.
[0081] Figure 5 It is a comparison diagram of the qualitative object detection effects of the quantized model of the post-training quantization data calibration method for providing visual services to IoT devices in the present invention. The prediction results of the true annotation, full-precision model, QDrop method, and the quantized model of the present invention are shown in sequence from left to right. The prediction label and its confidence score are marked in the upper left corner of the detection box. The first two rows are the W4A4 quantization settings, and the last two rows are the W2A4 quantization settings.
[0082] Figure 6 It is a comparison diagram of the effects of the without-replacement sampling strategy RSWOR of the calibration data and traditional random sampling on the object detection network model in the post-training quantization data calibration method for providing visual services to IoT devices in the present invention.
[0083] Figure 7 It is a comparison diagram of the Top-1 accuracy of the model on the ImageNet-1K dataset under multiple quantization bit widths in the post-training quantization data calibration method for providing visual services to IoT devices in the present invention. Detailed implementation manners
[0084] The following Figure 1 ~ Figure 7 is an explanation of the present invention as follows:
[0085] As Figure 1 , Figure 2 shown, the present invention provides a post-training quantization data calibration system and method for providing visual services to IoT devices. The data calibration system includes a deep neural network model pre-training module, a first data calibration module, a first quantization module, a second data calibration module, and a second quantization module; wherein:
[0086] The first data calibration module collects data samples with smooth pixel distribution from the training dataset as calibration data;
[0087] The first quantization module inputs the calibration data into the pre-trained model for forward propagation, and calculates the quantization parameter scale factors scale and zero_points zero_point of the weights and activation values according to the pre-trained weights and quantization bit width. Then, the weights W and activation values X of the full-precision model are mapped to a low-bit width (e.g., 8bit, 4bit), so as to obtain the quantized weights W Q and activation values X Q , to calculate the output Output of the quantization model Q ;
[0088] Output Q =W Q ·X Q ,
[0089] The second data calibration module uses a sampling strategy without replacement to draw batch-size samples from the calibration data multiple times, and removes the selected data after each sampling, iteratively calibrating the quantization model obtained by the first quantization module;
[0090] The second quantization module takes blocks as the basic unit, calculates the loss Loss according to the output before and after quantization i , and optimizes the weight rounding parameter V i of each block of the model and the activation value scale factor
[0091]
[0092] where: h(V i ) is the modified Sigmoid function, aiming to relax the Round() function in to a continuous form; V i is the rounding parameter of the weight; the parameter λ controls the weight of the rounding loss; τ is the temperature parameter;
[0093] Based on the rounding loss the gradient of V i and the reconstruction loss the gradient of , update V i and to minimize the rounding loss and the reconstruction loss, thereby minimizing the quantization loss:
[0094]
[0095] where: represents the learning rate of V i , representation of the learning rate.
[0096] The present invention can also adopt the following technical solutions:
[0097] Given a pre-trained neural network model weight file and a calibration dataset, a quantized model can be obtained by using the PTQ method. The specific technical solution of the present invention includes the following steps (as shown in Figure 1 and Figure 3 ):
[0098] S1. Calibration data selection: Take a small number of data samples with smooth pixel distribution from the original training dataset as calibration data, usually covering all categories of the training dataset, and the sampling quantity is adjusted according to different datasets;
[0099] S2. Input the calibration data into the pre-trained model for forward propagation, and obtain the quantized weights W Q and activation values X Q , and calculate the quantized model output, including:
[0100] 201. Initialize quantization parameters: Extract 16 samples from the calibration data selected in S1 and input them into the pre-trained model for forward propagation to optimize four quantization variables, namely scale w , zero_point w , scale x and zero_point x . Among them, scale w and zero_point w are the scale factor and zero point of the weights, and scale x and zero_point x are the scale factor and zero point of the activation values.
[0101] 202. Calculate the scale factor scale w and zero point zero_point w of the weights according to the following formula:
[0102] Let the maximum and minimum values of the model weights be w max and w min . The quantization bit width is n w , which determines the maximum value and minimum value of the quantization domain. The quantization parameters of the weights can be calculated by formulas (1) and (2):
[0103]
[0104] 203. Calculate the quantization parameter scale factor scale of the activation value according to the following formula x and the zero point zero_point x :
[0105] Let the maximum and minimum values of the model activation value be x max and w min . n x represents the quantization bit width, which determines the maximum value and the minimum value The quantization parameters of the activation value can be calculated by formulas (3) and (4):
[0106]
[0107] During the process of initializing the quantization parameters, the distribution of the activation value changes dynamically with the input data. Therefore, it is necessary to input the calibration data into the model for forward propagation to statistically calculate its floating-point range.
[0108] 204. Calculate the quantized weights and activation values: According to the quantization parameters of the weights and activation values obtained in steps 202 and 203, the present invention calculates the quantized weights W Q and the activation value X Q :
[0109]
[0110] where the weight of the model is W, the activation value is X; Round() represents the rounding function; Clamp() represents the truncation function, which restricts the quantized weight W Q and the activation value X Q within the quantization domain; n w and n x respectively represent the quantization bit widths of the weights and activation values.
[0111] 205. Based on the quantized weights W Q and the activation value X Q , calculate the output Output of the quantized model Q :
[0112] Output Q = W Q ·X Q ; (7)
[0113] S3. Adopt a sampling strategy without replacement for calibration data and calculate the quantization loss: Based on S2, a preliminary quantization model can be obtained. However, at this time, the performance of the quantization model drops severely and needs to be calibrated. The core of the calibration process is to input the calibration data into the model for forward propagation, calculate the quantization loss, and optimize the parameters of each block of the model through backpropagation. The number of iterations in the calibration process is usually set to 20,000. Each iteration requires continuous sampling batch-size times (only one image is drawn each time) from the calibration data, and a total of batch-size images are obtained for the calibration process. In order to minimize the usage of calibration data while ensuring the performance of the quantization model, the present invention proposes a sampling strategy without replacement for calibration data (RSWOR). Compared with the original random sampling (RS) method, RSWOR removes the selected data in each sampling process, avoiding the concentration of quantization loss on specific samples, thereby maintaining the quantization performance while reducing the quantity of calibration data.
[0114] S4. In each iteration of calibration, input the calibration data sampled by the RSWOR method into the model to perform quantization operations. Then, according to the weights and activation values before and after quantization, the present invention calculates the quantization loss Loss of each block independently with the block as the basic unit (a layer can be regarded as a special block). i . Usually, the quantization loss includes the rounding loss and the reconstruction loss Among them, the rounding loss stems from the loss of the fractional part when quantizing floating-point numbers to integers, and the reconstruction loss refers to the information loss caused by the inability of the quantized integer value to accurately recover the original data. Taking the i-th block of the model as an example, the calculation process of the quantization loss is as follows:
[0115]
[0116] Where: h(V i ) is the modified Sigmoid function, aiming to relax the Round() function in into a continuous form, V i is the rounding parameter of the weight; the parameter λ controls the weight of the rounding loss; τ is the temperature parameter.
[0117] Based on the rounding loss for the gradient of V i and the reconstruction loss for the gradient of , update the rounding parameter V i of the weight and the quantization parameter of the activation value. The present invention takes minimizing the quantization loss as the objective function, that is, minimizing the rounding loss and reconstruction loss of the i-th block. Among them, the rounding loss is minimized by optimizing V i , and to minimize the reconstruction loss:
[0118]
[0119] The present invention updates V using gradient optimization i and to minimize the rounding loss and the reconstruction loss. The update formula is as follows:
[0120]
[0121] where and represent the learning rates of V i and respectively.
[0122] Determine the quantization model: For the i-th block, based on the weight rounding parameter V i determined in S4 and the activation value quantization parameter quantize this block to minimize the rounding loss and the reconstruction loss, thereby minimizing the quantization loss. Perform the same operation on the remaining blocks of the model, and finally obtain the quantized model, reducing the usage of calibration data while improving the performance of the quantized model. In this way, the quantized model can be deployed under cloud-assisted training and run on IoT devices with high performance to provide visual services.
[0123] Through the above steps, the present invention obtains a quantized model with high performance and reduces the required number of calibration data, improving the visual service quality of the model deployed on the IoT device side, as Figure 5 shown. Each step is closely related, constituting a complete technical solution, effectively solving the problems of low performance of the quantized model and large demand for calibration data.
[0124] In the experiment, the hardware resources used in the experiment of the present invention are NVIDIA A100 Tensor Core GPU and Intel(R) Xeon(R) Platinum 8163 CPU. The environment is built on the PyTorch 1.8 framework and CUDA 11.1 is used for computing acceleration. During the quantization process, the weights are quantized channel-wise, while the activations are quantized tensor-wise. For the PTQ calibration process, the number of iterations for each layer or block is set to 20K, the loss function is MSE, and the output difference before and after quantization is minimized. Under this configuration, the present invention improves the visual service quality on the IoT device side. The experimental results of image classification and object detection services on the ImageNet-1K and MS COCO datasets show that the proposed method is superior to other state-of-the-art methods.
[0125] For example, when quantizing MobileNetV2 to W3A3 based on the ImageNet-1K dataset, the accuracy of CaPTQ is approximately 3.7% higher than that of BRECQ and approximately 3.4% higher than that of QDROP. In the object detection service on the MS COCO dataset, when quantizing the RetinaNet (with ResNet-50 as the backbone) model to W2A4, the accuracy of CaPTQ reaches an average accuracy of 34.4% (mean Average Precision, mAP), which is 1.3% higher than that of QDrop. When using the same amount of calibration data and quantizing ResNet-101 to W2A2, the accuracy of CaPTQ is approximately 4.6% higher than that of QDrop.
Claims
1. A post-training quantization data calibration system for providing visual services to Internet of Things devices, characterized in that, The data calibration system includes a deep neural network model pre-training module, a first data calibration module, a first quantization module, a second data calibration module, and a second quantization module; wherein: The first data calibration module collects data samples with smooth pixel distribution from the training dataset as calibration data; The first quantization module inputs the calibration data into the pre-trained module for forward propagation, and obtains the quantized weight W Q Q and the activation value X Q Q , and calculates the output of the quantized model; The second data calibration module adopts a sampling strategy without replacement to extract batch-size samples from the calibration data, and removes the selected data after each sampling, and iteratively calibrates the quantization model obtained by the first quantization module; The second quantization module takes blocks as the basic unit, calculates the loss Loss based on the outputs before and after quantization i , and optimizes the weight rounding parameter V and the activation value scale factor scale of each block of the model based on batch-size calibration samples x .
2. The post-training quantization data calibration system for providing visual services to Internet of Things devices according to claim 1, characterized in that: The first quantization module inputs the calibration data into the pre-trained model for forward propagation, and obtains the quantized weight W and activation value X according to the pre-trained weights and quantization bit width, including: Q and activation value X Q , including: Extract 16 samples from the calibration data and input them into the pre-trained module for forward propagation to optimize four quantization variables, namely scale w , zero_point w , scale x and zero_point x ; among them, scale w and zero_point w are the quantization parameter scale factor and zero point of the weights, and scale x and zero_point x are the quantization parameter scale factor and zero point of the activation values; Calculate the quantization parameter scale factor scale of the weight according to the following formula w and the zero point zero_point w : Where: w max is the maximum value of the weight, w min is the minimum value of the weight; n w is the quantization bit width, determining the maximum and minimum values Calculate the quantization parameter scale factor scale of the activation value according to the following formula x and the zero point zero_point x : Where: x max is the maximum value of the activation value, x min is the minimum value of the activation value; n x represents the quantization bit width, which determines the maximum and minimum values of the activation value quantization range Based on the weight quantization parameter scale obtained above w and zero_point w , calculate the quantized weight W of the neural network model according to the following formula Q :[[]]END]] Where: W is the weight of the full-precision model; Round() represents the rounding function; Clamp() represents the truncation function, which restricts the quantized weight W Q within the quantization domain; n w represents the quantization bit width of the weight; Based on the activation value quantization parameter scale obtained above x and zero_point x calculate the quantized activation value X of the neural network model according to the following formula Q : Where: X is the activation value of the full-precision model, Round() represents the rounding function; Clamp() represents the truncation function, which truncates the quantized activation value X Q within the quantization domain; n x represents the quantization bit width of the activation value; Based on the quantized weight W Q and the activation value X Q , calculate the output Output of the quantization model Q : Output Q = W Q · X Q 。 3. The post-training quantization data calibration system for providing visual services to Internet of Things devices according to claim 1, characterized in that: The second quantization module takes blocks as the basic unit and calculates the loss Loss based on the outputs before and after quantization i , and optimizes the weight rounding parameter V and the activation value scale factor scale of each block of the model based on the batch-size calibration samples x ; It includes: (1) Taking the block as the basic unit, calculate the quantization loss Loss based on the output Output before and after quantization i and calculate the quantization loss Loss i : where: h(V i ) is the modified Sigmoid function, which aims to relax the Round() function in into a continuous form; V i is the rounding parameter of the weights; the parameter λ controls the weight of the rounding loss; τ is the temperature parameter; Minimize the quantization loss Loss i as the objective function, that is, minimize the rounding loss of the i-th block and the reconstruction loss (2) Input batch-size calibration data samples into the model, and optimize the rounding parameter V of the weights based on the rounding loss for V i gradient, and optimize the rounding parameter V of the weights through the following formula i to minimize the rounding loss: Wherein: represents the learning rate of V i ; Based on the reconstruction loss For the gradient of, optimize the scale factor of the activation value through the following formula parameters to minimize the reconstruction loss: Wherein: represents learning rate; Based on the above steps, update V using gradient optimization i and to minimize the rounding loss and the reconstruction loss, thereby minimizing the quantization loss; Perform the same operations as above on the remaining blocks of the model obtained by the second quantization module to obtain a high-precision quantization model.
4. The post-training quantization data calibration method for providing visual services to Internet of Things devices by using the system according to claim 1, characterized in that It includes the following steps: S1. Collect data samples with smooth pixel distribution from the training dataset as calibration data; S2. Input the calibration data into the pre-trained model for forward propagation, and obtain the quantized weight W according to the pre-trained weights and the quantization bit width Q and the activation value X Q , and calculate the output of the quantized model, including:
201. Extract 16 samples from the calibration data and input them into the pre-trained model for forward propagation to optimize four quantization variables, namely scale w , zero_point w , scale x and zero_point x ; among them, scale w and zero_point w are the quantization parameter scale factor and zero point of the weights, and scale x and zero_point x are the quantization parameter scale factor and zero point of the activation values; 202. Calculate the quantization parameter scale factor scale of the weight according to the following formula w and the zero point zero_point w : Where: w max is the maximum value of the weight, w min is the minimum value of the weight; n w is the quantization bit width, which determines the maximum value and the minimum value 203. Calculate the quantization parameter scale factor scale of the activation value according to the following formula x and the zero point zero_point x : where: x max is the maximum value of the activation value, x min is the minimum value of the activation value; n x represents the quantization bit width, which determines the maximum and minimum values 204. Based on the weight quantization parameter scale w and zero_point w , calculate the quantized weight W of the model according to the following formula Q : Wherein: W is the weight of the full-precision model; Round() represents the rounding function; Clamp() represents the truncation function, which truncates the quantized weight W Q within the quantization domain; n w represents the quantization bit width of the weight; 205. Based on the activation value quantization parameter scale x and zero_point x obtained above, calculate the activation value X of the quantized model according to the following formula Q : Where: X is the activation value of the full-precision model, Round() represents the rounding function; Clamp() represents the truncation function that truncates the quantized activation value X Q within the quantization domain; n x represents the quantization bit width of the activation value; 206. Based on the quantized weight W Q and the activation value X Q , calculate the output Output of the quantization model Q : Output Q = W Q ·X Q , S3. Adopt a sampling strategy without replacement to extract batch-size samples from the calibration data multiple times, and remove the selected data after each sampling, and iteratively calibrate the quantization model obtained by the first quantization module; S4. Calculate the loss Loss with the block as the basic unit according to the outputs before and after quantization i , and optimize the weight rounding parameter V and the activation value scale factor scale of each block of the model based on the batch-size calibration samples x ; including:
401. Taking a block as the basic unit, calculate the quantization loss Loss based on the output Output before and after quantization i and calculate the quantization loss Loss i : where: h(V i ) is the modified Sigmoid function, aiming to relax the Round() function in into a continuous form; V i is the rounding parameter of the weights; the parameter λ controls the weight of the rounding loss; τ is the temperature parameter; 402. Minimize the quantization loss Loss i as the objective function, that is, minimize the rounding loss of the i-th block and the reconstruction loss Input batch-size calibration data samples into the model. Based on the rounding loss for the gradient of V i optimize the rounding parameter V of the weights through the following formula i to minimize the rounding loss: Wherein: represents the learning rate of V i ; Based on the reconstruction loss For the gradient of, optimize the scale factor of the activation value through the following formula parameters to minimize the reconstruction loss: Wherein: represents learning rate; Based on the above steps, use gradient optimization to update V i and to minimize the rounding loss and the reconstruction loss, thereby minimizing the quantization loss; 403. Perform the same operations as above on the remaining blocks of the model obtained by the second quantization module to obtain a high-precision quantization model.