Quantization parameter prediction method for improving precision of quantization network
By building a parameter prediction network and predicting the quantitative parameters of the neural network, the quantitative accuracy loss problem caused by large data differences in the existing technology is solved, and higher quantization model accuracy and recognition accuracy are achieved.
Patent Information
- Application Number
- CN202311787400.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
In the case of large data differences in the prior art, the quantization method leads to serious accuracy losses and cannot effectively adapt to the quantization range of different data.
A quantitative parameter prediction method is proposed. By constructing a parameter prediction network, using the network input as a subset of the training data set, the output is the predicted quantitative parameters, combined with the mse loss function for training, iterating to convergence, to adapt to the quantization parameters of different inputs.
The prediction network obtains quantitative parameters that are more in line with the data distribution, which improves the quantized model accuracy and enhances the correctness and recognition accuracy of the model.
Smart Images

Figure CN120197653A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network quantization, and particularly relates to a quantization parameter prediction method for improving the accuracy of a quantized network. Background Art
[0002] With the development of technology, especially the neural network applicable to fields such as image recognition, face recognition, video surveillance processing, etc., with the breakthrough development of artificial intelligence, it has increasingly become the trend of current technological development and is well-known and used by people. Due to the strict requirements for computational complexity, memory, and power consumption, modern deep neural networks (DNNs) cannot be effectively used in mobile and embedded devices. Quantization of weights and feature maps (activations) is a common method to solve this problem. Therefore, neural network quantization is a technology used to reduce the number of parameters in a neural network, thereby improving its operating efficiency and saving computational resources, and further improving the efficiency of the entire operating system, including reducing hardware redundancy and enhancing software efficiency. In a neural network, parameters are usually floating-point numbers, which require a large amount of memory and computational resources to store and calculate. By converting these parameters into low-bit numbers or fixed-point numbers, the storage and computational costs can be greatly reduced, thereby accelerating the operating speed of the system where the neural network is located. Specifically, it can be divided into post-quantization and training quantization.
[0003] Among them, neural network post-quantization refers to directly inserting quantization nodes based on a floating-point model. Training quantization is to introduce quantization into the training. It can convert high-precision parameters (such as 32-bit floating-point numbers) used in the training process into low-precision parameters (such as 8-bit integers), thereby reducing the storage space and memory occupancy of the network model and improving the speed and efficiency of model inference. On some edge devices, such as camera chips, development boards, AI chips and other terminal devices, this technology can significantly reduce power consumption and extend the battery life of the device.
[0004] In addition, in neural network quantization, in order to reduce the storage and computational amount of the model, it is necessary to convert the parameters (such as weights and biases) in the original model from high-precision (such as 32-bit floating-point numbers) to low-precision integers or fixed-point numbers, so as to store them in a smaller memory. This conversion process involves quantization parameters.
[0005] Among the quantization parameters, the quantization range is relatively important, which represents the distribution of the original floating-point data set. The quantization algorithm hopes that the quantized data is consistent with the original floating-point data distribution after dequantization.
[0006] However, for some networks, there may be a large difference in the quantization range for different data. At this time, quantization is very likely to cause losses, and the quantization methods in the prior art have serious accuracy losses in the case of large data differences. Summary of the Invention
[0007] To solve the above problems, the purpose of this application is to propose a method that uses a network to predict quantization parameters that better conform to the data distribution. For different inputs, different quantization parameters can be predicted to adapt to the quantization algorithm, improving the accuracy of the quantized model. The model accuracy represents the correctness of the model and can be manifested as an increase in the face recognition rate and more accurate object recognition. Generally, there are accuracy indicators for algorithm implementation, and a certain accuracy must be achieved to ensure normal use without being affected.
[0008] Specifically, the present invention provides a quantization parameter prediction method for improving the accuracy of a quantization network. The method includes:
[0009] S1. Select a subset of a batch of datasets. The batch of data refers to the training dataset of a neural network. Randomly select a small amount of datasets from it, which is a random sampling of the training dataset. Statistically calculate the quantization parameters of each input data. The quantization parameters refer to the scale and zero_point of each layer of feature map in the network. These two values are calculated from the quantization interval clip_min and clip_max of the feature map. The quantization interval indicates that most of the output data of this layer is within this interval. Record the input and the corresponding quantization parameters, that is, write the quantization parameters of each layer into a local file for storage.
[0010] S2. Construct a parameter prediction network. Among them, the network input is the input recorded in step S1, and the output is the predicted quantization parameters, that is, the scale and zero_point of each layer. Calculate the mse loss with the real quantization parameters.
[0011] The recorded input is the input data of the original network saved and the quantization parameters corresponding to this input. For the prediction network, its task is to predict the quantization parameters of the original network. Input the original input data and output the predicted quantization parameters. Use mse as the loss function during the training of the prediction network. Among them, mse is the mean squared error, that is, the sum of the squares of the differences of the corresponding data. Assume that the input of the original network is x, the recorded corresponding quantization parameter is t, and the prediction of the prediction network is y. Then loss = sum((t - y)^2). During training, the smaller the loss, the more accurate the prediction.
[0012] S3. Use the data in step S1 to train the parameter prediction network and iterate until convergence. Here, the recorded input and label are used as the training dataset for neural network training. Input the data into the network, the network calculates the result, compares the prediction result with the label, and updates the network parameters following the gradient backpropagation until the loss no longer decreases.
[0013] S4. After quantizing the original network, perform parameter prediction through the parameter prediction network, and load the predicted parameters into the quantization network for testing.
[0014] In step S2, first, the features of the input data are extracted through a convolution operator, and then the features are compressed again through pooling. Both average pooling and maximum pooling are used to extract multi-dimensional information, and then the extracted features are mapped to parameters through a fully connected layer.
[0015] In step S2, the kernel of the prediction network is not fixed. It changes according to the size of the input data. The larger the data, the more times of downsampling and the more layers of the network. If the input is an image, generally, the larger the resolution, the more times of downsampling and the deeper the network.
[0016] The sizes used are as follows:
[0017] The conditions that must be met are that the network output must be the same as the number of quantization parameters of the original network. Assuming that the original network has 10 quantization layers and each layer has two quantization parameters, the output of the prediction network should also be 20.
[0018] The structure of the parameter prediction network further includes:
[0019] S2.0, input data;
[0020] S2.1, perform the first layer of convolution;
[0021] S2.2, use the output as the input of the second layer of convolution and perform the second layer of convolution;
[0022] S2.3, use the output as the input of the third layer of convolution and perform the third layer of convolution;
[0023] S2.4, perform global maximum pooling and global average pooling on the output of the third layer of convolution respectively;
[0024] S2.5, perform a fully connected layer on the two results of the previous step S2.4; perform matrix multiplication matmul and output.
[0025] Step S2 further includes:
[0026] S2.1, perform the first layer of convolution, using: kernel 32*3*3*3, stride 2*2;
[0027] S2.2, use the output as the input of the second layer of convolution and perform the second layer of convolution; using: kernel64*32*3*3, stride 2*2;
[0028] S2.3, use the output as the input of the third layer of convolution and perform the third layer of convolution; using: kernel128*64*3*3, stride 2*2;
[0029] S2.4. The pooling in the third layer uses global pooling, namely globalavgpool and globalmaxpool. Assume there is a tensor with a shape of 1*3*64*64. Then globalavgpool calculates the average value of 64*64 numbers on each channel respectively, and the output is 1*3*1*1. The same applies to Globalmaxpool.
[0030] S2.5. Matmul is a fully connected layer and is matrix multiplication in terms of calculation.
[0031] Therefore, the advantages of this application are as follows: Appropriate quantization parameters can be predicted, and the final quantization effect is also higher than that of previous quantization methods. This application also proves that there is a correlation between the network input and the quantization parameters. When the neural network model is deployed on a device, due to the high computational density of the neural network, float operations are generally converted into integer operations, but this has a risk of precision loss. The process of converting into integer operations is called quantization, and the quantization parameters have a great impact on the model precision. By using the method of this application, better quantization parameters can be obtained, thereby improving the precision of the quantized model. The quantized model has a precision no worse than that of the float model, and is faster in operation speed, less in memory consumption, and lower in bandwidth occupancy than the float model. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not constitute a limitation to the present invention.
[0033] Figure 1 It is a schematic diagram of the data distribution involving quantization using the quantization range of t1 in this application.
[0034] Figure 2 It is a schematic diagram of the data distribution involving quantization using the quantization range of t2 in this application.
[0035] Figure 3 It is a schematic diagram of the structure of the parameter prediction network in this application.
[0036] Figure 4 It is a schematic diagram of the flow of the method in this application.
[0037] Figure 5 It is a schematic diagram of the assumed existing network structure in the embodiment.
[0038] Figure 6 It is a diagram for a specific input.
[0039] Figure 7 It is Figure 6 a schematic diagram of the corresponding quantization value range.
[0040] Figure 8It is a diagram for another specific input.
[0041] Figure 9 is Figure 8 a schematic diagram of the corresponding quantization value range.
[0042] Figure 10 is a schematic diagram of constructing a prediction network by saving the above input image and its parameters.
[0043] Figure 11 is based on Figure 10 the prediction network, using the saved data for training. After training is completed, it is a schematic diagram of the input in testing the prediction network effect.
[0044] Figure 12 is based on Figure 10 the prediction network, using the saved data for training. After training is completed, it is a schematic diagram of the output of the prediction network in testing the prediction network effect. Detailed implementation manners
[0045] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0046] The present invention belongs to the improvement of the neural network quantization method, mainly for improving the generalization ability of the neural network.
[0047] For the quantization technology of converting the existing float model into an integer model, the quantization technology is widely used in model deployment. In the post-quantization process of the existing technology, it includes:
[0048] 1. Prepare a trained high-precision neural network model. This model can be trained from scratch or fine-tuned using a pre-trained model.
[0049] 2. Select the quantization precision and scheme. Selecting an appropriate quantization precision and scheme can balance between reducing the model size and improving the inference speed. Generally, selecting a low-bitwidth quantization precision can reduce the model size and accelerate the inference speed, but may reduce the model accuracy.
[0050] 3. Convert the high-precision parameters to low-precision parameters. In this step, the parameters in the model (such as weights, biases, etc.) will be quantized to low-precision parameters. Different quantization schemes require different conversion methods. For example, weights use per-channel symmetry, and activation quantization can use non-channel-specific symmetric / non-symmetric methods.
[0051] 4. Perform inference on the quantized model. Deploy the quantized model to the target device and perform inference on the test dataset. Various evaluation metrics (such as accuracy, latency, etc.) can be used to evaluate the performance of the quantized model.
[0052] For the above point 3, when quantifying the output of each layer in the quantization network, the following algorithm is generally adopted: 3.1. Determine the quantization bit width and quantization range:
[0053] First, it is necessary to determine the quantization bit width and quantization range, that is, map the range of floating-point parameters to the range of integer parameters. Common quantization bit widths include 8 bits, 16 bits, 32 bits, etc., and the quantization range usually takes integer ranges such as [-128, 127] or [-32768, 32767].
[0054] 3.2. Calculate the asymmetric quantization parameters scale and zero_point:
[0055] According to the quantization bit width and quantization range, the symmetric quantization parameters scale and zero_point can be calculated. Among them, scale represents the ratio of the range of the original parameters to the range of the quantization parameters, that is, scale = (max_range - min_range) / (quant_max - quant_min), where max_range and min_range are the maximum and minimum values of the original parameters respectively, and quant_max and quant_min are the maximum and minimum values of the quantization parameters respectively. zero_point represents that the center point of the original parameters is mapped to the zero point of the quantization parameters, that is, zero_point = round(-min_range / scale).
[0056] 3.3. Perform quantization and dequantization:
[0057] Quantize the original parameters, using the symmetric quantization formula for quantization, that is, quantized_value = round(float_value / scale) + zero_point, where quantized_value represents the quantized integer parameter and float_value represents the original floating-point parameter. During inference, restore the quantized parameters to the original parameters and use the dequantization formula for dequantization, that is, float_value = (quantized_value - zero_point) * scale.
[0058] Taking 8-bit quantization as an example, quant_max = 127, quant_min = -128, and max_range and min_range are the value ranges of the tensors to be quantized. Such as Figure 1 、 Figure 2As shown, there are two original data distributions; the ranges of the two tensors are significantly inconsistent. If the range of t1 is used for quantization, many values outside the quantization range will be truncated for t2, affecting the accuracy. If the quantization range of t2 is used for quantization, the quantization granularity for t1 is not fine enough, and a lot of bit widths are wasted.
[0059] The algorithms of the prior art are consistent with the quantization algorithm used in this application, and both use quantization parameters to quantize the float model. However, the calculation cost of the quantization parameters in the prior art is relatively large, and it is not possible to calculate them for each input. Generally, a fixed value is taken. This application uses a prediction network to obtain the quantization parameters for each input, achieving the effect of improving the quantization accuracy.
[0060] The best quantization is to use the corresponding quantization range for each tensor. Without re-statistical the tensor range, this application uses a network to predict the quantization parameters, achieving the adaptability of the quantization parameters of each tensor to its own data distribution.
[0061] As Figure 4 shown, the steps of a quantization parameter prediction method for improving the accuracy of a quantization network according to the present invention include:
[0062] S1. Select a subset of a batch of data sets. Different tasks have different data set types, not limited to images, audio, and text sequences; here, a batch of data refers to the training data set of a neural network. A small amount of data sets are randomly selected from it, that is, random sampling of the training data set. It can be said that the selected data set can represent the entire training set to a certain extent. For example, there are 100 numbers from 0 to 1, and 10 numbers are randomly selected. The data distribution is still from 0 to 1. The quantization parameters of each input data are statistically calculated. The quantization parameters refer to the scale and zero_point of each layer of feature map in the network. These two values are calculated from the quantization interval clip_min and clip_max of the feature map. The quantization interval indicates that most of the output data of this layer is in this interval. The input and the corresponding quantization parameters are recorded, and the quantization parameters of each layer are written into a local file for storage.
[0063] S2. Construct a parameter prediction network, where the network input is the input recorded in step S1, and the output is the predicted quantization parameters, namely the scale and zero_point of each layer; calculate the mse loss with the true quantization parameters; the input data of the original network and the corresponding quantization parameters are saved above; for the prediction network, its task is to predict the quantization parameters of the original network, input the original input data, and output the predicted quantization parameters; use mse as the loss function during the training of the prediction network; where mse is the mean square error, that is, the sum of the squares of the differences of the corresponding data; assume: the original network input is x, the recorded corresponding quantization parameter is t, and the prediction of the prediction network is y; then loss = sum((t - y)^2); the smaller the loss during training, the more accurate the prediction.
[0064] S3. Use the data in step S1 to train the parameter prediction network and iterate until convergence. Here, the recorded input and labels are used as the training dataset for the general neural network training process; input the data into the network, the network calculates the result, compares the prediction result with the label, and updates the network parameters following the gradient backpropagation; until the loss no longer decreases.
[0065] S4. After quantizing the original network, perform parameter prediction through the parameter prediction network, and load the predicted parameters into the quantization network for testing; the quantization network is still the quantization of the original network. Because calculating the quantization parameters is time-consuming, a new network is trained to predict the quantization parameters of the original network based on the input data.
[0066] In practical applications, the method of the present application can predict appropriate quantization parameters, and the final quantization effect is also higher than the previous quantization methods. The present application also proves that there is a correlation between the network input and the quantization parameters.
[0067] As Figure 3 shown, step S2 further includes:
[0068] First, extract the features of the input data through the convolution operator, then compress the features again through pooling, and use average pooling and max pooling simultaneously to extract multi-dimensional information. Then use the fully connected layer to map the extracted features into parameters. The quantization parameters can be effectively predicted through the input data.
[0069] Among them, the structure of the parameter prediction network further includes:
[0070] S2.0, input data;
[0071] S2.1, perform the first layer of convolution; in this article, kernel 32*3*3*3 stride 2*2;
[0072] S2.2, Output it as the input of the second - layer convolution and perform the second - layer convolution; in this paper, the kernel is 64*32*3*3 and the stride is 2*2;
[0073] S2.3, Output it as the input of the third - layer convolution and perform the third - layer convolution; in this paper, the kernel is 128*64*3*3 and the stride is 2*2;
[0074] S2.4, Perform global max - pooling and global average - pooling on the output of the third - layer convolution respectively;
[0075] S2.5, Fully connect the two results of the previous step S2.4; perform matrix multiplication matmul and output.
[0076] Specifically, an example is given below: Suppose there is an image classification network classmodel with 10 quantizable layers and its training set train_set. When quantizing classmodel, for each input image, the optimal quantization parameters of the network are different. And according to calculating the optimal quantization parameters of each image, it is necessary to first run the float classmodel, count the data distribution of each layer's output to obtain clip_min and clip_max. Then calculate the quantization parameters according to clip_min and clip_max. This process is too complex and time - consuming in calculation. Now construct a parameter prediction network predmodel, which can predict the quantization parameters required for each layer of classmodel according to the original input image.
[0077] Actual example: Suppose the existing network structure is as Figure 5 shown. The network input is the cifar10 dataset, and its quantizable layers are 4 layers.
[0078] For a specific input, a specific input refers to a certain input that meets the task requirements. Here, it refers to an image in the cifar10 dataset. Suppose it is Figure 6 shown: A cat. Its corresponding quantization value range is Figure 7 shown.
[0079] For a specific input, suppose it is Figure 8 shown: A yacht. Its corresponding quantization value range is Figure 9 shown.
[0080] Save the input image and its parameters such as the above:
[0081] Construct a prediction network, as Figure 10 shown.
[0082] Use the saved data for training. After training is completed, test the effect of the prediction network: Input, such asFigure 11 As shown. The inputs of the original network and the prediction network are the same. That is Figure 6 and Figure 11 are the same. The prediction network outputs, such as Figure 12 as shown. The output data are the quantization parameters of each layer respectively, and there is not much difference compared with the original ones. That is, the output of the prediction network should not have much difference from the quantization parameters of the same input of the original network.
[0083] Thus, a prediction model that can quickly obtain the quantization parameters of the original model is obtained. Every time the original model needs to perform quantization inference, first run the prediction model to obtain the parameters, and then load the parameters into the quantized original model for inference.
[0084] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A quantization parameter prediction method for improving the accuracy of a quantization network, characterized in that The method includes: S1. Select a subset of a batch of datasets. The batch of data refers to the training datasets of a neural network. Randomly select a small amount of datasets from it, which is a random sampling of the training datasets. Statistically calculate the quantization parameters of each input data. The quantization parameters refer to the scale and zero_point of the feature map of each layer in the network. These two values are calculated from the quantization interval clip_min and clip_max of the feature map. The quantization interval indicates that most of the output data of this layer is within this interval. Record the input and the corresponding quantization parameters, that is, write the quantization parameters of each layer into a local file for storage. S2. Construct a parameter prediction network. Among them, the network input is the input recorded in step S1, and the output is the predicted quantization parameters, that is, the scale and zero_point of each layer. Calculate the mseloss with the real quantization parameters. The recorded input is the input data of the original network saved and the quantization parameters corresponding to this input. For the prediction network, its task is to predict the quantization parameters of the original network. Input the original input data and output the predicted quantization parameters. Use mse as the loss function during the training of the prediction network. Among them, mse is the mean square error, that is, the sum of the squares of the differences of the corresponding data. Assume that the original network input is x, the recorded corresponding quantization parameter is t, and the prediction of the prediction network is y, then loss = sum((t - y)^2). During training, the smaller the loss, the more accurate the prediction. S3. Use the data in step S1 to train the parameter prediction network and iterate until convergence. Here, the recorded input and label are used as the training datasets for neural network training. Input the data into the network, the network calculates the result, compares the prediction result with the label, and updates the network parameters following the gradient backpropagation until the loss no longer decreases. S4. After quantizing the original network, perform parameter prediction through the parameter prediction network, and load the predicted parameters into the quantization network for testing.
2. A quantization parameter prediction method for improving the accuracy of a quantization network according to claim 1, characterized in that, In step S2, first extract the features of the input data through a convolution operator, then compress the features again through pooling, and use both average pooling and maximum pooling to extract multi-dimensional information, and then use a fully connected layer to map the extracted features into parameters.
3. The quantization parameter prediction method for improving the accuracy of a quantization network according to claim 2, characterized in that, In step S2 The kernel of the prediction network is not fixed. It changes according to the size of the input data. The larger the data, the more downsampling times and the more network layers. The sizes used are as follows: The condition that must be met is that the network output must be the same as the number of quantization parameters of the original network. Assume that the original network has 10 quantization layers and each layer has two quantization parameters, then the output of the prediction network should also be 20. The structure of the parameter prediction network further includes: S2.0, input data; S2.1, perform the first layer of convolution; S2.2, use the output as the input of the second layer of convolution and perform the second layer of convolution; S2.3, use the output as the input of the third layer of convolution and perform the third layer of convolution; S2.4, perform global maximum pooling and global average pooling on the output of the third layer of convolution respectively. S2.5, fully connect the two results of the previous step S2.4; perform tensor matrix multiplication matmul and output.
4. A quantization parameter prediction method for improving the accuracy of a quantization network according to claim 3, characterized in that The step S2 further includes: S2.1, perform the first layer of convolution, using: kernel 32*3*3*3, stride 2*2; S2.2, use the output as the input of the second layer of convolution and perform the second layer of convolution; use: kernel64*32*3*3, stride2*2; S2.3, use the output as the input of the third layer of convolution and perform the third layer of convolution; use: kernel128*64*3*3, stride2*2; S2.4, the pooling in the third layer uses global pooling, globalavgpool, globalmaxpool; assume: there is a tensor with a shape of 1*3*64*64, then globalavgpool calculates the average value of 64*64 numbers on each channel respectively, and the output is 1*3*1*1; Globalmaxpool is the same; S2.5, matmul is a fully connected layer, and the calculation is matrix multiplication.