A quantization method and system for generating standard image data and improving the accuracy of a neural network model using the standard image data
By scaling, standardizing and interpolation processing of image data, calculating quantization coefficients with KL divergence, optimizing the FPGA model structure, the performance degradation caused by image preprocessing is solved, and efficient image inference and accuracy improvement is achieved.
Patent Information
- Application Number
- CN202411228468.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-09-03
AI Technical Summary
When the prior art performs image inference on FPGA, the image preprocessing method causes the model inference performance to degrade, which cannot meet the input requirements of high-resolution images, and the quantization coefficient calculation is complex, resulting in increased accuracy loss and resource requirements.
By scaling, standardizing and interpolation processing of the input image data, standard image data is generated, and quantization coefficients are calculated in combination with KL divergence, convolution and batch standardization layers are combined, model structure is optimized, and input size and calculation requirements of FPGA are adapted.
It improves the inference performance of the model on FPGA, reduces quantization loss, maintains the accuracy and efficiency of the model, and reduces the computing resource requirements.
Smart Images

Figure CN119090710B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and relates to a quantization method and system for generating standard image data and improving the accuracy of a neural network model by using the standard image data. Background Art
[0002] Existing symmetric range-based linear quantization coding techniques cannot be well adapted to FPGAs. The reasons are that, in addition to the need to recalculate quantization coefficients online for each input image data and the quantization coefficients being floating-point numbers, different preprocessing methods for image data also affect the performance of final model inference. For example, the Chinese invention patent application "A quantization method and system for optimizing int8 and its process" (Patent Application No.: CN202010863091.7). This method mainly provides a quantization method and system for optimizing int8, using the floating-point model saved after training a neural network; calculating the weight quantization scaling factors for each channel in each layer of the neural network; calculating the activation value quantization scaling factors for each layer of the neural network using the KL divergence algorithm; determining the optimal weight quantization scaling factors and optimal activation value quantization scaling factors for each layer of the neural network according to the cosine distance; obtaining the int8 integer result based on the optimal weight quantization scaling factors and optimal activation value quantization scaling factors. According to the method of the present invention, according to the optimal weight quantization scaling factors and optimal activation value quantization scaling factors, automatic fine-tuning is performed to obtain the int8 integer result, which can avoid the influence of extreme values; at the same time, the KL divergence is used to calculate the activation value quantization scaling factors, and automatic fine-tuning is performed using the cosine distance, which can reduce the quantization error and avoid the influence of extreme values. The Chinese invention patent "A method for improving the accuracy of a low-bit neural network model through weight preprocessing" (Patent Application No.: CN202010497769.4). This invention provides a method for improving the accuracy of a low-bit neural network model through weight preprocessing. First, the full-precision data of the neural network is obtained, and then the full-precision weights are preprocessed, that is, the weights are preprocessed before quantization; the data to be quantized is quantized to obtain low-bit data; this method takes into account the bit-width limitation before training the low-bit model, so when training the full-precision model, it will be considered that the values represented in low bits are limited. In order to make full use of each value, the weights will be preprocessed, so as to make full use of each value and thus improve the accuracy of the model.
[0003] The FPGA requires the input size of the image to be a power of 2, and the quantization coefficient also needs to be an integer power of 2. At the same time, due to the limitation of the memory size of the FPGA, the input size of the image cannot be too large. Therefore, for images with higher resolutions, image preprocessing is required before they can meet the input size requirements of the FPGA. Before feeding the network model for training, in order to reduce the size of the original image, the general process of image preprocessing is as follows: First, perform data normalization, and then perform proportional scaling of the image size. To meet the requirement that the input size is a square size, fill in the missing part with a gray border. However, this image preprocessing method will cause a large decline in the inference performance of the quantized model in the FPGA, resulting in the unavailability of the model inference results. Summary of the Invention
[0004] In order to solve the deficiencies of the existing technology, the purpose of the present invention is to provide an image preprocessing method that improves the image inference effect, can generate standard image data, improves the accuracy of the low-bit neural network model through image preprocessing, and is a use of image processing normalization technology in FPGA inference.
[0005] The present invention provides a method for generating standard image data. The method obtains standard image data by scaling and normalizing the input image data and then performing interpolation processing, including the following steps:
[0006] Step 1: Input an image and scale the image data to the interval [0, 1];
[0007] Step 2: Normalize and process the scaled data obtained in Step 1 to obtain normalized data y t ;
[0008]
[0009] where y is the value after normalizing the image data, ∈ is the standard normal distribution, t is the uniform distribution in the interval [0, 1], and π(t) is a monotonically increasing function from 0 to 1; in the specific implementation process, the monotonically increasing function can be a linear function, a logarithmic function, a Sigmoid function, or other higher-order functions, etc., and different monotonically increasing functions can be selected according to actual needs;
[0010] Step 3: Perform interpolation processing on the normalized data obtained in Step 2 to obtain image data z of the target size.
[0011] In Step 1, divide the pixel values in the input image by 256 one by one to scale the image data;
[0012] and / or,
[0013] In step 2, the scaled data obtained in step 1 is normalized to a distribution with a mean of 0 and a variance of 1; where x is the pixel value of the input image, and x mean is the average value of the pixels of the input image, and x σ is the variance of the pixels of the input image;
[0014] and / or
[0015] In step 3, the interpolation processing includes at least one of bilinear interpolation, trilinear interpolation, and nearest neighbor interpolation to match the image size to the model input.
[0016] The present invention also proposes a quantization method for improving the accuracy of a neural network model using standard image data, and the method includes the following steps:
[0017] Step I. Obtain standard image data according to the above method for generating standard image data;
[0018] Step II. Input the standard image data into a deep learning model for training to obtain an inference network model;
[0019] Step III. Convert the inference network model into an ONNX model, and merge the batch normalization (BN) layer and the convolutional layer in the ONNX model;
[0020] Step IV. Use a calibration image to calculate the quantization coefficient by KL divergence to obtain the threshold T of the activation function;
[0021] Step V. Calculate the quantization coefficient based on the quantization formula using the information including weights, biases, variances, and means of each network layer, and verify the accuracy of the quantization coefficient.
[0022] The standard image data adapts to the input size of the deep learning model; and / or
[0023] In step II, the deep learning model is a neural network model applicable to image processing (such as ResNet, etc.) for inference on an FPGA; during the training process, grouped convolution or separable convolution is used to replace the original convolution operator; after training, it also includes the steps of cross-validation and testing of the model; and / or
[0024] In step III, export the inference network model from its original framework to an ONNX format file and perform compatibility adjustment; extract the trained parameters of the inference network model from the BN layer, recalculate the weights and biases of the convolutional layer, so that the convolution operation can directly include the effect of BN, and replace the model originally including the BN layer and the convolutional layer with a model only including the merged convolutional layer; and / or
[0025] In step IV, a calibration image is used to generate an output distribution through a network layer. The output distribution is quantized, and the KL divergence between the quantized distribution and the original floating-point distribution is calculated to minimize the KL divergence, thereby determining the threshold T of the activation function; and / or,
[0026] In step V, according to the quantization formula x f is the input data, n is the bit width of the quantized value, is the quantization coefficient, and the quantization coefficient is obtained by using the weight information, bias information, variance information, and mean information of each network layer.
[0027] When exporting the inference network model from its original framework to an ONNX format file, the architecture and weight information of the model are converted and saved; the model structure is adapted to the ONNX format requirements through compatibility adjustment; and / or,
[0028] The ONNX model is parsed for the network. The input layer and output layer of the network are read in sequence and labeled in sequence to ensure that subsequent processing of the model (such as quantization, optimization, etc.) can be based on a clear hierarchical structure without hierarchical confusion or processing errors.
[0029] After step V, a quantization self-check step may further be included: comparing the quantized output of each layer of the network with the output of its previous layer to ensure the consistency of the output and the accuracy of the quantization coefficient; if the output is found to be inconsistent, the quantization coefficient is automatically adjusted until the output meets the consistency requirements, ensuring that the inference accuracy of the quantized model is close to that of the unquantized model.
[0030] The present invention also provides a system for implementing the above method. The system includes: a data preprocessing module, a model training module, a model conversion module, a network parsing module, a calibration and quantization module, and a model verification module;
[0031] Data preprocessing module: This module is responsible for scaling, normalizing, and interpolating the original image data, converting the image data into a distribution with a mean of 0 and a standard deviation of 1, and adjusting it to the target size. The output of this module is the standardized and adjusted image data, ready to enter the model training.
[0032] Model training module: The preprocessed data is input into the deep learning model for training. This module may include the configuration and training of a convolutional neural network, the optimization of the model structure (such as using grouped convolution or separable convolution), and generate a preliminary inference network model.
[0033] Model conversion module: This module converts the trained model into the ONNX format for cross-platform deployment. During the conversion process, the module also merges the batch normalization (BN) layer with the convolutional layer to simplify the model structure and improve the inference efficiency.
[0034] Network parsing module: Responsible for reading and analyzing the structure of the ONNX model, extracting the input-output relationships and weight parameters of the network, and marking the connection order between layers to prepare for subsequent optimization and quantization processes.
[0035] Calibration and quantization module: Calculates the KL divergence using the calibration image set, minimizes the error to determine the optimal threshold of the activation function. Subsequently, according to the quantization formula, parameters such as weights, biases, variances, and means of each layer are quantized to generate a low-bitwidth model suitable for FPGA inference.
[0036] Model verification module: Finally, this module verifies the quantized model to ensure that its inference output is consistent with the unquantized model, guaranteeing the efficient operation of the model on hardware.
[0037] The present invention also provides the above method for generating standard image data, the above quantization method for improving the accuracy of the neural network model, or the application of the above system in standard image generation for matching models, improving model inference effects, accelerating model training, reducing model quantization losses, etc.
[0038] The present invention also provides a hardware system for implementing the above method. The hardware system includes: a memory and a processor; a computer program is stored on the memory, and when the computer program is executed by the processor, the above method for generating standard image data and the quantization method for improving the accuracy of the neural network model using the standard image data are implemented.
[0039] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method for generating standard image data and the quantization method for improving the accuracy of the neural network model using the standard image data are implemented.
[0040] The beneficial effects of the present invention include: Existing symmetric range-based linear quantization coding techniques such as distiller need to recalculate the quantization coefficients online for each input image data, and the quantization coefficients are floating-point numbers. The former will result in relatively large accuracy losses, and the latter will lead to an increase in the storage and computing resource requirements of the FPGA. By performing targeted normalization processing on the original image data and combining multiple images to calculate the quantization coefficients offline, the detection performance of the quantized model is compared with that of directly quantizing without these processes. The average intersection over union on the test set differs by about 10 points. At the same time, replacing the floating-point numbers with the nearest power-of-two integers greatly accelerates the inference calculation of the FPGA. Brief Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 It is a flowchart of the method for the present invention to improve the accuracy of the neural network model using standard image data.
[0043] Figure 2 It is a schematic diagram of the detection result in the prior art (Solution A).
[0044] Figure 3 It is a schematic diagram of the detection result after generating standard image data in the present invention (Solution B). Detailed Embodiments
[0045] Combined with the following specific embodiments and drawings, the present invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and well-known common sense in the art, and the present invention has no special limitations.
[0046] Aiming at the problems in the prior art, the present invention first performs standardization processing on the original data, transforms the data into a distribution with a mean of 0 and a standard deviation of 1, then performs interpolation methods such as standardization and bilinear interpolation on the data to a specified size and then sends it into a deep learning network for training, then converts the trained model into an onnx model, then reads the onnx model and merges the Conv layer and BN layer in the network, analyzes the network connection sequence and weight and other parameter information of the onnx model, then uses the calibration image to calculate the KL divergence to obtain the activation function threshold, and finally calculates the quantization coefficient according to the quantization formula from the weight, bias and other information of each layer of the network.
[0047] Specifically,
[0048] The present invention provides a new image data preprocessing process. Inappropriate image preprocessing for a deep learning model can lead to precision loss during the model quantization process, resulting in a significant decrease in the model's precision-recall rate. A new image data preprocessing technique is proposed to reduce the precision loss of model quantization caused by the need to adapt to the inference requirements of the edge-side model, so that the precision-recall rate during model inference remains stable and does not decrease. And the fact that the precision-recall rate remains basically unchanged after model quantization is the key to achieving the actual application effect of using the model for downstream tasks. For example, in the autonomous driving perception module, if the detection box of the obstacle detection model is too large or too small, it will cause a large error in the estimation of the actual boundary area of the obstacle, resulting in an incorrect division of the drivable area boundary, and further affecting the incorrect output of the final control and decision-making, increasing the risk of collision.
[0049] In a specific embodiment, the entire method process of the present invention is as follows: First, perform standardization processing on the original image data to transform the data into a data distribution with a mean of 0 and a standard deviation of 1. Then, perform standardization and bilinear interpolation on the data to the required size, and then send it into a deep learning network for model training. Convert the trained model into an onnx model, read the onnx model and merge the BN network layer in the model with the previous Conv network layer. Sequentially read the input layer and output layer of the network and number them in sequence. Use the calibration image to calculate the KL divergence, obtain the threshold of the activation function according to the minimization of the KL divergence, and then obtain the quantization coefficient according to the quantization formula using the weight information, bias information, variance information, mean information, etc. of each network layer. Finally, verify whether the quantization coefficient meets the requirement that the output of the previous layer is equal to the input of the current layer.
[0050] The differences between this quantization technique and the prior art include the image data preprocessing process, which is specifically as follows:
[0051] The image data preprocessing process first scales the input image data to the interval [0,1] to improve the convergence speed of the model, and then normalizes it according to the formula to a distribution with a mean of 0 and a variance of 1, where x is the pixel value of the input image, x mean is the average value of the pixels of the input image, and x σ is the variance of the pixels of the input image, where ∈ is the standard normal distribution, t is the uniform distribution on the interval [0,1], and π(t) is a monotonically increasing function from 0 to 1. Then, perform bilinear interpolation on y t to obtain the image data z of the target size.
[0052] After generating the standard image data in the method of the present invention, according to the existing related technologies for model training, the image data z is sent into a deep learning model for training, the network model structure is modified, and grouped convolution and / or separable convolution are used to replace the original convolution operator for model training.
[0053] In a specific embodiment, taking the yolov5 model as an example, for the convolution operator therein, grouped convolution and / or separable convolution are used to replace the operator. Note that only the convolution operator with a channel number ≥ 256 is replaced here, and the 7*7 pooling layer operator is replaced with multiple 3*3 pooling operators. After the replacement is completed, model training is carried out.
[0054] As Figure 1 shown, the specific steps implemented by the flowchart of the quantization method for improving the accuracy of the neural network model using the standard image data of the present invention are as follows:
[0055] Step 1, input the training image, divide the data value of the image by 256, and scale it to the interval [0,1]. The purpose of data normalization is to accelerate the training convergence of the model, and at the same time, it is convenient for the FPGA to use shifting to replace division. However, normalization will affect the original distribution form of the data. Therefore, the present invention performs the operation of Step 2 on the basis of Step 1 to obtain the standardized data y t while trying to keep the original data distribution intact as much as possible while accelerating the training convergence of the model. Step 1 belongs to the well-known technology;
[0056] Step 2, standardize the data obtained in Step 1 according to the formula to a distribution with a mean of 0 and a variance of 1, where x is the pixel value of the input image, x mean is the average value of the pixels of the input image, and x σ is the variance of the pixels of the input image. Then, according to the formula the standardized data y t is obtained. Step 2 belongs to the non-well-known technology;
[0057] Among them, y is the value after the image data is standardized, ∈ is the standard normal distribution, t is the uniform distribution in the interval [0,1], and π(t) is a monotonically increasing function from 0 to 1;
[0058] Step 3, perform interpolation methods such as bilinear interpolation and / or trilinear interpolation and / or nearest neighbor interpolation on the y t obtained in Step 2 to obtain the image data z with the target size. Step 3 belongs to the well-known technology;
[0059] Step 4: Feed the preprocessed data into a deep learning model for model training to obtain the required inference network model. Step 4 belongs to the prior art; the overall operation of using the images generated according to Steps 1 to 3 as the input data for the training model belongs to non-prior art; feeding the data generated in Steps 1 to 3 into the deep learning model for training also belongs to non-prior art;
[0060] Step 5: Convert the model obtained in Step 4 into an onnx model. Step 5 belongs to the prior art;
[0061] Step 6: Read the onnx model in Step 5, and merge the BN network layer in the onnx model with the upper Conv network layer of this network layer. Step 6 belongs to the prior art;
[0062] Step 7: Parse the merged model, read the input layer and output layer of the network in sequence and label them in sequence. Step 7 belongs to the prior art;
[0063] Step 8: Use the calibration image to calculate the quantization coefficient by KL divergence, and obtain the threshold T of the activation function according to the minimized KL divergence. For example, for the output y of a certain layer of the network, its q layer The calculation formula for the quantization coefficient is n is the bit width of the quantized value, and the quantization coefficient of the weight The quantization coefficient of the bias where T is obtained according to the minimized KL divergence. Step 8 belongs to the prior art;
[0064] Step 9: Calculate the quantization results of each layer of the network according to the quantization formula to obtain the output result of the quantized model, where x f is the input data, is the quantization coefficient. Step 9 belongs to the prior art;
[0065] The key point of the present invention is the preprocessing process of the image data, and the points to be protected are two. The first point is: the first point is the process of obtaining the preprocessed data y t ; the second point is the entire quantization model process from the image preprocessing data to the model training quantization.
[0066] Example 1
[0067] In this example, it is mainly necessary to detect the cars in the original image. Use the existing image preprocessing technology to scale the original image (1920×1080) to 640*360, and then fill in the black edges to a size of 640*512 as the input of the quantization model, which is denoted as Scheme A here. The result is as Figure 2 shown, and use the image preprocessing technology proposed in this article as the input, which is denoted as Scheme B here. The result is asFigure 3 As shown, after the same model quantization, the model inference in the FPGA obtains the following detection results respectively. It can be seen that the preprocessing technology proposed by the present invention makes the detection boxes of the inference results more accurate. mAP is the mean average precision. In the method of the present invention, this value is significantly higher than that of the images processed by the existing image preprocessing technology methods.
[0068]
[0069]
[0070] Embodiment 2
[0071] This embodiment mainly detects multiple targets in the original image, including cars, buses, pedestrians, tricycles, bicycles, etc. Compared with the existing direct quantization scheme, the scheme of the present invention has obvious advantages in both the mean average precision and the comparison of average time-consuming data.
[0072] Direct quantization The quantization method of the present invention quantizes Category Floating point (%) Fixed point (%) Fixed point (%) Car 81.0 80.1 80.6 Bus 83.2 82 90 Pedestrian 84.1 82.1 83.3 Tricycle 72.3 69.1 70.5 Bicycle 45.5 44.4 45.0 mAP value 53.20 51.7 53.9 Average time-consuming comparison 300ms 30ms 21ms
[0073] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The solutions in the embodiments of the present application can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0074] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0075] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one or more of the procedures Figure 1 or more procedures and / or blocks Figure 1 or more blocks.
[0076] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the procedures Figure 1 or more procedures and / or blocks Figure 1 or more blocks.
[0077] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0078] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
[0079] The protection scope of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept of the present invention, all changes and advantages that can be conceived by those skilled in the art are included in the present invention, and the appended claims are used as the protection scope.
Claims
1. A method for generating standard image data, characterized in that, The method obtains standard image data by scaling and normalizing the input image data and then performing interpolation processing, including the following steps: Step 1: Input an image and scale the image data to the interval [0, 1]; Step 2. Standardize and process the scaled data obtained in Step 1 to obtain transformed data y t ; Where y is the value after normalizing the image data, ∈ is the standard normal distribution, t is the uniform distribution in the interval [0, 1], and π(t) is a monotonically increasing function from 0 to 1; the monotonically increasing function is a linear function, a logarithmic function, a Sigmoid function, or a higher-order function; Step 3: Perform interpolation processing on the conversion data obtained in Step 2 to obtain image data z with the target size.
2. The method according to claim 1, wherein In Step 1, the pixel values in the input image are divided by 256 one by one to scale the image data; and / or In step two, the scaled data obtained in step one is passed through normalized to a distribution with a mean of 0 and a variance of 1; where x is the pixel value of the input image, and x mean is the average of the pixels of the input image, and x σ is the variance of the pixels of the input image; and / or In Step 3, the interpolation processing includes at least one of bilinear interpolation, trilinear interpolation, and nearest neighbor interpolation to match the image size with the model input.
3. A quantization method for improving the accuracy of a neural network model using standard image data, characterized in that, The method includes the following steps: Step I. Obtain standard image data according to the method described in Claim 1 or 2; Step II. Input the standard image data into a deep learning model for training to obtain an inference network model; Step III. Convert the inference network model into an ONNX model and merge the batch normalization layer and the convolutional layer in the ONNX model; Step IV. Use a calibration image to calculate the quantization coefficient by KL divergence to obtain the threshold T of the activation function; Step V. Calculate the quantization coefficient based on the quantization formula using the information including weights, biases, variances, and means of each network layer, and verify the accuracy of the quantization coefficient.
4. The quantization method according to claim 3, wherein The standard image data adapts to the input size of the deep learning model; and / or In Step II, the deep learning model is a neural network model suitable for image processing and is used for inference on FPGA; during the training process, grouped convolution or separable convolution is used to replace the original convolution operator; after training, it also includes the steps of cross-validation and testing of the model; and / or In Step III, export the inference network model from its original framework to an ONNX format file and perform compatibility adjustment; extract the trained parameters of the inference network model from the batch normalization layer, recalculate the weights and biases of the convolutional layer, so that the convolution operation directly includes the effect of batch normalization, and replace the model originally including the batch normalization layer and the convolutional layer with a model only including the merged convolutional layer; and / or In Step IV, use a calibration image to generate an output distribution through the network layer, perform quantization processing on the output distribution, calculate the KL divergence between the quantized distribution and the original floating-point distribution, so that the KL divergence reaches the minimum value, and determine the threshold T of the activation function; and / or In step V, according to the quantization formula x f is the input data, n is the bit width of the quantized value, is the quantization coefficient, and the quantization coefficient is obtained by using the weight information, bias information, variance information, and mean information of each network layer.
5. The quantization method according to claim 4, wherein When exporting the inference network model from its original framework to an ONNX format file, convert and save the architecture and weight information of the model; make the model structure adapt to the requirements of the ONNX format through compatibility adjustment.
6. The quantization method according to claim 3, wherein After Step V, it also includes a quantization self-check step: compare the quantized output of each layer of the network with the output of its previous layer to ensure the consistency of the output and the accuracy of the quantization coefficient; If output inconsistencies are detected, automatically adjust the quantization coefficients until the output meets the consistency requirements, ensuring that the inference accuracy of the quantized model is close to that of the unquantized model.
7. A system for implementing the method according to claim 1 or 2, or the quantization method according to any one of claims 3-6, characterized in that, The system includes: a data preprocessing module, a model training module, a model conversion module, a network parsing module, a calibration and quantization module, and a model verification module; The data preprocessing module is responsible for performing scaling, normalization, and linear interpolation on the original image data; The model training module inputs the preprocessed data into a deep learning model for training to generate a preliminary inference network model; The model conversion module converts the trained model into the ONNX format and merges the batch normalization layer and the convolutional layer; The network parsing module is responsible for reading and analyzing the structure of the ONNX model, extracting the input-output relationship and weight parameters of the network, and marking the connection order between layers; The calibration and quantization module calculates the KL divergence using the calibration image set, determines the optimal threshold of the activation function by minimizing the error; quantizes the weights, biases, variances, and means of each layer according to the quantization formula to generate a low-bitwidth model suitable for FPGA inference; The model verification module verifies the quantized model.
8. A hardware system for implementing the method according to any one of claims 1-6, characterized in that, The hardware system includes: a memory and a processor; a computer program is stored on the memory, and when the computer program is executed by the processor, the method described in any one of claims 1-6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Quantification method and system for optimizing int8
CN111950716A
A method to improve the accuracy of low-bit neural network models through weight preprocessing
CN113762494B
Threshold sensing low-order quantization method based on medical hyperspectral detection network
CN115700806A
Convolutional neural network hybrid computing post-training quantization algorithm for embedded system
CN116341639A