A CNN Custom Network Quantization Acceleration Method for FPGA
By building a lightweight neural network model and combining the improved PACT and CEOCO methods, the activation value and weight value are quantized, which solves the storage and resource consumption problems of convolutional neural network inference on FPGA, and achieves efficient quantitative inference acceleration.
Patent Information
- Application Number
- CN202311200741.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-09-18
AI Technical Summary
The prior art faces the problems of large storage space requirements, high resource consumption and accuracy losses when using FPGA accelerated convolutional neural network inference. Especially in complex network structures, many operators are not quantized.
A lightweight neural network model is built, and the activation value and weight value are quantized using improved PACT method and perceptual quantization technology. The CEOCO performs network retraining and fusion operations through convolutional continuous execution method, and deploys it to ARM+FPGA heterogeneous chip for quantitative inference acceleration.
Without significantly losing accuracy, the storage space requirements and resource consumption of the network model are reduced, the inference speed of edge smart terminals is accelerated, and the network quantization acceleration effect is improved.
Smart Images

Figure CN117151178B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network compression, and specifically relates to a method for quantifying and accelerating a CNN customized network for FPGA. Background Art
[0002] Convolutional neural network is one of the most commonly used network models in artificial intelligence applications, and is widely used in computer vision, natural language processing, embedded systems and other aspects. Current networks improve the network's ability to learn features through complex model structures and continuously stacked convolutional layers. Their complex computational operators and huge number of parameters bring heavy storage and computational pressure to hardware devices with limited resources. With the increasing popularity of edge intelligent terminals based on ARM+FPGA, realizing local inference poses higher requirements for model lightweighting. In this context, many researchers have proposed methods such as approximation, quantization, and pruning to compress the network volume. Compression technology reduces the model size by removing redundancy and irrelevance, and plays a crucial role in reducing memory bandwidth.
[0003] As one of the key technologies for network compression, quantization converts weight parameters from 32-bit floating-point precision to 8 bits or lower, reducing the computational intensity, parameters, and memory consumption of the network. However, network quantization often leads to a loss of accuracy. How to reduce the network size without significantly sacrificing accuracy has become a research hotspot in this field. To compensate for the quantization loss, DoReFa-Net proposed a method using bit convolutional kernels for training and inference. Choi et al. introduced the PACT method to optimize activation quantization during training without significant performance degradation. Post-training quantization quantizes the trained network parameters to low precision using part of the validation data to achieve network compression. However, the errors introduced by these approximations accumulate during the forward propagation calculation process, resulting in a significant performance decline. Jacob proposed the quantization-aware training method, which injects a quantizer into the network graph during training, quantizes the network parameters, and uses straight-through estimation to approximate the gradient. However, most quantization-aware training methods only inject the quantizer before the convolutional operation, and in more complex network structures, many other operators remain unquantized.
[0004] During the training process of convolutional neural networks, it is usually necessary to use a graphics processing unit (GPU) to provide support. However, after training, using a GPU during the network inference phase increases the overall system performance and resource consumption. To efficiently deploy and execute the network, Kotlar et al. studied various deployment locations and suitable underlying hardware architectures, including multi-core processors, many-core processors, field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs). Due to the low power consumption and flexible and configurable hardware resources of FPGAs, using FPGAs for deployment and inference has gradually become a new research hotspot. However, deploying CNNs on FPGAs also faces some challenges. Traditional networks such as Vgg16 have 138 million parameters, which require 255 MB of storage space when stored in 32-bit format. Transferring these values to / from off-chip memory also incurs performance and energy overhead. In addition, due to the unique structure and limited resources of FPGAs, numerical adjustment of the network's inference deployment is required. Summary of the Invention
[0005] To solve a series of problems faced when using FPGAs to accelerate convolutional neural network inference, the following measures are mainly implemented in this paper:
[0006] S1: Construct a neural network and train a full-precision neural network model;
[0007] S2: Introduce the improved PACT method and use the perceptual quantization method to quantize the model;
[0008] S3: According to the different network model structures, adopt the convolution continuous execution method CEOCO to quantize specific non-convolution operators, and perform a fusion operation after obtaining the quantized model;
[0009] S4: Deploy lightweight CNNs to the target hardware FPGA for inference acceleration verification;
[0010] Preferably, the process of constructing and training the neural network includes: constructing a lightweight neural network model, preprocessing the image classification dataset, dividing it into a training set and a validation set, and randomly flipping, adjusting the brightness, and randomly cropping the image data to a unified size. Input the processed pictures into the neural network for full-precision model training.
[0011] Preferably, the improved PACT method is introduced, and the process of quantifying the model using the perceptual quantization method includes: analyzing and improving the PACT method, and improving the values of the activation values less than 0 in the PACT method; performing quantization-aware training on the neural network, and using the quantizer to preprocess the activation values with PACT during forward propagation, and then performing layer-by-layer symmetric quantization and dequantization. For the weight values, more fine-grained per-channel symmetric quantization and dequantization are used. By quantifying and dequantifying the values to introduce the error brought by the quantization values, during the convolution calculation, the dequantized floating-point values are used, so that the neural network can learn the error brought by the quantization. During backpropagation, since the round-to-nearest function is used, this will cause all gradient calculations to be 0. The straight-through estimator STE is used to solve this problem, that is, directly skip the quantization calculation formula and pass the value to the upper layer for calculation.
[0012] Furthermore, the formula of the improved PACT method is:
[0013]
[0014] Among them, x is the input activation value, y is the value truncated by the PACT function, α and β are two trainable parameters, the initial value of α is set to 20, and the initial value of β is set to 3. During the quantization training of the neural network, the activation values are truncated by changing the values of α and β to remove outliers, so that the values to be mapped are compact.
[0015] Furthermore, the formula for per-channel symmetric quantization of weights is:
[0016]
[0017]
[0018] Among them, W i is the weight of the i-th convolution kernel in the convolutional layer, is the weight value after quantization and dequantization by Q(·), is the scaling factor of this weight. Since it is symmetrically quantized to 8 bits, the bias z is always equal to 0, and the truncated function clamp is used to truncate the quantized value to (-127, 127). round is the round-to-nearst function, that is, the rounding function, and abs and max are the absolute value and maximum value of the tensor respectively.
[0019] Furthermore, the formula for layer-by-layer symmetric quantization of the feature map / activation value is:
[0020]
[0021] Among them, X is the input floating-point activation value, which is the value after quantization and dequantization. The difference lies in the calculation formula of the scaling factor for the activation value. The scaling factor is calculated by sampling the numerical MSE using the moving average absolute maximum sampling strategy:
[0022] moving_avg_max = moving_avg_max * β + max(abs(X)) * (1 - β)
[0023]
[0024] where moving_avg_max is the absolute maximum of the average value, β is the momentum of moving_avg_max, with an initial value of 0.9, and β will change according to training. Adopting the moving average absolute maximum sampling strategy can reduce the sensitivity of the model to noise and redundant information, thereby improving the generalization ability of the model.
[0025] Preferably, according to the different network model structures, the convolution continuous execution method CEOCO is used to quantize specific non-convolution operators and perform subgraph fusion. The main process includes: analyzing the network model structure obtained through traditional perceptual quantization training, and analyzing whether the input and output data types of the convolution operator are both INT8. If not, a pseudo-quantization node needs to be inserted in front of the operator for quantization and dequantization operations, and then through retraining, these values can be converted to the INT8 type during quantization inference, so that the convolution operator can be continuously executed on the FPGA chip. The neural network framework will form a subgraph of operators that can be continuously executed on the target hardware. Therefore, before deployment, the convolution operators can be combined into a subgraph operator.
[0026] Preferably, the quantization inference formula is:
[0027]
[0028] After arrangement, we get
[0029]
[0030] where S can be obtained during quantization training X , S Y . Assume M = 2 -n M o , where M is a floating-point number, and M o is a fixed-point number. During inference quantization, M can be replaced by M o using bit shift. At this time, the data types of the formula are all fixed-point, and full INT8 fixed-point inference on the chip can be completed.
[0031] The beneficial effects of the present invention are as follows: The present invention compresses the neural network based on the fixed-point scalar quantization technology. Specifically, a lightweight neural network is constructed and quantization-aware training is carried out. During the forward propagation of the training, the improved parametric clipped activation (PACT) is used to preprocess the activation values, and outliers are removed through range truncation to make the calculation of the scaling factor more accurate. Then, the convolutional execution-only continuous optimization (CEOCO) method is used to retrain the network. After the retraining is completed, various fusion operations are performed on the network to reduce the resource consumption of the network model for reading data from the register during the deployment phase. Finally, the fused network is deployed on the ARM+FPGA heterogeneous chip Xilinx Zynq Ultrascale+MPSoC 3EG for quantization inference acceleration. On the premise that the accuracy loss of the neural network is acceptable, the present invention accelerates the inference speed of the neural network on the edge intelligent terminal and has a better network quantization acceleration effect than the prior art. Description of the Drawings
[0032] Figure 1 is the flowchart of the present invention;
[0033] Figure 2 is the flowchart of the customized network quantization method proposed in the present invention;
[0034] Figure 3 is the schematic diagram of the deployment of the network operator of the present invention on the heterogeneous chip, Figure 3 (a) is the network structure composed of operators, Figure 3 (b) is the schematic diagram of the frequent platform switching when the network is deployed on the heterogeneous chip for inference;
[0035] Figure 4 is the schematic diagram of the CEOCO strategy proposed in the present invention, Figure 4 (a) is the original network model structure diagram, Figure 4 (b) is the network model structure diagram after adding pseudo-quantization nodes, Figure 4 (c) is the network model structure diagram if FPGAs support other non-convolution operators, Figure 4 (d), Figure 4 (e) and Figure 4 (f) are respectively Figure 4 (a), Figure 4 (b) and Figure 4 (c) after the subgraph fusion network model structure diagrams;
[0036] Figure 5 is the convergence accuracy schematic diagram of the traditional quantization method and the quantization algorithm using the present invention in the present invention, Figure 5 (a), Figure 5 (b), Figure 5 (c), Figure 5(d) are the comparison graphs of the convergence of the Top-1 accuracy of the quantization algorithm of the present invention and the traditional quantization method for the four networks of mobilenetv1, mobilenetv3, pplcnet, and pplcnetv2 on the validation set. Detailed implementation manners
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0038] The present invention proposes a quantization acceleration method for a CNN custom network for FPGA, as Figure 1 and Figure 2 shown. The method includes the following: constructing and performing quantization-aware training on a lightweight neural network, and preprocessing the activation values using an improved parametric clipped activation (PACT) during the forward propagation of the training; retraining the network using the convolution execution optimization (CEOCO) method, and after the retraining is completed, performing various fusion operations on the network, and deploying the fused network to the ARM+FPGA heterogeneous chip Xilinx Zynq Ultrascale+MPSoC 3EG for quantization inference acceleration;
[0039] S1: Construct a neural network and train a full-precision neural network model;
[0040] The process of constructing and training the neural network includes: constructing a lightweight neural network model, preprocessing the image classification data set, dividing it into a training set and a validation set, and randomly flipping, adjusting the brightness, and randomly cropping the image data to a unified size (3×224×224), and inputting the processed pictures into the neural network for full-precision model training.
[0041] S2: Introduce the improved PACT method and use the quantization-aware method to quantize the model;
[0042] The operation process includes: analyzing and improving the PACT method, and improving the values of the PACT method with activation values less than 0; performing quantization-aware training on the neural network. During forward propagation, the quantizer is used to first preprocess the activation values with PACT and then perform layer-by-layer symmetric quantization and dequantization. For the weight values, more fine-grained per-channel symmetric quantization and dequantization are adopted. By quantizing and dequantizing the values, the error brought by the quantization values is introduced. When performing convolution calculations, the dequantized floating-point values are used, so that the neural network can learn the error caused by quantization. During backpropagation, since the round-to-nearest function is used, this will cause all gradient calculations to be 0. The straight-through estimator (STE) is used to solve this problem, that is, directly skip the quantization calculation formula and pass the value to the upper layer for calculation. The formula for the improved PACT method is:
[0043]
[0044] Among them, x is the input activation value, y is the value after truncation using the PACT function, α and β are two trainable parameters. The initial value of α is set to 20, and the initial value of β is set to 3. During the quantization training of the neural network, the activation values are truncated by changing the values of α and β to remove outliers, so that the values to be mapped are compact. The formula for per-channel symmetric quantization of weights is:
[0045]
[0046]
[0047] Among them, W i is the weight of the i-th convolution kernel in the convolutional layer, is the weight value after quantization and dequantization through Q(·), is the scaling factor of this weight. Since it is symmetrically quantized to 8 bits, the bias z is always equal to 0, and the truncation function clamp is used to truncate the quantized value to (-127, 127). round is the round-to-nearst function, that is, the rounding function, and abs and max are the absolute value and maximum value of the tensor respectively. The formula for layer-by-layer symmetric quantization of the feature map / activation value is:
[0048]
[0049] Among them, X is the input floating-point activation value, is the value after quantization and dequantization. The difference is that for the calculation formula of the scaling factor of the activation value, the scaling factor is calculated by sampling the mean squared error (MSE) of the values sampled by the moving average absolute maximum sampling strategy:
[0050] moving_avg_max = moving_avg_max * β + max(abs(X)) * (1 - β)
[0051]
[0052] Where moving_avg_max is the absolute maximum value of the average, β is the momentum of moving_avg_max, with an initial value of 0.9, and β will change according to training. The sampling strategy of using the moving average absolute maximum value can reduce the sensitivity of the model to noise and redundant information, thereby improving the generalization ability of the model.
[0053] S3: According to the different network model structures, use the convolution continuous execution method CEOCO to quantize specific non-convolution operators, and perform a fusion operation after obtaining the quantized model;
[0054] Figure 3 This is a schematic diagram of the deployment of network operators in the heterogeneous chip in the present invention, Figure 3 (a) is the network structure composed of operators, Figure 3 (b) is a schematic diagram of the frequent switching of platforms when the network is deployed on the heterogeneous chip for inference;
[0055] Figure 4 This is a schematic diagram of the CEOCO strategy proposed in the present invention, Figure 4 (a) is the original network model structure diagram, Figure 4 (b) is the network model structure diagram after adding pseudo-quantization nodes, Figure 4 (c) is the network model structure diagram if FPGAs support other non-convolution operators, Figure 4 (d), Figure 4 (e) and Figure 4 (f) are respectively Figure 4 (a), Figure 4 (b) and Figure 4 (c) after sub-graph fusion of the network model structure diagram;
[0056] Figure 5 This is a schematic diagram of the convergence accuracy of the traditional quantization method and the quantization algorithm using the present invention in the present invention, Figure 5 (a), Figure 5 (b), Figure 5 (c), Figure 5 (d) are respectively the comparison diagrams of the quantization algorithm of the present invention and the traditional quantization method in terms of the Top-1 accuracy convergence on the validation set for the four networks of mobilenetv1, mobilenetv3, pplcnet, and pplcnetv2.
[0057] The main process includes: analyzing the network model structure obtained through traditional perceptual quantization training, and analyzing whether the input and output data types of the convolution operator are both INT8. As Figure 3 shown, if it is not INT8, it is necessary to insert a pseudo-quantization node in front of its operator for quantization and de-quantization operations, as Figure 4 shown. Then, through retraining, these values can be converted to the INT8 type during quantization inference, so that the convolution operator can be continuously executed on the FPGA chip. The neural network framework will form a subgraph with operators that can be continuously executed on the target hardware. Therefore, the convolution operators can be combined into a subgraph operator before deployment.
[0058] S4: Deploy lightweight CNNs to the target hardware FPGA for inference acceleration verification;
[0059] Deploy the network model to the Xilinx Zynq Ultrascale+ MPSoC 3EG chip for quantization inference. The formula is:
[0060]
[0061] After arrangement:
[0062]
[0063] Among them, S can be obtained during quantization training X , S Y . Assume M = 2 -n M o , where M is a floating-point number, and M o is a fixed-point number. During inference quantization, M is replaced by M o using bit shift.
[0064] Table 1 shows the accuracy comparison and model size comparison between the full-precision floating-point model and the optimized model in this paper on the validation set. Among them, *-F is the full-precision floating-point model, and *-O is the optimized model in this paper. It can be seen that the average accuracy loss of the optimized model in this paper drops by 1.2%, still meeting the accuracy recognition requirements for image classification. In addition, the volume of the optimized model in this paper is compressed by nearly a quarter, reducing the model volume size, thus facilitating the deployment of neural networks on resource-constrained smart terminals.
[0065] Table 1 Comparison between the full-precision model and the optimized model in this paper
[0066]
[0067] Table 2 shows the accuracy comparison between the traditional quantization model and the optimized model of this paper on the validation set, as well as the inference time comparison when deployed on the chip. Here, *-C represents the traditional quantization model, and *-O represents the optimized model of this paper. It can be seen that compared with the traditional quantization method and the optimized method of this paper, the accuracy of the model is not much different, but the inference speed of the optimized method of this paper has been improved. Figure 5 This is a schematic diagram of the convergence accuracy of the traditional quantization method and the quantization algorithm using the present invention in the present invention.
[0068] Comparison between the ordinary quantization model and the optimized model of this paper in Table 2
[0069]
[0070]
[0071] The above examples further elaborate on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above examples are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A CNN custom network quantization acceleration method for FPGA, characterized in that: The method includes the following steps: S1: Construct a neural network and train a full-precision neural network model; S2: Introduce the improved PACT method and use the perceptual quantization method to quantize the model, specifically including: analyzing and improving the PACT method, and improving the processing of numerical values with activation values less than 0 in the PACT method; performing quantization-aware training on the neural network. During forward propagation, use the quantizer to first perform PACT preprocessing on the activation values, and then perform layer-by-layer symmetric quantization and dequantization. For the weight values, use a finer-grained per-channel symmetric quantization and dequantization; introduce the error brought by the quantization values by quantizing and dequantizing the numerical values. When performing convolution calculations, use the dequantized floating-point numerical values to learn the error brought by quantization; during backpropagation, since the round-to-nearest function is used, this will cause all gradient calculations to be 0. Use the straight-through estimator (STE) to solve this problem, that is, directly skip the quantization calculation formula and pass the numerical values to the upper layer for calculation; The formula for the improved PACT method is: where x is the input activation value, y is the value truncated by the PACT function, α and β are two trainable parameters. Set the initial value of α to 20 and the initial value of β to 3. During the neural network quantization training, truncate the activation values by changing the values of α and β to remove outliers, so that the numerical values to be mapped are compact; The formula for per-channel symmetric quantization of weights is: Among them, W i is the weight of the i-th convolutional kernel of the convolutional layer, is the weight value after quantization and dequantization through Q(·), is the scaling factor of this weight. Since it is symmetrically quantized to 8 bits, the bias z is always equal to 0, and the truncated function clamp is used to truncate the quantized value to (-127, 127). round is the round-to-nearst function, that is, the rounding function, and abs and max are the functions to take the absolute value and the maximum value of the tensor respectively; The formula for layer-by-layer symmetric quantization of feature maps / activation values is: where X is the input floating-point activation value, is the value after quantization and dequantization; differently, for the calculation formula of the scaling factor of the activation value, the numerical MSE is sampled by the sampling strategy of moving average absolute maximum to calculate the scaling factor: moving_avg_max = moving_avg_max * β + max(abs(X)) * (1 - β) where moving_avg_max is the absolute maximum value of the average value, β is the momentum of moving_avg_max, with an initial value of 0.9, and β will change according to training. Adopt the sampling strategy of moving average absolute maximum value to reduce the sensitivity of the model to noise and redundant information and improve the generalization ability of the model; S3: According to the different network model structures, use the convolution continuous execution method (CEOCO) to quantize specific non-convolution operators, and perform a fusion operation after obtaining the quantized model; S4: Deploy lightweight CNNs to the target hardware FPGA for inference acceleration verification.
2. The CNN custom network quantization acceleration method for FPGA according to claim 1, characterized in that: The specific content of S1 includes: constructing a lightweight neural network model, preprocessing the image classification dataset, dividing it into a training set and a validation set, and randomly flipping, adjusting the brightness and randomly cropping the images to a unified size; inputting the processed images into the neural network for full-precision model training.
3. A CNN custom network quantization acceleration method for FPGA according to claim 1, characterized in that: The specific operations of S3 and S4 include: analyzing the network model structure obtained through traditional perception quantization training, determining whether the input and output data types of the convolution operator are both INT8. If not, inserting a pseudo-quantization node in front of the operator for quantization and de-quantization operations, and then converting these values to the INT8 type during quantization inference through retraining, enabling the convolution operator to be continuously executed on the FPGA chip; the neural network framework will form a subgraph of the operators that are continuously executed on the target hardware, and combine the convolution operators into a subgraph operator before deployment. The quantization inference formula is: After rearrangement, we get Among them, S is obtained during quantization training X , S Y Assumptions M=2 -n M o , where M is a floating point number, M o For fixed-point numbers, use bit shifting to replace M with M during inference quantization. o Instead, the data types of the formulas are all fixed-point, completing full INT8 fixed-point reasoning on the chip.
Citation Information
Patent Citations
Quantitative compression method of convolutional neural network
CN114118406A
Improved neural network hardware acceleration method and device based on FPGA
CN115564035A