Method for processing target detection model in quantitative perception training
By adopting post-trained quantized BN fusion strategy and quantized awareness training based on BN layer folding in the quantized perception training of deep neural networks, the problem of binarized weights being sensitive to parameter distribution changes is solved, the target detection accuracy is improved and efficient inference is supported on resource-constrained platforms.
Patent Information
- Application Number
- CN202411971456.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
In the quantitative perception training of deep neural networks, binarized weights are more sensitive to changes in parameter distribution, resulting in the performance of traditional BN fusion strategies in binarized models.
A processing method is proposed, through the quantized BN fusion strategy BQ-P after training and the quantized perception training strategy BNF-QAT based on BN layer folding, so that the final accuracy of the binarized weight of deep neural networks is not affected by the fusion of BN layer parameters, maintaining the improvement of the calculation speed while eliminating the adverse impact of BN layer fusion on task performance.
It improves the accuracy of object detection in image recognition, ensures that deep neural networks can perform inference tasks efficiently and quickly on resource-constrained hardware platforms, and is suitable for real-time calculation of intelligent perception algorithms on low-power devices.
Smart Images

Figure CN119942070A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and in particular to a method for processing a target detection model in quantization perception training. Background Art
[0002] With the rapid development of deep learning technology, neural network models have achieved remarkable results in fields such as image recognition, and the demand for deploying and using deep neural networks on various embedded and low-power devices has also arisen. However, deep models usually have a large number of parameters and complex calculation processes, which limits their deployment and application on resource-constrained platforms such as mobile devices and embedded devices. To solve this problem, quantizing the model to 1-bit is called binarization. Binarizing floating-point numbers to 1-bit greatly reduces storage space and can replace addition and multiplication calculations with XOR and bit 1 counts, thereby improving the calculation speed of inference.
[0003] DoReFa-Net is a quantization-aware training algorithm proposed by Megvii Technology in 2016. It is one of the mainstream algorithms for the deployment of neural network quantization. This method can quantize the weights, activations, and gradient values of the model at different bit widths, significantly reducing the model parameter storage requirements and computational complexity. In order to reduce the amount of calculation, the BN layer is usually fused with the convolutional layer, which will cause weight deviations and affect performance. In the scenario of model binarization, the binary weights are more sensitive to changes in parameter distribution. If the DoReFa-Net algorithm is used for traditional BN fusion, the distribution of binary weights will change, which will inevitably lead to significant performance degradation. To this end, it is necessary to adjust the BN fusion strategy to make it more suitable for binary quantization. Summary of the invention
[0004] The present invention provides a method for processing a target detection model in quantization perception training. Based on the target detection task in image recognition, the present invention uses the BN fusion strategy BQ-P of post-training quantization and the quantization perception training strategy BNF-QAT based on BN layer folding, so that the final accuracy of the binarization weights of the deep neural network will not be affected by the fusion of BN layer parameters, while maintaining the calculation speed improvement brought by BN fusion, eliminating the adverse effects of BN layer fusion on task performance, thereby improving the accuracy of target detection in image recognition.
[0005] In a first aspect, a method for processing a target detection model in quantization-aware training is provided, wherein the target detection model includes multiple convolution modules, and the convolution module performs the calculation of convolution layer + BN layer + activation function. The processing method is for any convolution module, and the processing method includes:
[0006] The single batch input to the convolution module is m input feature maps A = {a1, a2, ..., a m};
[0007] Binary quantize the weights of the convolution layer and calculate a convolution;
[0008] Calculate the mean and variance of the current batch of samples in the BN layer, and update the mean and variance of the sliding statistics of the BN layer based on the mean and variance of the current batch of samples;
[0009] Fuse the BN layer parameters into the convolution layer and calculate the Conv+BN fusion layer output Z;
[0010] The Conv output is scaled using the ratio of the sliding statistical variance to the current batch variance and added to the bias term. The output result is input to the activation function SiLU, which then outputs m output feature maps B = {b1, b2, ..., b m}.
[0011] In combination with the first aspect, in some implementations of the first aspect, the binary quantization weight is to quantize the target convolution layer weight to 1 bit, and the quantization method is:
[0012] W binary =sign(W conv )×E F (|W conv |)
[0013] Where W binary is a binary weight, W conv is a floating point weight, sign is a sign function; E F (|W conv |) is the scaling factor for the floating point weights.
[0014] In combination with the first aspect, in some implementations of the first aspect, the convolution layer Conv is calculated as follows:
[0015] Y=W binary A+bias conv
[0016] W binary is a binary weight, bias conv is the bias weight, Y={y1,y2,…,y m} is the output data of the convolution layer, y j is the jth sample of the output data.
[0017] In combination with the first aspect, in some implementations of the first aspect, the mean and variance of the current batch samples of the BN layer satisfy:
[0018]
[0019]
[0020] μ and σ 2 is the mean and variance of all input feature vectors of the current batch, m is the total number of input feature vectors of the current batch, and j represents the jth feature map of the current batch.
[0021] In combination with the first aspect, in some implementations of the first aspect, the mean μ of the sliding statistics of the BN layer is BN and variance σ 2 BN Updates meet:
[0022] μ BN =μ BN (1-υ)+μυ
[0023] σ 2 BN =σ 2 BN (1-υ)+σ 2 υ
[0024] μ and σ 2 is the mean and variance of all input feature vectors of the current batch, v∈(0,1).
[0025] In combination with the first aspect, in some implementations of the first aspect, the BN layer parameters are integrated into the convolution layer to satisfy:
[0026]
[0027] W fused-conv is the weight of the fused Conv+BN layer, bias fused-conv is the bias term of the fused Conv+BN layer; σ 2 BN is the sliding statistical variance of the BN layer, μ and σ 2 is the mean and variance of all input feature vectors in the current batch, γ and β are the scaling coefficient and translation coefficient respectively, and ∈ is a constant.
[0028] In combination with the first aspect, in some implementations of the first aspect, the Conv+BN fusion layer outputs Z, satisfying:
[0029] Z=sign(W fused-conv )×E F (|W fused-conv |)A
[0030] W fused-conv is the weight of the fused Conv+BN layer, sign is the sign function; E F (|W fused-conv|) is the scaling factor of the Conv+BN layer weights.
[0031] In combination with the first aspect, in some implementations of the first aspect, the Conv output is scaled using the ratio of the sliding statistical variance to the current batch variance and added to the bias term, satisfying:
[0032]
[0033] σ 2 BN is the sliding statistical variance of the BN layer, σ 2 is the variance of all input feature vectors in the current batch, ∈ is a constant, bias fused-conv It is the bias term of the fused Conv+BN layer.
[0034] In a second aspect, a processing method using a target detection model is provided, wherein the target detection model is trained by the processing method described in any one of the implementations of the first aspect, and the processing method using the target detection model includes:
[0035] Input a single input feature map X to the convolution module in a single batch;
[0036] Binary quantize weights of convolutional layers;
[0037] Get the sliding statistical mean μ determined after the target detection model training is completed BN and variance σ 2 BN ;
[0038] According to the sliding statistical mean μ BN and variance σ 2 BN Integrate the BN parameters into the convolutional layer;
[0039] Calculate the output Z of the Conv+BN fusion layer, input the output result Z into the activation function SiLU, and then output the output feature map B′=SiLu(Z).
[0040] In conjunction with the second aspect, in some implementations of the second aspect, is the weight of the fused Conv+BN layer, is the bias term of the fused Conv+BN layer; W conv is a floating point weight, bias conv is the bias weight, γ and β are the scaling coefficient and translation coefficient respectively, and ∈ is a constant.
[0041] Compared with the prior art, the solution provided by the present invention includes at least the following beneficial technical effects:
[0042] The present invention develops a neural network binarization method based on BN fusion quantization, which can support the binarized deep neural network to perform efficient and fast reasoning tasks on resource-constrained hardware platforms, and can be applied to the real-time calculation of intelligent perception algorithms on low-power devices. The present invention has been verified on a variety of data sets and different model architectures. Experiments have shown that the proposed method is suitable for networks and data sets of different complexities, and the accuracy loss caused by binarization on different network structures and data sets does not exceed 3%, which has high promotion value and reference significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 Schematic diagram of the inference strategy for binary network training.
[0044] Figure 2 This is a schematic diagram of the YOLOv5 series network architecture.
[0045] Figure 3 Schematic diagram of the main components of the YOLOv5 backbone network. DETAILED DESCRIPTION
[0046] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0047] The present invention provides a method for processing a target detection model. The target detection model includes multiple convolution modules. For any convolution module, during training, multiple input feature maps are input in a single batch, and the convolution module performs the calculation of convolution layer + batch normalization + activation function (Conv + BN + Activation), satisfying: output = Act (BN (Conv (X))). In the actual application of the target detection model, if a single input feature map is input in a single batch, the convolution module performs the calculation of convolution layer + activation function, satisfying: output = Act (Conv (X)).
[0048] The present invention first introduces the processing method of the target detection model during the training process, including the processing method for the BN layer statistical data in the quantization-aware training. The specific steps are as follows.
[0049] Step 1: Input a single batch of m input feature maps A = {a1, a2, ..., a m}.
[0050] Step 2: Binary quantize the weights and calculate a convolution.
[0051] The target convolution layer weight is quantized to 1 bit using the following method:
[0052] W binary =sign(W conv )×E F (|Wconv |)
[0053] Where W binary is a binary weight, W conv is a floating point weight, and sign is a sign function. F (|W conv |) is the scaling factor for the floating point weights.
[0054] m input feature maps A={a1,a2,…,a m}, the convolutional layer Conv is calculated as:
[0055] Y = Conv(X) binary =W binary A+bias conv
[0056] bias conv is the bias weight.
[0057] A is the convolutional layer input data with a batch size of m, a j is the jth sample of the input data, Y={y1,y2,…,y m} is the output data of the convolution layer, y j is the jth sample of the output data.
[0058] Step 3: Calculate the mean and variance of the current batch of samples in the BN layer:
[0059]
[0060] μ and σ 2 is the mean and variance of all input feature vectors of the current batch, m is the total number of input feature vectors of the current batch, and j represents the jth feature map of the current batch.
[0061] Step 4: Update the mean and variance of the sliding statistics of the BN layer
[0062] μ BN =μ BN (1-υ)+μυ
[0063] σ 2 BN =σ 2 BN (1-υ)+σ 2 υ
[0064] v∈(0,1)
[0065] That is, the mean μ of all input feature vectors of the current batch determined in step 3 and the sliding statistical mean μ corresponding to the previous batch BN , calculate μ BN(1-υ)+μυ and assign it to μ BN , μ after assignment BN Represents the sliding statistical mean corresponding to the current batch. Similarly, the variance σ of all input feature vectors of the current batch determined in step 3 is 2 , and the sliding statistical variance σ corresponding to the previous batch 2 BN , calculate σ 2 BN (1-υ)+σ 2 υ and assign it to σ 2 BN , after the assignment 2 BN Indicates the sliding statistical variance corresponding to the current batch.
[0066] Step 5: BN parameters are integrated into the convolutional layer:
[0067]
[0068] γ and β are learnable scaling and translation coefficients, respectively. ∈ = 0.001 is a small constant introduced to avoid the variance being zero. It should be noted that during the training process, the BN layer is calculated as follows: Therefore, bias fused-conv Use the current batch mean μ and variance σ 2 calculate.
[0069] Step 6: Calculate the Conv+BN fusion layer output Z, satisfying:
[0070] Z=Conv(X) binary-fused =sign(W fused-conv )×E F (|W fused-conv |)A
[0071] Step 7: Scale the Conv output using the ratio of the sliding statistical variance to the current batch variance and add it to the bias term to satisfy:
[0072]
[0073] Step 8: Calculate the activation function SiLU, satisfying:
[0074] B=SiLu(Z′)
[0075] The output is m output feature maps B = {b1, b2, ..., b m}.
[0076] The present invention further introduces a processing method of the target detection model in the reasoning process or the actual application process. The specific steps are as follows.
[0077] Step 1: Input a single input feature map X into the convolution module in a single batch.
[0078] Step 2: Binary quantization weights.
[0079] The target convolution layer weight is quantized to 1 bit using the following method:
[0080] W binary =sign(W conv )×E F (|W conv |)
[0081] Where W binary is a binary weight, W conv is a floating point weight, and sign is a sign function. F (|W conv |) is the scaling factor for the floating point weights.
[0082] Step 3: Obtain the sliding statistical mean μ determined after the target detection model training is completed BN and variance σ 2 BN .
[0083] Since the input data may not be input in batches during the inference process, the sliding mean μ saved during training needs to be used at this time. BN and variance And use this to estimate the mean and variance of the entire sample set. Here, v=0.1.
[0084] Step 4: According to the sliding statistical mean μ BN and variance σ 2 BN Determine the BN parameters and integrate them into the convolutional layer.
[0085] In order to reduce the amount of calculation and improve the calculation speed, the calculation of BN and the weight of the convolution layer can be fused. The fusion method is:
[0086]
[0087] Among them, Conv(X) fused Calculate the fused Conv+BN layer. is the weight of the fused Conv+BN layer, is the bias term of the fused Conv+BN layer; γ and β are the learnable scaling coefficient and translation coefficient, respectively, and ∈=0.001 is a small constant introduced to avoid the variance being zero.
[0088] Step 5: Calculate the Conv+BN fusion layer output Z, satisfying:
[0089] Z=Conv(X) binary-fused
[0090] =sign(W fused-conv )×E F (|W fused-conv |)X+bias fused-conv
[0091] Step 6: Calculate the activation function SiLU, satisfying:
[0092] B′=SiLu(Z)
[0093] To ensure the performance of binary training, the present invention uses the scaling factor of the original floating-point weight to represent the statistical information of the weight according to the idea in the DoReFa-Net quantization method. After the training is completed, the floating-point weight is fixed. If the traditional method of fusion of convolution layer and BN layer is used to fuse the BN parameters to the weight of the convolution layer, the distribution of the floating-point weight will be changed, and the distribution of the weight value will be changed during binarization, which will have a greater impact on the performance.
[0094] Since the binary weights after quantization-aware training are limited to {-1,1}, this type of weights is different from floating-point weights. They are very sensitive to the input of the previous layer and their own original distribution. Some slight disturbances may change the output of the entire network. Although a simple post-training quantization strategy can complete the task of fusing BN and convolutional layers, it will still change the input distribution learned by the network during training through the quantization scaling factor, which will have a certain impact on the final accuracy of the network.
[0095] In the present invention, BN is directly integrated into the weights of the convolutional layer during training. Figure 1 As shown in Figure 2, during the training phase, following the floating-point BN fusion method during inference, the sliding mean and variance of BN are first calculated and then fused into W. fused-conv In the process of binary quantization, the output after Conv and BN fusion is calculated; in the test phase, the inference process is consistent with the traditional method. There is no post-training quantization in this method, and the distribution of weight values will not be changed during binary quantization. Therefore, the present invention calls it BN Folded based Quantization Aware Training (BNF-QAT).
[0096] Figure 2It is the YOLOv5 series network architecture, which is mainly composed of three parts: Backbone, Neck and Head. The main components of YOLOv5 are the basic convolution (Conv+BN+SiLU, CBS) module, the cross-stage (Cross Stage Partial, CSP) module and the fast pyramid pooling (Spatial Pyramid Pooling-Fast, SPPF) module. Figure 3 The details of these modules are shown in . CBS is the main module in YOLOv5, which consists of a convolutional layer (Conv), batch normalization (BatchNorm) and a SiLU activation function. In the target detection task training phase, the present invention can be applied to the Conv convolutional layer and the BN layer in the CBS module to fuse, and provide a forward calculation strategy for the CBS module in the training phase. In order to avoid the impact of the fluctuation of the batch mean and variance on the training process, this strategy chooses to fuse the sliding statistics rather than the current batch statistics into the weight of the convolutional layer, and uses the ratio of the sliding statistics variance to the current batch variance to scale the convolution output for restoration after the convolution. Through this strategy of fusing BN during training, the present invention can eventually obtain a neural network composed of quantized convolutional layers that simulate the behavior of BN. No quantization operation is required before reasoning, and the data distribution disturbance caused by the quantization operation after training will not affect the final reasoning performance.
[0097] Although the present invention is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be based on the scope defined by the claims of the present invention.
Claims
1. A method for processing a target detection model in quantization-aware training, characterized in that: The target detection model includes multiple convolution modules, and the convolution modules perform the calculation of convolution layer + BN layer + activation function. The processing method is for any convolution module, and the processing method includes: The single batch input to the convolution module is m input feature maps A = {a1, a2, ..., a m }; Binary quantize the weights of the convolution layer and calculate a convolution; Calculate the mean and variance of the current batch of samples in the BN layer, and update the mean and variance of the sliding statistics of the BN layer based on the mean and variance of the current batch of samples; Fusion the BN layer parameters into the convolution layer and calculate the Conv+BN fusion layer output Z; The Conv output is scaled using the ratio of the sliding statistical variance to the current batch variance and added to the bias term. The output result is input to the activation function SiLU, which then outputs m output feature maps B = {b1, b2, ..., b m }.
2. The processing method according to claim 1, characterized in that: The binary quantization weight is to quantize the target convolution layer weight to 1 bit. The quantization method is: W binary =sign(W conv )×E F (|W conv |) Where W binary is a binary weight, W conv is a floating point weight, sign is a sign function; E F (|W conv |) is the scaling factor for the floating point weights.
3. The processing method according to claim 1, characterized in that: The convolution layer Conv is calculated as follows: Y=W binary A+bias conv W binary is a binary weight, bias conv is the bias weight, Y={y1,y2,…,y m } is the output data of the convolution layer, y j is the jth sample of the output data.
4. The processing method according to claim 1, characterized in that: The mean and variance of the current batch of samples in the BN layer satisfy: μ and σ 2 is the mean and variance of all input feature vectors of the current batch, m is the total number of input feature vectors of the current batch, and j represents the jth feature map of the current batch.
5. The processing method according to claim 1, characterized in that: The mean μ of the sliding statistics of the BN layer BN and variance σ 2 BN Updates meet: m BN =μ BN (1-u)+mu s 2 BN =s 2 BN (1-y)+s 2 u μ and σ 2 is the mean and variance of all input feature vectors of the current batch, v∈(0,1).
6. The processing method according to claim 1, characterized in that: The BN layer parameters are integrated into the convolution layer to meet the following requirements: W fused-conv is the weight of the fused Conv+BN layer, bias fused-conv is the bias term of the fused Conv+BN layer; σ 2 BN is the sliding statistical variance of the BN layer, μ and σ 2 is the mean and variance of all input feature vectors in the current batch, γ and β are the scaling coefficient and translation coefficient respectively, and ∈ is a constant.
7. The processing method according to claim 1, characterized in that: The Conv+BN fusion layer outputs Z, which satisfies: Z=sign(W fused-conv )×E F (|W fused-conv |)A W fused-conv is the weight of the fused Conv+BN layer, sign is the sign function; E F (|W fused-conv |) is the scaling factor of the Conv+BN layer weights.
8. The processing method according to claim 1, characterized in that: Use the ratio of the sliding statistical variance to the current batch variance to scale the Conv output and add it to the bias term to satisfy: σ 2 BN is the sliding statistical variance of the BN layer, σ 2 is the variance of all input feature vectors in the current batch, ∈ is a constant, bias fused-conv It is the bias term of the fused Conv+BN layer.
9. A processing method using a target detection model, characterized in that: The target detection model is obtained by training the processing method according to any one of claims 1 to 8, and the processing method using the target detection model includes: Input a single input feature map X to the convolution module in a single batch; Binary quantize weights of convolutional layers; Get the sliding statistical mean μ determined after the target detection model training is completed BN and variance σ 2 BN ; According to the sliding statistical mean μ BN and variance σ 2 BN Integrate the BN parameters into the convolutional layer; Calculate the output Z of the Conv+BN fusion layer, input the output result Z into the activation function SiLU, and then output the output feature map B′=SiLu(Z).
10. The processing method according to claim 9, characterized in that: is the weight of the fused Conv+BN layer, is the bias term of the fused Conv+BN layer; W conv is a floating point weight, bias conv is the bias weight, γ and β are the scaling coefficient and translation coefficient respectively, and ∈ is a constant.