Complex data set-oriented binarization target detection method
By adopting BNF-QAT and multi-grained scaling factor selection strategies in the YOLOv5m network, the detection accuracy of the binarized YOLOv5s model on complex data sets is solved, and the effect of efficient and fast object detection on complex data sets is achieved.
Patent Information
- Application Number
- CN202411971469.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
The binary YOLOv5s model has significantly reduced detection accuracy on complex data sets, making it difficult to meet the high requirements of target characteristics, noise and changes in complex data sets.
Based on the YOLOv5m network, the BNF-QAT based on BN layer folding is used to directly fuse the BN parameters into the weight of the convolution layer, and the network structure is adjusted to balance the calculation time and detection accuracy through the multi-grained scaling factor selection strategy.
It realizes the high detection accuracy on complex data sets while minimizing computing time. It is suitable for resource-constrained hardware platforms and supports efficient and fast object detection by lightweight deep neural networks on low-power devices.
Smart Images

Figure CN119942071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and in particular to a binary target detection method for complex data sets. Background Art
[0002] With the rapid development of deep learning technology, object detection algorithms are playing an increasingly important role in the field of computer vision. In particular, binary object detection algorithms have broad application prospects in various performance-constrained embedded and low-power devices. Currently, the YOLOv5s (You Only Look Once) algorithm has attracted widespread attention due to its advantages such as fast detection speed and high accuracy.
[0003] BNF-QAT, a quantization-aware training strategy based on BN layer folding, is a quantization-aware training method. By directly fusing BN parameters into the weights of the convolutional layer during binarization network training, the final accuracy of the binarization weights of the deep neural network will not be affected by the fusion of BN layer parameters. The binarized YOLOv5s model using BNF-QAT maintains the computational speed improvement brought by BN fusion while eliminating the adverse effects of binarization on target detection accuracy. The binarized YOLOv5s model based on BNF-QAT achieves good task performance on simple datasets.
[0004] However, the target characteristics, noise and changes in complex data sets place higher requirements on the binarization network. When the binary YOLOv5s model is applied to complex data sets, such as optical remote sensing data sets, it often leads to a significant decrease in the model detection accuracy. Summary of the invention
[0005] The present invention provides a binary target detection method for complex data sets. Based on the modification of the YOLOv5s network architecture, a multi-granularity scaling factor selection strategy is proposed to maximize the balance between calculation time and detection accuracy under the condition of limited device performance.
[0006] First, a binary target detection method for complex data sets is provided, which is applied to the YOLOv5m network target detection task, including:
[0007] In the first step, a quantization-aware training strategy BNF-QAT based on BN layer folding is adopted to directly integrate BN into the weights of the convolutional layer during training;
[0008] The second step is to select the granularity scaling factor;
[0009] The third step is to train the network according to the selected scaling factor granularity and the corresponding data set;
[0010] The fourth step is to verify the test accuracy of the network. If the accuracy does not meet the standard, replace one layer in the backbone network with a finer scaling factor according to the layer index order, and the scaling factors of other layers remain unchanged;
[0011] Step 5: Repeat steps 3 and 4 until the accuracy reaches the target.
[0012] In combination with the first aspect, in some implementations of the first aspect, in the YOLOv5m network, the quantization method of the binary weights of all convolutional layers Conv in the 1-bit layer is:
[0013] W binary =sign(W)×E(|W|)
[0014] The calculation method of the convolution layer Conv in the CBS block is:
[0015] Conv(X) binary =W binary A
[0016] Where W represents the floating point weight in the convolutional layer, W binary represents the binary weight after binary quantization, sign is the sign function, E(|W|) is the scaling factor of the convolutional layer, A={a1,a2,…,a i}, is i input feature maps.
[0017] In conjunction with the first aspect, in some implementations of the first aspect, the granularity scaling factor is calculated by layer-by-layer calculation, satisfying:
[0018]
[0019] Among them, |x1|…|x9| represents the weight on a convolution kernel, a=1,2,…,M represents the convolution kernel corresponding to the output channel, b=1,2,…,N represents the dimension of the convolution kernel corresponding to the input channel, c=1,2,…,9 represents the absolute value of the weight of the cth position of the bth channel dimension of the ath convolution kernel.
[0020] In combination with the first aspect, in some implementations of the first aspect, the granularity scaling factor is calculated by channel-by-channel calculation, satisfying:
[0021]
[0022]
[0023] in, represents the scaling factor of the convolution kernel of the first output channel, represents the scaling factor of the convolution kernel of the Mth output channel, E per-channel(|W|) represents the set of channel-by-channel scaling factors, and there are M scaling factors in a layer.
[0024] In combination with the first aspect, in some implementations of the first aspect, the granularity scaling factor is calculated by convolution kernel by convolution kernel, satisfying:
[0025]
[0026] in, Represents the scaling factor at the cth position of the scaling factor matrix of the convolution kernel for the ath input channel.
[0027] In combination with the first aspect, in some implementations of the first aspect, selecting a granularity scaling factor includes:
[0028] Before the formal training of the binary network, a single sample from different data sets is used to perform a training and forward process under each granularity scaling factor to obtain the forward reasoning calculation time corresponding to the three granularity scaling factors, and the granularity scaling factor is selected according to the calculation time limit.
[0029] In combination with the first aspect, in some implementations of the first aspect, the scaling factor is calculated as follows:
[0030]
[0031] Among them, W is the floating point weight, W binary is a binary weight, sign() is a sign function; scaling factor E F (|W|) is the mean of the absolute values of the weights before quantization.
[0032] In combination with the first aspect, in certain implementations of the first aspect, a channel-by-channel calculation mode is adopted for the SSDD data set scaling factor.
[0033] In combination with the first aspect, in some implementations of the first aspect, a channel-by-channel, convolution kernel-by-convolution kernel mixed calculation mode is adopted for the HRSC2016 dataset scaling factor.
[0034] In combination with the first aspect, in some implementations of the first aspect, the YOLOv5m network includes weights in some layers quantized to 8 bits, and weights in the remaining layers quantized to 1 bit; in the fourth step, if the accuracy is not up to standard, one layer of the 1-bit layers in the backbone network is replaced with a finer scaling factor according to the layer index order, and the scaling factors of other layers remain unchanged.
[0035] Compared with the prior art, the solution provided by the present invention includes at least the following beneficial technical effects:
[0036] The present invention develops a binary target detection method for complex data sets, which can support lightweight deep neural networks to perform efficient and fast target detection tasks on resource-constrained hardware platforms, and can be applied to real-time calculations of intelligent perception algorithms on low-power devices. The present invention has been verified on data sets of different complexity, and experiments have shown that the proposed method is applicable to data sets of different complexity, and has high promotion value and reference significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a schematic diagram of the YOLOv5m network architecture.
[0038] Figure 2 Schematic diagram of the convolution kernel used in a layer of convolution calculation.
[0039] Figure 3 Schematic diagram of the scaling factor results obtained under different granularity calculation methods. DETAILED DESCRIPTION
[0040] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0041] The present invention provides a binary target detection method for complex data sets, which is applied to the YOLOv5m network target detection task. The overall process is as follows.
[0042] In the first step, we use the quantization-aware training strategy BNF-QAT based on BN layer folding, which directly integrates BN into the weights of the convolutional layer during training, so that the network can cope with the changes in data distribution brought about by the integration of BN.
[0043] Figure 1 It is the YOLOv5 series network architecture, which mainly consists of three parts: Backbone, Neck and Head. The main components of YOLOv5 are the basic convolution (Conv+BN+SiLU, CBS) module, the cross-stage (Cross Stage Partial, CSP) module and the fast pyramid pooling (Spatial Pyramid Pooling-Fast, SPPF) module. Among them, CBS is the main module in YOLOv5, which includes convolution layer (Conv), batch normalization (BatchNorm) and SiLU activation function.
[0044] like Figure 1 As shown in the figure, the weights in some layers of the YOLOv5m network are quantized to 8 bits, the weights in the remaining layers are quantized to 1 bit, the activation values are uniformly quantized to 8 bits, and the BN fusion strategy is BNF-QAT.
[0045] The quantization method of the binary weights of all convolutional layers Conv in the layer quantized to 1 bit is:
[0046] W binary =sign(W)×E(|W|)
[0047] The calculation method of the convolution layer Conv in the CBS block is:
[0048] Conv(X) binary =W binary A
[0049] Where W represents the floating point weight in the convolutional layer, W binary represents the binary weight after binary quantization, sign is the sign function, E(|W|) is the scaling factor of the convolutional layer, A={a1,a2,…,a i}, is i input feature maps.
[0050] In the second step, before the formal training of the binarization network, a single sample from different data sets is used to perform a training and forward process under each granularity scaling factor (the same granularity scaling factor is used for all 1-bit layers) to obtain the forward reasoning calculation time corresponding to the three granularity scaling factors, and select the scaling factor with the finest granularity possible based on the calculation time limit.
[0051] In the process of binary network training, a scaling factor is set to count the information of weights before quantization in order to improve network performance. Generally, the scaling factor is calculated as follows:
[0052]
[0053] Among them, W is the floating point weight, W binary is a binary weight, sign() is a sign function. Scaling factor E F (|W|) is the mean of the absolute values of the weights before quantization.
[0054] The scaling factor can represent the information of the original weight to a certain extent, and then adjust the result of the binary convolution calculation, and at the same time, it can also make the network reach a convergence state faster and better. In the present invention, three granularity calculation methods are proposed for the scaling factor, namely, per-layer calculation, per-channel calculation, and per-kernel calculation. The granularity of these three calculation methods ranges from coarse to fine, and the richness of statistical information also increases with the refinement of the granularity, and the occupancy rate of the corresponding computing resources also gradually increases.
[0055] like Figure 2As shown in the figure, assuming that the number of input channels of a layer of convolution calculation is N, the number of output channels is M, and the size of the convolution kernel is K×K=3×3, there are M convolution kernels of size (K×K×N). First, the simplest process of calculating the scaling factor layer by layer (per-layer) is explained. The purpose of calculating the scaling factor layer by layer (per-layer) is to calculate a representative scaling factor for the convolution calculation of the current layer. The calculation method is:
[0056]
[0057] Among them, |x1|…|x9| represents the weight on a convolution kernel, a=1,2,…,M represents the convolution kernel corresponding to the output channel, b=1,2,…,N represents the dimension of the convolution kernel corresponding to the input channel, c=1,2,…,9 represents the absolute value of the weight of the cth position of the bth channel dimension of the ath convolution kernel.
[0058] This scaling factor is calculated in the simplest way. The average of all weights in all convolution kernels in the current layer is calculated once, and a scaling factor E is obtained for each layer. per-layer ,like Figure 3 As shown on the left, the granularity level is the coarsest, and the additional resources occupied by the scaling factor are O(1).
[0059] The second granularity statistical method is to calculate the scaling factor channel by channel. This calculation method requires calculating a scaling factor representative of the current output channel in each output channel. The calculation method is:
[0060]
[0061] in, represents the scaling factor of the convolution kernel of the first output channel, represents the scaling factor of the convolution kernel of the Mth output channel, E per-channel (|W|) represents the set of channel-by-channel scaling factors, such as Figure 3 As shown, there are a total of M scaling factors for one layer.
[0062] In the per-channel level statistics, the present invention needs to calculate the mean once for each output channel to obtain M different scaling factors. This statistical method has a moderate granularity and occupies O(M) additional resources.
[0063] Finally, the most granular per-convolution kernel calculation method requires that in each output channel, the mean of the elements at the corresponding positions of each input channel dimension of the convolution kernel be counted one by one to obtain a scaling factor matrix of size K×K. The size of the scaling factor matrix is the same as the size of the convolution kernel. The scaling factor at each position is calculated as follows:
[0064]
[0065] in, Represents the scaling factor at the cth position of the scaling factor matrix of the convolution kernel for the ath input channel.
[0066] like Figure 3 As shown on the right side of , in the per-kernel level statistics, the present invention calculates the scaling factor for the weights of the K×K positions of the N input channel dimensions of the convolution kernel in each output channel, and finally obtains M scaling factor matrices of size K×K, which contains detailed statistical information inside each convolution kernel in each output channel. This statistical method has the finest granularity and occupies the most additional resources, which is O(M×K×K).
[0067] The third step is to train the network based on the selected scaling factor granularity and the corresponding data set.
[0068] The fourth step is to verify the test accuracy of the network. If the accuracy does not meet the standard, one of the 1-bit layers in the backbone network is replaced with a finer scaling factor according to the layer index order, and the scaling factors of other layers remain unchanged.
[0069] Step 5: Repeat steps 3 and 4 until the accuracy reaches the target.
[0070] In the sixth step, forward reasoning calculations are performed on the embedded side to verify the recognition accuracy and computing performance of the deployed target recognition algorithm.
[0071] This embodiment takes the YOLOv5m network and the SSDD and HRSC2016 datasets as examples. The SSDD dataset is a radar imaging grayscale image, and the HRSC2016 dataset is an optical remote sensing image. The optical remote sensing dataset contains more information, the image composition is more complex, and the detection difficulty is higher. For this reason, the two datasets should select scaling factors of different granularities. According to the above process, the per-channel calculation mode is used for the scaling factor of the SSDD dataset, and the per-channel, per-kernel (3-layer) mixed calculation mode is used for the scaling factor of the HRSC2016 dataset.
[0072] Although the present invention is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be based on the scope defined by the claims of the present invention.
Claims
1. A binary target detection method for complex data sets, characterized in that: Applied to YOLOv5m network target detection tasks, including: In the first step, a quantization-aware training strategy BNF-QAT based on BN layer folding is adopted to directly integrate BN into the weights of the convolutional layer during training; The second step is to select the granularity scaling factor; The third step is to train the network according to the selected scaling factor granularity and the corresponding data set; The fourth step is to verify the test accuracy of the network. If the accuracy does not meet the standard, replace one layer in the backbone network with a finer scaling factor according to the layer index order, and the scaling factors of other layers remain unchanged; Step 5: Repeat steps 3 and 4 until the accuracy reaches the target.
2. The method according to claim 1, characterized in that In the YOLOv5m network, the binary weights of all convolutional layers Conv in the 1-bit layer are quantized as follows: W binary =sign(W)×E(|W|) The calculation method of the convolution layer Conv in the CBS block is: Conv(X) binary =W binary A Where W represents the floating point weight in the convolutional layer, W binary represents the binary weight after binary quantization, sign is the sign function, E(|W|) is the scaling factor of the convolutional layer, A={a1,a2,…,a i }, is i input feature maps.
3. The method according to claim 1, characterized in that The granularity scaling factor is calculated layer by layer to satisfy: Among them, |x1|…|x9| represents the weight on a convolution kernel, a=1,2,…,M represents the convolution kernel corresponding to the output channel, b=1,2,…,N represents the dimension of the convolution kernel corresponding to the input channel, Represents the absolute value of the weight of the cth position of the bth channel dimension of the ath convolution kernel.
4. The method according to claim 1, characterized in that The granularity scaling factor is calculated by channel-by-channel, satisfying: in, represents the scaling factor of the convolution kernel of the first output channel, represents the scaling factor of the convolution kernel of the Mth output channel, E per-channel (|W|) represents the set of channel-by-channel scaling factors, and there are M scaling factors in a layer.
5. The method according to claim 1, characterized in that The calculation method of the granularity scaling factor includes convolution kernel calculation, which satisfies: in, Represents the scaling factor at the cth position of the scaling factor matrix of the convolution kernel for the ath input channel.
6. The method according to claim 1, characterized in that Select a granularity scaling factor, including: Before the formal training of the binary network, a single sample from different data sets is used to perform a training and forward process under each granularity scaling factor to obtain the forward reasoning calculation time corresponding to the three granularity scaling factors, and the granularity scaling factor is selected according to the calculation time limit.
7. The method according to claim 1, characterized in that The scaling factor is calculated as: Among them, W is the floating point weight, W binary is a binary weight, sign() is a sign function; scaling factor E F (|W|) is the mean of the absolute values of the weights before quantization.
8. The method according to claim 1, characterized in that The scaling factor for the SSDD dataset is calculated channel by channel.
9. The method according to claim 1, characterized in that: For the HRSC2016 dataset scaling factor, a channel-by-channel, convolution kernel-by-convolution kernel mixed calculation mode is adopted.
10. The method according to claim 1, characterized in that The YOLOv5m network includes weights quantized to 8 bits in some layers and 1 bit in the remaining layers; in the fourth step, if the accuracy is not up to standard, one of the 1-bit layers in the backbone network is replaced with a finer scaling factor according to the layer index order, and the scaling factors of other layers remain unchanged.