Object Detection Method Based on Model Lightweighting
By introducing packet convolution, multi-layer small convolution kernel and Fire module into the YOLOv5 network, a lightweight object detection model is built, which solves the problems of high computing costs and large storage on portable devices, and realizes real-time detection and efficient operation on embedded devices.
Patent Information
- Application Number
- CN202211612259.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-12-15
AI Technical Summary
The existing target detection network is difficult to achieve real-time detection on portable devices, mainly due to the high computing cost and large storage volume, which limits its application on embedded devices.
Based on YOLOv5, a lightweight object detection model is built by introducing packet convolution, multi-layer small convolution kernel, Fire module and shuffling operations, including multi-layer small convolution kernel replacing large convolution kernels, limiting the number of channels of intermediate features, and using grouping convolution and shuffling operations to reduce the computational volume and storage requirements.
It realizes real-time operation of object detection on embedded devices, reduces the use of software and hardware resources, and improves the accuracy and computing efficiency of the model.
Smart Images

Figure CN116229199B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of object detection, and particularly to an application of model lightweighting based on a convolutional neural network. Technical Background
[0002] The detection and tracking of various objects have always been a research hotspot in the field of computer vision, and have very wide applications in video surveillance, military technology, intelligent traffic management, etc. With the development of computer vision, object detection technology based on deep learning has made breakthrough progress in computing accuracy and has been widely applied in real-life scenarios. In practical applications, although neural networks can achieve real-time object detection on large computers such as servers, it is difficult to achieve this effect when transplanting the network structure to embedded devices because the processor performance of embedded devices is much lower than that of servers, etc. Moreover, existing object detection networks have the disadvantages of high computational cost and large model storage, which are not conducive to their deployment in devices with strict computational time requirements and low memory resources, and this also severely restricts the development and application of deep neural networks on portable devices.
[0003] In order to improve the efficiency and ability of portable devices to process image and video data, while meeting the limitations of storage space and power consumption, designing a lightweight deep neural network architecture suitable for portable devices is the key to solving this problem.
[0004] Researchers mainly conduct research on lightweighting from three different directions, namely, manually designing lightweight neural network models, automatically designing neural network architectures based on neural network architecture search, and compressing neural network models.
[0005] (1) The main idea of manually designing lightweight neural networks lies in designing more efficient network computing methods, mainly targeting convolutional computing methods. Existing deep convolutional neural networks, in order to achieve better performance, set a large number of feature channels and the number of convolutional kernel sizes, but there are often a large amount of redundancies. Manually designing lightweight neural networks constructs a more effective neural network structure by reasonably reducing the number of convolutional kernels, reducing the number of channels of target features, and combining the design of more efficient convolutional operations, etc., which can significantly reduce the parameters and computational amount of the network while maintaining the performance of the neural network, and realize the training and application of deep neural networks on portable devices.
[0006] (2)Automated neural network architecture design based on neural network architecture search means that by taking the set of all candidate neural network architectures as the search space, using the learned search strategy to construct the optimal neural network architecture from the search space, using the performance evaluation strategy to measure the performance of the network architecture, and during the training phase, as a reward to guide the learning of the search strategy, through repeated iterations, the optimal neural network architecture for solving a specific task is obtained, realizing the automatic search of the deep neural network model. The neural network architecture search method has significant overlap with hyperparameter optimization and meta-learning. The neural network architecture search method mainly consists of three parts: the search space, the search strategy, and the performance evaluation strategy.
[0007] (3)Neural network model compression is achieved by cutting the full-precision floating-point numbers of the network weights, further quantifying the intermediate feature outputs of the network, as well as pruning, weight sharing, low-rank decomposition, knowledge distillation and other methods according to the redundancy degree of each layer in the neural network. By compressing the neural network model, the occupied storage space is reduced, the power consumption limit is met, and it is embedded on the chip of the portable device to achieve real-time operation. Summary of the Invention
[0008] In view of the technical background, the present invention is based on the YOLOv5 deep learning framework, and by introducing group convolution into the YOLOv5 network model, the number of channels of the target features is reduced, and a more effective neural network structure is constructed to solve the problems of large computational amount and large storage amount, so that it can run perfectly on the embedded development board.
[0009] To solve the above technical problems, the present invention adopts the following scheme:
[0010] A target detection method based on model lightweighting, the steps include: first collecting pictures; then sending the pictures into the target detection method model for detection; finally obtaining the target information in the pictures; the feature is that the construction steps of the target detection method model include:
[0011] Step 1: Process the training set, convert the picture dataset format into the xml format suitable for YOLOv5 training, and divide it into two major parts: the training set and the test set;
[0012] Step 2: Build the PyTorch deep learning framework, and the configuration of the deep learning model uses the YOLOv5 algorithm, the steps include:
[0013] 2.1) Use Mosaic data augmentation, adaptive anchor box calculation and adaptive picture scaling at the input end to perform preprocessing operations on the image;
[0014] 2.2) Use the Focus structure, CSP structure, and SPP structure in the Backbone backbone network for network feature extraction, and then enhance the network feature fusion ability through the FPN+PAN structure of Neck;
[0015] 2.3) Use the output end Prediction to output three feature sizes, and obtain the parameter matrix according to the image grid division;
[0016] Step 3: Modify the YOLOv5 network. The method is as follows:
[0017] 3.1) Use multiple small convolutional kernels to replace one large convolutional kernel
[0018] Use a 3×3 convolutional kernel to replace the large convolutional kernels of 5×5 and 7×7 sizes; for a receptive field of size 5×5, it is achieved through two convolutional layers of size 3×3; for one convolutional kernel, it is achieved through three convolutional layers;
[0019] 3.2) Limit the number of channels of intermediate features. The method is as follows:
[0020] Adopt the Fire moudle, which includes a compression squeeze layer and an expansion expand layer; reduce the computational amount required by the entire model by reducing the number of channels in the squeeze layer;
[0021] 3.3) Use grouped convolution to replace standard convolution:
[0022] The feature is that in step 3.3), grouped convolution is used to replace standard convolution:
[0023] The method of grouped convolution is as follows: First, group the input feature map feature map, and then convolve each group of feature maps separately;
[0024] The connection method of grouped convolution is as follows: The number of output feature maps of the first group group1 is 2, there are 2 convolutional kernels, and the channel number of each convolutional kernel is 4, which is the same as the channel number of the input feature map of the first group group1. The convolutional kernels only convolve with the input feature maps of the same group, and do not convolve with the input feature maps of other groups;
[0025] And so on;
[0026] Adopt the shuffle Shuffle operation to rearrange the features from different groups so that the groups contain features from previous groups;
[0027] Step 4: Train the network;
[0028] Step 5: Use the trained model to perform a detect operation to obtain a target detection method model based on model lightweighting.
[0029] In the process of constructing the target detection method model, this method modifies the YOLOv5 network through various measures to lightweight the model, so as to reduce the occupation of software and hardware resources during use and is suitable for embedding on a chip to achieve real-time operation. Description of the Drawings
[0030] Figure 1a Decompose the 5×5 convolution into two layers of 3×3 convolutions;
[0031] Figure 1b Decompose the 3×3 convolution into consecutive 1×3 and 3×1 convolutions;
[0032] Figure 2 Schematic diagram of Fire Moudle;
[0033] Figure 3a Schematic diagram of standard convolution;
[0034] Figure 3b Schematic diagram of grouped convolution;
[0035] Figure 4a For general grouped convolution (causing information not to flow between groups);
[0036] Figure 4b and Figure 4c For channel shuffle operation;
[0037] Figure 5 Schematic diagram of the Shuffle process;
[0038] Figure 6 Schematic diagram of YOLOv5. Detailed Implementation Manner
[0039] The present invention will be further described below with reference to the accompanying drawings:
[0040] Step 1: Process the training set, convert the open-source dataset format into the xml format suitable for YOLOv5 training, add some difficult sample pictures, screen out some pictures with poor quality, and divide them into two major parts: the training set and the test set.
[0041] Step 2: Build the PyTorch deep learning framework. The configuration of the deep learning model uses the YOLOv5 algorithm. First, preprocess the image using Mosaic data augmentation, adaptive anchor box calculation, and adaptive image scaling at the input end. Subsequently, use the Focus structure, CSP structure, and SPP structure in the Backbone backbone network for network feature extraction, and then enhance the network feature fusion ability through the FPN+PAN structure of Neck. Finally, use Prediction at the output end to output three feature sizes and obtain the parameter matrix according to the image grid division.
[0042] Step 3: Modify the YOLOv5 network:
[0043] (1) Replacing a single large convolutional kernel with multiple small convolutional kernels can effectively reduce the parameters of the network
[0044] Use a 3×3 convolutional kernel instead of large convolutional kernels of sizes 5×5 and 7×7. For a receptive field of size 5×5, it can be achieved by two 3×3 convolutional layers. For example Figure 1a .
[0045] In terms of the number of parameters, the number of parameters of a 5×5 convolutional kernel is 25, while the number of parameters of two 3×3 convolutional kernels is 18, a reduction of 28% in the number of parameters. From the perspective of floating-point operations (FLOPs), for a feature of size H×W×C in and an output feature map of size H×W×C out , FLOPs 5×5 = H×W×C in ×C out ×5 2 , while the FLOPs of two convolutional layers (3×3)×2 = 2×H×W×C in ×C out ×3. The amount of computation is also reduced accordingly, and two 3×3 convolutions can combine two non-linear layers, increasing the non-linearity ability more than a single large convolutional kernel.
[0046] (2) Limit the number of channels of intermediate features
[0047] For a standard convolution operation without bias, FLOPs = H×W×C in ×C out ×K 2 , and the amount of computation is affected by the number of input channels C in and the number of convolutional kernels C out . Generally, the number of convolutional kernels represents the number of extracted features. Reducing the number of convolutional kernels will affect the accuracy of the network. Therefore, the number of input channels C in can be selected to reduce the amount of computation.
[0048] Therefore, based on this research, the present invention proposes a Fire module, as Figure 2 shown, which reduces the amount of computation while ensuring accuracy. The Fire module consists of two parts, the squeeze layer and the expand layer, and reduces the amount of computation required for the entire model by reducing the number of channels in the squeeze layer. Specifically, first, through the expand layer, a 1×1 convolutional kernel is used to increase the dimension of the input features and perform convolutional operations on the high-dimensional features; then, through the squeeze layer, a 1×1 convolutional kernel is used to reduce the dimension of the feature map and reduce the number of channels of the intermediate features.
[0049] (3) Using grouped convolution instead of standard convolution
[0050] If the size of the input feature map is C×H×W, there are N convolutional kernels, and the number of output feature maps is the same as the number of convolutional kernels, which is also N. The size of each convolutional kernel is C×K×K, and the total number of parameters of the N convolutional kernels is N×C×K×K. The connection method between the input map and the output map is as Figure 3a shown,
[0051] As Figure 3b shown, grouped convolution is to group the input feature map and then perform convolution on each group separately. Assume that the size of the input feature map is still C×H×W, and the number of output feature maps is N. If it is set to be divided into G groups, then the number of input feature maps in each group is , and the number of output feature maps in each group is , the size of each convolutional kernel is , the total number of convolutional kernels is still N, and the number of convolutional kernels in each group is . The convolutional kernel only convolves with the input map in its same group, and the total number of parameters of the convolutional kernel is . It can be seen that the total number of parameters is reduced to of the original, and its connection method is as follows Figure 3b shown. The number of output maps of group1 is 2, there are 2 convolutional kernels, and the number of channels of each convolutional kernel is 4, which is the same as the number of channels of the input map of group1. The convolutional kernel only convolves with the input map in the same group and does not convolve with the input maps of other groups.
[0052] Using grouped convolution will cause the information between groups to not flow, affecting the expressive ability of the network. To enable grouped convolution to obtain the features generated by other groups, the present invention proposes a shuffle operation to rearrange the features from different groups, so that the new groups contain the features from the previous groups, ensuring the information flow between groups.
[0053] The specific method of Shuffle is as follows:
[0054] 1) Expand the Feature Map into a four-dimensional matrix of g×n×w×h (for simplicity of understanding, here w×h is reduced to one dimension and denoted as s);
[0055] 2) Transpose along the g-axis and n-axis of the matrix of size g×n×s;
[0056] 3) After tiling the g-axis and n-axis, obtain the shuffled Feature Map;
[0057] 4) Perform 1×1 convolution within the group.
[0058] Step 4: Train the network, the method is as follows:
[0059] (1) Use the stochastic gradient descent method for training, set the learning rate to 0.01, add the momentum set to 0.937 to accelerate the convergence speed, and set the weight decay to 5e-4 to avoid overfitting.
[0060] (2) Adjust the learning rate by cosine annealing: To avoid getting stuck in local minima in the gradient descent algorithm, use the pre-annealing algorithm to jump out of the local optimal solution and thus restart the training.
[0061] (3) In each iteration, input a batch of labeled training data into the network and then update the parameters.
[0062] Step 5: Use the trained model for detection to obtain the object detection method model based on model lightweighting.
[0063] This method is compared and verified with the current mainstream lightweight object detection networks. Using the COCO dataset, evaluate indicators such as mAP, Params, FLOPs, etc., as shown in the following table.
[0064] Model Size Params(M) FLOPs(G) mAP(0.5:0.95) mAP(0.5) YOLOv3-Tiny 416 8.86 5.62 16.6 33.1 YOLOv4-Tiny 416 6.06 6.96 21.7 40.2 MobileDet 320 3.85 1.02 24.2 - NanoDet-M 416 0.95 1.2 23.5 - YOLOX-Nano 416 0.91 1.07 25.8 - YOLOX-Tiny 416 5.06 6.45 32.8 - Ours 320 0.99 0.73 27.1 41.4 Ours 416 0.99 1.24 30.6 45.5
[0065] As can be seen from the experimental results, the model proposed in this paper achieves a better trade-off between accuracy and parameter calculation complexity. Compared with YOLOX-Nano, our model achieves a 30.6% mAP with only 0.99M parameters, an increase of 4.8%. Compared with NanoDet-M, our model improves the mAP by 7.1% with comparable computational complexity. In summary, the model in this paper achieves a better balance between accuracy and lightweight.
Claims
1. A target detection method based on model lightweighting, the steps include: First, collect images; Then, send the images into the target detection method model for detection; Finally, obtain the target information in the images. Its feature is that the construction steps of the target detection method model include: Step 1: Process the training set, convert the format of the image dataset into the xml format suitable for YOLOv5 training, and divide it into two major parts: the training set and the test set; Step 2: Build the PyTorch deep learning framework. The configuration of the deep learning model uses the YOLOv5 algorithm. The steps include: 2.1) Use Mosaic data augmentation, adaptive anchor box calculation, and adaptive image scaling at the input end to preprocess the images; 2.2) Use the Focus structure, CSP structure, and SPP structure in the Backbone backbone network to extract network features, and then use the FPN+PAN structure of Neck to enhance the ability of network feature fusion; 2.3) Use Prediction at the output end to output three feature sizes, and obtain the parameter matrix according to the image grid division; Step 3: Modify the YOLOv5 network. The method is: 3.1) Use multiple small convolutional kernels to replace one large convolutional kernel Use 3×3 convolutional kernels to replace the large convolutional kernels of 5×5 and 7×7 sizes; for a receptive field of size 5×5, it is implemented by two convolutional layers of size 3×3; 3.2) Limit the number of channels of the intermediate features. The method is: Adopt the Fire moudle. The Fire moudle includes a compression squeeze layer and an expansion expand layer; reduce the amount of computation required for the entire model by reducing the number of channels in the squeeze layer; 3.3) Use grouped convolution to replace standard convolution: The feature is that in step 3.3), grouped convolution is used to replace standard convolution: The method of grouped convolution is: First, group the input feature map feature map, and then convolve each group of feature maps separately; The connection method of grouped convolution is: The number of output feature maps of the first group group1 is 2, there are 2 convolutional kernels, and the channel number of each convolutional kernel is 4, which is the same as the channel number of the input feature map of the first group group1. The convolutional kernels only convolve with the input feature maps of the same group, and do not convolve with the input feature maps of other groups; And so on; Adopt the shuffle Shuffle operation to rearrange the features from different groups, so that the groups contain features from the previous groups; Step 4: Train the network; Step 5: Use the trained model to perform the detect operation, so as to obtain the target detection method model based on model lightweighting.
2. The object detection method based on model lightweight as claimed in claim 1, wherein In step 1, add difficult sample images to the image dataset and screen out some images with poor quality.
3. The object detection method based on model lightweighting according to claim 1, characterized in that In step 3.2), first, through the expand layer, use a 1×1 convolutional kernel to perform dimensionality increase on the input features and perform convolutional operations on the high-dimensional features; then, through the squeeze layer, use a 1×1 convolutional kernel to perform dimensionality reduction on the feature map and reduce the number of channels of the intermediate features.
4. The object detection method based on model lightweighting according to claim 1, characterized in that In the grouped convolution of step 3.3), assume the size of the input feature map is C×H×W, the number of output feature maps is N, and assume it is divided into G groups. Then the number of input feature maps in each group is The number of output feature maps in each group is The size of each convolutional kernel is The total number of convolutional kernels is still N, and the number of convolutional kernels in each group is The convolutional kernels only perform convolution with the input feature maps in the same group. The total number of parameters of the convolutional kernels is 5. The object detection method based on model lightweight according to claim 1, characterized in that The steps of the shuffle Shuffle in step 3.3) include: First, expand the feature map into a four-dimensional matrix of g×n×w×h. Here, reduce w×h to one dimension and denote it as s; Next, transpose along the g-axis and n-axis of the matrix of size g×n×s; Then, tile the g-axis and n-axis to obtain the shuffled feature map; Finally, perform a 1×1 convolution within the group.
6. The object detection method based on model lightweighting according to claim 1, characterized in that The method of step 4 is as follows: 4.1) Use the stochastic gradient descent method for training. Set the learning rate to 0.01, add momentum set to 0.937 to accelerate the convergence speed, and set the weight decay to 5e-4 to avoid overfitting; 4.2) Adjust the learning rate by cosine annealing: To avoid getting stuck in local minima in the gradient descent algorithm, use the pre-annealing algorithm to jump out of the local optimum and thus restart the training; 4.3) In each iteration, input a batch of labeled training data into the network and then update the parameters.
Citation Information
Patent Citations
Human body posture estimation method based on bidirectional serialization modeling
CN112633220A
YOLO v5-based attached marine organism type identification method
CN113688948A