A cross-layer fusion improved YOLOv4 road object recognition algorithm

By introducing the Dense-SPP module and the Dense-feature fusion module into the YOLOv4 network and performing parameter pruning, a lightweight Dense-YOLOv4-Small network model is designed, which solves the problems of slow detection speed and large calculations in the existing technology, and achieves efficient road target recognition.

CN114565896BActive Publication Date: 2025-05-23XIDIAN UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210006574.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-05
Publication Date
2025-05-23
Estimated Expiration
2042-01-05

AI Technical Summary

Technical Problem

The prior art has problems such as slow detection speed and large calculation amount in road target recognition, and it is difficult to improve detection speed while maintaining detection accuracy.

Method used

Based on the YOLOv4 network, the Dense-SPP module and the Dense-feature fusion module are designed to reduce the computing amount of the network, and the lightweight Dense-YOLOv4-Small network model is designed through parameter pruning and cutting.

Benefits of technology

While maintaining the detection accuracy almost does not decrease, the detection speed is greatly improved and the model parameter quantity and weight file size are significantly reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114565896B_ABST
    Figure CN114565896B_ABST
Patent Text Reader

Abstract

The present invention proposes a cross-layer fusion improved YOLOv4 road target recognition algorithm. The data set used is the KITTI road target data set. In order to make the model more lightweight while maintaining the detection accuracy, the method of the present invention takes YOLOv4 as the basic network, draws on the idea of ​​DenseNet, designs a Dense-SPP cross-layer spatial pooling module and a Dense-feature fusion module, and performs parameter pruning and reduction on the original model, designs a lightweight Dense-YOLOv4-Small network model, reduces the CSP module in the skeleton network CSPDarknet-53, and The number of ResUnits in the original CSP module was uniformly set to 1. The network was pruned to eliminate redundant calculations of the network. The “Misc” and “Dontcare” in the KITTI road target dataset were removed to obtain the KITTI‑7classes road target dataset. The YOLOv4, Dense‑YOLOv4, and Dense‑YOLOv4‑Small models were trained on this dataset, and the detection speed and detection performance of the three models were compared. The detection results show that the detection speed of Dense‑YOLOv4‑Small is greatly improved, and the detection accuracy remains almost unchanged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of image recognition, and is a cross-layer fusion improved YOLOv4 road target recognition algorithm, which shows good detection performance on general standard datasets. Background Art

[0002] In order to manage traffic roads more efficiently and maintain social stability, it is necessary to detect pedestrians and vehicles on the road. The detection task of vehicles and pedestrians occupies an important position in the field of unmanned driving. Intelligent vehicle and pedestrian recognition can assist traffic police in effective management and traffic flow control, and can timely predict the next traffic conditions and prevent traffic congestion.

[0003] With the continuous development of computer technology and the continuous improvement of computing power, computer vision and target detection have become popular directions in recent years. Target detection can be used to identify and locate specific objects, and has broad development prospects in driving assistance systems, military early warning systems, etc. Target detection technology includes traditional target detection technology and target detection technology based on deep learning. The latter has become the mainstream algorithm in the current target detection field because it is superior to the former in terms of performance and complexity. In order to manage traffic roads more efficiently and maintain social stability, it is necessary to detect targets such as pedestrians and vehicles on the road. The detection task of vehicles and pedestrians occupies an important position in the field of unmanned driving. Intelligent vehicle and pedestrian recognition can assist traffic police in effective management and traffic flow control, and can timely predict the next traffic conditions and prevent traffic congestion.

[0004] Based on the YOLOv4 network, the present invention uses the KITTI road target data set to build a higher performance road target detection and recognition algorithm. Based on the YOLOv4 network, the present invention draws on the idea of ​​DenseNet to design the Dense-SPP module and the Dense-feature fusion module, which can effectively perform multi-scale pooling of high-level features to increase the receptive field and more fully integrate the high-level features of the network, while also reducing the amount of network calculation.

[0005] The data set used in the present invention is the KITTI road target data set. In order to make the model more lightweight while basically maintaining the detection accuracy, a Dense-YOLOv4-Small network model is designed. The three models of YOLOv4, Dense-YOLOv4, and Dense-YOLOv4-Small are trained on the KITTI-7classes road target data set, and the performance of the three models in detection speed, mAP and Recall indicators are compared. The present invention designs a lightweight Dense-YOLOv4-Small network model by designing a Dense-SPP cross-layer spatial pooling module and a Dense-feature fusion module, and pruning and reducing the parameters of the original model. Pruning the network can eliminate redundant calculations of the network, without causing a significant decrease in accuracy, but the detection speed can be greatly improved. Summary of the invention

[0006] In view of the above problems, the purpose of the present invention is to provide an improved YOLOv4 road target recognition algorithm based on a cross-layer fusion module for the YOLOv4 network structure.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] A cross-layer fusion improved YOLOv4 road target recognition algorithm, based on YOLOv4 as the basic network, draws on the idea of ​​DenseNet, designs the Dense-SPP cross-layer spatial pooling module and the Dense-feature fusion module, and prunes and reduces the parameters of the original model, designs a lightweight Dense-YOLOv4-Small network model, eliminates redundant calculations of the network, and greatly improves the detection speed with almost no decrease in test accuracy;

[0009] The road target recognition algorithm comprises the following steps:

[0010] Step 1: Download the KITTI road target dataset, a general dataset in the current target detection field, remove the "Misc" and "Dontcare" data in the original KITTI dataset, and create the KITTI-7Classes road target dataset. Using this dataset can ensure that the algorithm detection effect is consistent with the general dataset disclosed in this field, and construct the road target dataset used in the present invention; divide the test set, validation set and training set in a ratio of 6:2:2;

[0011] The KITTI dataset is currently the largest dataset for autonomous driving scenarios. KITTI contains real image data collected from various road scenes. The KITTI dataset contains nine categories, namely Car, Van, Truck, Pedestrian, Person (sitting), Cyclist, Tram, Misc and Dontcare. Since there are two categories in KITTI, namely "Misc" and "Dontcare", which are "disorganized" and "don't care" respectively, these two categories are meaningless, and since these two categories have no specific target features, the objects that may be contained in the "Misc" class in different pictures are different. The present invention removes "Misc" and "Dontcare" from the original KITTI dataset to form the KITTI-7Classes dataset. The present invention will be trained and tested on KITTI-7Classes.

[0012] Step 2. Use the standard YOLOv4 network to train and identify and locate road targets; Use the standard YOLOv4 network to train the road target dataset based on step 1, download the standard YOLOv4 network and compile it. The download address of the standard YOLOv4 network is: https: / / github.com / AlexeyAB / darknet; For the road target data kitti-7classes, change the training set, validation set, and test set directories in the kitti7.data file in the cfg folder to the address of the downloaded dataset, specify the number of categories and category names, and delete the "Misc" and "Dontcare" entries in kitti7.name; In the command executed by training, set the number of iterations (epoch) to 100 according to the accuracy requirements, load kitti7.data according to this experimental dataset, and load yolov4.cfg at the same time, and the program can start training; Save the weight file Q of each layer during the training process 1 , as the weight input file for detection after training; using the weight file Q 1 The test was conducted to obtain the mean average precision (mAP), recall rate (Recall) and frame rate (Frame Per Second, FPS) during detection. When the target occupies more than half of the entire image, the network's detection effect is not good due to the limitation of the actual effective receptive field.

[0013] The training process is as follows:

[0014] 1) Build a YOLOv4 network model and use the Initialization function to initialize the weight parameters of each layer of the neural network;

[0015] YOLOv4 consists of four parts: (1) Input: refers to the original sample data input into the network; (2) Backbone network: refers to the convolutional neural network structure that performs feature extraction operations; (3) Neck: fuses the image features extracted by the backbone network and passes the fused features to the prediction layer; (4) Head: predicts the target object of interest in the image and generates a visual prediction box and target category;

[0016] After downloading the standard YOLOv4 network, use the make command to compile the YOLOv4 network to form an executable file darknet; edit the kitti7.data file in the cfg folder for the road target dataset KITTI-7classes, and change the class, train, valid, and names strings to the directory and parameters of the corresponding dataset. In this way, the parameters required for the Input part of the standard YOLOv4 network are edited. After setting epoch in the command line for training execution, load kitti7.data according to the experimental dataset and load yolov4.cfg at the same time, and the program can start training; the program will use the Initialization function to initialize the weight parameters of each layer of the neural network when it is running;

[0017] 2) Input image data from the Input part, pass through the Backbone part, and finally output feature maps of two scales, and use the classifier to output the prediction box Pb 1 and classification probability CP x ;

[0018] The image data is input from the Input part, passed through the Backbone part, and finally outputted feature maps of two scales. The feature maps of two different scales are sent to the Neck part composed of the Feature Pyramid Network (FPN), and the fused features are passed to the prediction layer. At the same time, the Head part completes the classification of the target and outputs the prediction box Pb 1 and classification probability CP x , where x is the index of each category;

[0019] 3) Perform IoU and NMS post-processing on these data and convert the predicted box Pb 2 Compare with the real frame Gtb, and use the Adam algorithm to update the weights of each layer of the neural network;

[0020] The prediction box Pb generated by the Backbone network 1The number is too large, and there are a large number of detection boxes for the same object in the picture, resulting in redundant detection results; the Head part of YOLOv4 will simultaneously complete the prediction box and its corresponding classification probability; IoU and NMS post-processing are performed on these data to obtain processed data; the Iou and NMS used here are the CIoU_loss and NMS of the standard YOLOv4; after these post-processing, the prediction box Pb of the target of interest can be obtained 2 The corresponding classification probability CP x ; At the same time, the Adam algorithm is used to update the weights of each layer of the neural network using the loss obtained in the post-processing process;

[0021] 4) Loop through steps 2) and 3) and continue iterating until the epoch value specified in the command is reached, stop training, and output a file Q recording the weights and offsets of each layer. 1 ; Using Q 1 The obtained weights and offsets are used to detect the test set, and the mAP, Recall and frame rate FPS during detection are calculated;

[0022] The present invention sets the iteration threshold epoch=100 according to the accuracy requirement. When the number of iterations is less than the threshold, the Adam algorithm is used to update the weights of each layer of the network until the threshold epoch=100 stops training, calculates mAP and Recall, and outputs a file Q recording the weights and offsets of each layer. 1 ;

[0023] Step 3: Design a Dense-YOLOv4 network model; design two cross-layer fusion modules in the model, namely the Dense-SPP module and the Dense-feature fusion module; take YOLOv4 as the basic network and introduce the above two modules into the YOLOv4 model;

[0024] On the Backbone network, CSPDarknet-53 is used as the feature extraction network. On the feature fusion network, the original SPP module is replaced with a Dense-SPP module. At the same time, the single-channel five-layer convolution module on the path aggregation network (PAN) structure that fuses the output feature map of the previous scale is replaced with a Dense-feature fusion module designed by the present invention. As for the detector of the network, the original three-scale detection is adopted. The input size of the network is 640×640×3, and the sizes of the feature maps of the final detection layer are 20×20, 40×40 and 80×80, respectively, to detect large, medium and small targets.

[0025] (1) A Dense-SPP module is designed. A cross-layer connection module is introduced on the basis of the original SPP module. In this way, the feature map is divided into two branches, one of which is subjected to the convolution and pooling operations of the original SPP module, and the other is subjected to a 1×1×512 single convolution. Then, the output feature maps of the two branches are concatenated. The number of convolution kernels of the Dense-SPP module is 11264, while the number of convolution kernels of the original SPP module is 20480, which reduces the number of parameters by 45%. The Dense-SPP module designed by the present invention also adopts 5 The difference of the CBL module is that it adds a cross-layer connection, which integrates the previous convolution information and retains more original information. Secondly, the network is hierarchical, that is, for the same task, different samples may use different types of features to complete the detection. The shallow network extracts simple features, such as texture features, while different samples may require features of different complexity to make judgments. Without the Concat splicing operation, the network does not save the features extracted by the previous shallow network. After adding the Concat splicing operation, it is equivalent to splicing the feature information of the first layer of the module on the output, which becomes effective.

[0026] (2) Design a Dense-feature fusion module. Entering the Dense-feature fusion module, the features are divided into two branches. One feature map undergoes four convolutions with kernel sizes of 1×1×256, 3×3×256, 1×1×256, and 3×3×256, respectively. The other feature map undergoes a single convolution of 1×1×256, and then the features of the two branches are concat-joined. Compared with the single-channel quintuple convolution in the original YOLOv4, the Dense-feature fusion module has less computational complexity. The number of convolution kernel parameters is 5376, while the number of convolution kernel parameters of the single-channel quintuple convolution is 9984, which reduces the computational complexity by 40%. Since each convolution wastes some information, such as the inhibitory effect of the activation function and the randomness of the convolution kernel parameters, This cross-layer connection is equivalent to directly taking the previously processed information and processing it together now, which has the effect of reducing loss; and the Concat splicing operation is equivalent to splicing the feature information of the first layer of the module on the output, realizing feature reuse; the output of each layer of the dense cross-layer connection module will establish an input-output relationship with all subsequent layers, and the input of each layer is the accumulation of all previous layers. This mode can retain the simple features of the shallow layer of the network to the deep layer of the network and fuse them with high semantic features to achieve feature reuse; this model can reduce the number of network parameters; subsequent network layers obtain the gradient of the loss function and the original input signal, so that the network contains implicit deep supervision, which makes it easy to train deeper networks and has a regularization effect, alleviating the gradient disappearance problem in the training process to a certain extent;

[0027] Step 4: Compare the test results of the model performance in step 2 and step 3, including model detection accuracy, model detection speed, model detection recall rate, and model weight file size, and check the images in the data sets actually detected in step 2 and step 3 to analyze the test results.

[0028] The present invention takes YOLOv4 as the basic network, draws on the idea of ​​DenseNet, designs the Dense-SPP cross-layer spatial pooling module and the Dense-feature fusion module, and performs parameter pruning and reduction on the original model to design a lightweight Dense-YOLOv4-Small network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0030] Figure 1 is a flow chart of the method of the present invention;

[0031] Figure 2 This is a flowchart for training using YOLOv4;

[0032] Figure 3 It is the structural diagram of Dense-SPP and Dense-feature fusion module; 3(a) Dense-SPP module, 3(b) Dense-feature fusion module;

[0033] Figure 4 It is the structure diagram of the cross-layer fusion module after lightweighting; 4(a) Dense-SPP-Small module, 4(b) Dense-Feature Fusion-Small module;

[0034] Figure 5 It is a bar chart comparing the performance of YOLOv4 and the model with the cross-layer fusion module; 5(a) mAP comparison bar chart, 5(b) recall rate comparison bar chart;

[0035] Figure 6 The following is a comparison chart of the detection speed of YOLOv4 and the model with the cross-layer fusion module; 6(a) total parameter comparison bar chart, 6(b) model weight file size comparison bar chart, 6(c) FPS comparison bar chart, 6(d) single image detection comparison bar chart;

[0036] Figure 7It is a comparison chart of the detection effects of the Dense-YOLOv4, Dense-YOLOv4-Small and the original YOLOv4 algorithm; 7(a) original YOLOv4 detection effect chart, 7(b) Dense-YOLOv4 detection effect chart, 7(c) Dense-YOLOv4-Small detection effect chart;

[0037] Figure 8 It is the performance analysis of the model with YOLOv4 and the introduction of cross-layer fusion module;

[0038] Fig. 9 It is the model detection speed analysis of YOLOv4 and the introduction of cross-layer fusion module; DETAILED DESCRIPTION

[0039] In order to make the above and other purposes, features and advantages of the present invention more obvious, the embodiments of the present invention are specifically cited below, and the accompanying drawings are used to provide a detailed description as follows:

[0040] Figure 1 The specific flow chart of this method can be divided into four steps:

[0041] Step 1: Download the KITTI road target dataset, a general dataset in the current target detection field, remove the "Misc" and "Dontcare" data in the original KITTI dataset, and create the KITTI-7Classes road target dataset. Using this dataset can ensure that the algorithm detection effect is consistent with the general dataset disclosed in this field, and construct the road target dataset used in the present invention; divide the test set, validation set and training set in a ratio of 6:2:2;

[0042] The KITTI dataset is currently the largest dataset for autonomous driving scenarios. KITTI contains real image data collected from various road scenes. The KITTI dataset contains nine categories, namely Car, Van, Truck, Pedestrian, Person (sitting), Cyclist, Tram, Misc and Dontcare. Since there are two categories in KITTI, namely "Misc" and "Dontcare", which are "disorganized" and "don't care" respectively, these two categories are meaningless, and since these two categories have no specific target features, the objects that may be contained in the "Misc" class in different pictures are different. The present invention removes "Misc" and "Dontcare" from the original KITTI dataset to form the KITTI-7Classes dataset. The present invention will be trained and tested on KITTI-7Classes.

[0043] Step 2. Use the standard YOLOv4 network to train and identify and locate road targets; Use the standard YOLOv4 network to train the road target dataset based on step 1, download the standard YOLOv4 network and compile it. The download address of the standard YOLOv4 network is: https: / / github.com / AlexeyAB / darknet; For the road target data kitti-7classes, change the training set, validation set, and test set directories in the kitti7.data file in the cfg folder to the address of the downloaded dataset, specify the number of categories and category names, and delete the "Misc" and "Dontcare" entries in kitti7.name; In the command executed by training, set the number of iterations (epoch) to 100 according to the accuracy requirements, load kitti7.data according to this experimental dataset, and load yolov4.cfg at the same time, and the program can start training; Save the weight file Q of each layer during the training process 1 , as the weight input file for detection after training; using the weight file Q 1 The test was conducted to obtain the mean average precision (mAP), recall rate (Recall) and frame rate (Frame Per Second, FPS) during detection. When the target occupies more than half of the entire image, the network's detection effect is not good due to the limitation of the actual effective receptive field.

[0044] Reference Figure 2 , the training process is as follows:

[0045] 1) Build a YOLOv4 network model and use the Initialization function to initialize the weight parameters of each layer of the neural network;

[0046] YOLOv4 consists of four parts: (1) Input: refers to the original sample data input into the network; (2) BackBone network: refers to the convolutional neural network structure that performs feature extraction operations; (3) Neck: fuses the image features extracted by the backbone network and passes the fused features to the prediction layer; (4) Head: predicts the target object of interest in the image and generates a visual prediction box and target category;

[0047] After downloading the standard YOLOv4 network, use the make command to compile the YOLOv4 network to form an executable file darknet; edit the kitti7.data file in the cfg folder for the road target dataset KITTI-7classes, and change the class, train, valid, and names strings to the directory and parameters of the corresponding dataset. In this way, the parameters required for the Input part of the standard YOLOv4 network are edited. After setting epoch in the command line for training execution, load kitti7.data according to the experimental dataset and load yolov4.cfg at the same time, and the program can start training; the program will use the Initialization function to initialize the weight parameters of each layer of the neural network when it is running;

[0048] 2) Input image data from the Input part, pass through the Backbone part, and finally output feature maps of two scales, and use the classifier to output the prediction box Pb 1 and classification probability CP x ;

[0049] The image data is input from the Input part, passed through the Backbone part, and finally outputted feature maps of two scales. The feature maps of two different scales are sent to the Neck part composed of the Feature Pyramid Network (FPN), and the fused features are passed to the prediction layer. At the same time, the Head part completes the classification of the target and outputs the prediction box Pb 1 and classification probability CP x , where x is the index of each category;

[0050] 3) Perform IoU and NMS post-processing on these data and convert the predicted box Pb 2 Compare with the real frame Gtb, and use the Adam algorithm to update the weights of each layer of the neural network;

[0051] The prediction box Pb generated by the Backbone network 1 The number is too large, and there are a large number of detection frames for the same object in the picture, resulting in redundant detection results; the Head part of YOLOv4 will simultaneously complete the prediction frame and its corresponding classification probability; these data are post-processed with Iou and NMS to obtain processed data; the Iou and NMS used here are the CIoU_loss and NMS of the standard YOLOv4; after these post-processing, the prediction frame Pb of the target of interest can be obtained 2 The corresponding classification probability CP x ; At the same time, the Adam algorithm is used to update the weights of each layer of the neural network using the loss obtained in the post-processing process;

[0052] 4) Loop through steps 2) and 3) and continue iterating until the epoch value specified in the command is reached, stop training, and output a file Q recording the weights and offsets of each layer. 1 ; Using Q 1 The obtained weights and offsets are used to detect the test set, and the mAP, Recall and frame rate FPS during detection are calculated;

[0053] The present invention sets the iteration threshold epoch=100 according to the accuracy requirement. When the number of iterations is less than the threshold, the Adam algorithm is used to update the weights of each layer of the network until the threshold epoch=100 stops training, calculates mAP and Recall, and outputs a file Q recording the weights and offsets of each layer. 1 ;

[0054] The most basic network performance evaluation indicators are divided into four categories, namely TP (True Positives): positive samples are correctly identified as positive samples, that is, dogs are correctly identified as dogs; TN (True Negatives): negative samples are correctly identified as negative samples, that is, cats are correctly identified as cats; FP (False Positives): negative samples are incorrectly identified as positive samples, that is, cats are incorrectly identified as dogs; FN (False Negatives): positive samples are incorrectly identified as negative samples, that is, dogs are incorrectly identified as cats; Accuracy represents the ratio of the number of correctly predicted samples to the total number of samples, which is used to evaluate the overall accuracy of the algorithm model. The calculation method is Precision is the ratio of the number of correctly identified samples to the total number of identified samples. The calculation method is: The recall rate is the ratio of samples correctly identified as positive examples to all positive examples. The calculation method is: An algorithm model with good performance should maintain a high recall rate while ensuring a high accuracy rate. The Precision-Recall (PR) curve is used to show the trade-off between the accuracy and recall rate of the algorithm model. AP refers to the area enclosed by the PR curve plotted by the accuracy and recall rate obtained at a certain threshold and the horizontal and vertical axes, which measures the detection performance of the model in each category. mAP refers to the average AP of multiple target categories, which is used to measure the detection performance of the algorithm model on all tested categories. If there are N categories, the calculation method of mAP is The present invention mainly uses the model overall evaluation index mAP and Recall as the main evaluation indicators;

[0055] Step 3: Design a Dense-YOLOv4 network model. In this model, two cross-layer fusion modules are designed, namely the Dense-SPP module and the Dense-feature fusion module. The structure is as follows: Figure 3 (a) and 3(b); YOLOv4 is used as the basic network, and the above two modules are introduced into the YOLOv4 model; CSPDarknet-53 is used as the feature extraction network on the skeleton network, and the original SPP module is replaced with a Dense-SPP module on the feature fusion network; at the same time, the single-channel five-layer convolution module on the path aggregation network (PAN) structure that fuses the output feature map of the previous scale is replaced with a Dense-feature fusion module designed by the present invention; in terms of the network detector, the original three-scale detection is adopted, the input size of the network is 640×640×3, and the final detection layer feature map sizes are 20×20, 40×40 and 80×80, respectively, to detect large, medium and small targets; the CSP module in the skeleton network CSPDarknet-53 is reduced, and the number of ResUnits in the original CSP module is uniformly set to 1, called the CSP1 module, and the Dense-feature fusion and Dense-SPP modules are lightweighted. After lightweighting, the structure is as follows Figure 4 (a) and 4(b); the adjusted network is called Dense-YOLOv4-Small model, and its parameter count is greatly reduced

[0056] (1) A Dense-SPP module is designed. A cross-layer connection module is introduced on the basis of the original SPP module. In this way, the feature map is divided into two branches, one of which is subjected to the convolution and pooling operations of the original SPP module, and the other is subjected to a 1×1×512 single convolution. Then, the output feature maps of the two branches are concatenated. The number of convolution kernels of the Dense-SPP module is 11264, while the number of convolution kernels of the original SPP module is 20480, which reduces the number of parameters by 45%. The Dense-SPP module designed by the present invention also adopts 5 CBL modules, but the difference is that a cross-layer connection is added, which integrates the previous convolution information and retains more More original information; secondly, the network is hierarchical, that is, for the same task, different samples may use different types of features to complete the detection; for example: for a bottle of pure water and a bottle of beverage, if a sample is yellow, it is judged to be a beverage by color features, and if the sample is colorless, it needs to be judged by more complex features such as its taste; the shallow network extracts simple features, such as texture features, and different samples may require features of different complexity to be judged. Without the Concat splicing operation, the network does not save the features extracted by the previous shallow network; after adding the Concat splicing operation, it is equivalent to splicing the feature information of the first layer of the module on the output, which becomes effective;

[0057] (2) Design a Dense-feature fusion module, such as Figure 3As shown in (b), the Dense-feature fusion module is entered, and the features are divided into two branches. One feature map undergoes four convolutions with kernel sizes of 1×1×256, 3×3×256, 1×1×256, and 3×3×256 respectively, and the other feature map undergoes a single convolution of 1×1×256, and then the features of the two branches are concat-joined. Compared with the single-channel quintuple convolution in the original YOLOv4, the Dense-feature fusion module has less computational complexity. The number of convolution kernel parameters is 5376, while the number of convolution kernel parameters of the single-channel quintuple convolution is 9984, which reduces the computational complexity by 40%. Since each convolution wastes some information, such as the inhibitory effect of the activation function and the randomness of the convolution kernel parameters, this cross-layer connection is equivalent to directly processing the previously processed information together with the current one. It has the effect of reducing loss; and the Concat splicing operation is equivalent to splicing the feature information of the first layer of the module on the output, realizing feature reuse; the output of each layer of the dense cross-layer connection module will establish an input-output relationship with all the subsequent layers, and the input of each layer is the accumulation of all the previous layers. This mode can retain the simple features of the shallow layer of the network to the deep layer of the network and fuse them with high semantic features to realize feature reuse; this model can reduce the number of network parameters; the CSP module in the skeleton network CSPDarknet-53 is reduced, and the number of ResUnits in the original CSP module is uniformly set to 1, called the CSP1 module, and the two designed cross-layer fusion modules are lightweight adjusted. The adjusted network is called the Dense-YOLOv4-Small model, and its parameter volume is greatly reduced;

[0058] Step 4: Compare the test results of the model performance in step 2 and step 3, including model detection accuracy, model detection speed, model detection recall rate, and model weight file size, and check the images in the data set actually detected in step 2 and step 3 to analyze the test results;

[0059] The present invention takes YOLOv4 as the basic network, draws on the idea of ​​DenseNet, designs the Dense-SPP cross-layer spatial pooling module and the Dense-feature fusion module, and performs parameter pruning and reduction on the original model to design a lightweight Dense-YOLOv4-Small network model.

[0060] The invention is further described below in conjunction with a simulation example.

[0061] Simulation example:

[0062] The present invention uses the original YOLOv4 as a comparison sample, and both the training data set and the test data set are from the general road target data set KITTI-7classes to verify the universality of the algorithm to different data sets.

[0063] The AP and mAP on the dataset are as follows Figure 8 As shown, from Figure 8 As can be seen from the figure, the Dense-YOLOv4 model has the best mAP value, reaching 89.9%, which is 0.8% higher than the original YOLOv4 algorithm. Figure 4 It can be seen that the Recall of the Dense-YOLOv4 model is also the highest, reaching 90.6%, which is 1.1% higher than the original YOLOv4 model. For each category, the Dense-YOLOv4 model is basically higher than the AP value of the original YOLOv4. It can be seen that the cross-layer fusion Dense module proposed by the method of the present invention can reuse the features extracted by the feature extraction network, and this cross-layer fusion module is equivalent to splicing the input feature information in the output feature map, making the maximum use of the extracted features, and improving the detection effect and accuracy. The mAP value of the model Dense-YOLOv4-Small after a large amount of pruning reaches 87.7%, and its accuracy has not dropped significantly. Its Recall value reaches 90.1%, which is 0.6% higher than the original YOLOv4 model. This shows that pruning the network can eliminate the redundant calculation of the network, which will not cause a significant drop in accuracy, but can greatly improve the detection speed.

[0064] Comparison of model parameter quantity, FPS and detection time performance of three models: YOLOv4, Dense-YOLOv4 and Dense-YOLOv4-Small Figure 5 , Figure 6 and Fig. 9 As shown. The number of parameters of the Dense-YOLOv4 model is 52221792, while that of the original YOLOv4 is 64165728, a decrease of about 20%. Therefore, the cross-layer fusion module designed by the present invention can effectively reduce the number of parameters, and its FPS is also increased from 43 to 49. Dense-YOLOv4-Small has the fastest detection speed, with an FPS of 93, which is twice that of the original YOLOv4. The memory occupied by the model is only 31M, and the number of parameters and the memory size occupied by the model are about one tenth of the original YOLOv4. It can be seen that the lightweight network structure of Dense-YOLOv4-Small designed by the present invention reduces the number of residual structures of the CSP module in the original YOLOv4. The original function of the residual structure is to prevent the network from being too deep and causing a decrease in network performance. However, for the KITTI-7classes data set used in the present invention, reducing the number of residual structures does not cause a significant decrease in the overall performance of the network, but the detection speed of the model is greatly improved.

[0065] The actual detection effects of the two algorithm models Dense-YOLOv4 and Dense-YOLOv4-Small proposed in this paper are compared with the original YOLOv4 model. Figure 7 As shown in the figure, the same picture is tested on the original YOLOv4, Dense-YOLOv4 and Dense-YOLOv4-Small models. Taking the leftmost vehicle in the figure as an example, the detection confidence of the original YOLOv4 is 93%, and the detection confidence of Dense-YOLOv4 is 97%, which is an increase of 4%. This shows that the Dense cross-layer fusion module designed in this section can fuse the information of the previous convolution and retain more original information. It retains some original simple features for the subsequent prediction of the high-level network, which can make the network more hierarchical and improve the detection accuracy and effect. After pruning the Dense-YOLOv4 network, the detection confidence of the leftmost vehicle of the Dense-YOLOv4-Small network is 94%, which is 1% higher than that of the original YOLOv4. This shows that after pruning the network, redundant calculations are cut off and effective calculations are retained.

[0066] In summary, the simulation results show that compared with the original YOLOv4 algorithm model, the cross-layer fusion Dense module proposed in the present invention can reuse the features extracted by the feature extraction network, and this cross-layer fusion module is equivalent to splicing the input feature information in the output feature map, making the maximum use of the extracted features, and can improve the detection effect and accuracy of the Dense-YOLOv4 and Dense--YOLOv4-Small algorithm models. Pruning the network can eliminate redundant calculations of the network, which will not cause a significant decrease in accuracy, but can greatly improve the detection speed.

Claims

1. A cross-layer fusion improved YOLOv4 road object recognition method, based on the road object recognition of the KITTI general dataset, Features: Step 1: Download the KITTI road target dataset, a general dataset in the current target detection field, remove the "Misc" and "Dontcare" data in the original KITTI dataset, and create the KITTI-7Classes road target dataset. Using this dataset can ensure that the algorithm detection effect is consistent with the public general dataset in this field, and construct the road target dataset used; divide the test set, validation set, and training set in a ratio of 6:2:2; The KITTI dataset is currently the largest dataset for autonomous driving scenarios. KITTI contains real image data collected from various road scenes. The KITTI dataset contains nine categories, namely Car, Van, Truck, Pedestrian, Personsitting, Cyclist, Tram, Misc, and Dontcare. Since there are two categories in KITTI, "Misc" and "Dontcare", which are "disorganized" and "don't care" respectively, these two categories are meaningless, and since these two categories have no specific target features, the objects contained in the "Misc" class are different in different pictures. The "Misc" and "Dontcare" in the original KITTI dataset are removed to form the KITTI-7Classes dataset, which will be trained and tested on KITTI-7Classes. Step 2: Use the standard YOLOv4 network to train, identify and locate road targets; Use the standard YOLOv4 network to train the road target data set based on step 1, download the standard YOLOv4 network and compile it; For the road target data kitti-7classes, change the training set, validation set, and test set directories in the kitti7.data file in the cfg folder to the address of the downloaded data set, specify the number of categories and category names, and delete the "Misc" and "Dontcare" entries in kitti7.name; In the command executed by training, set the number of iterations to 100 according to the accuracy requirements, load kitti7.data according to this experimental data set, and load yolov4.cfg at the same time, and the program can start training; Save the weight file Q of each layer during the training process 1 , as the weight input file for detection after training; Using the weight file Q 1 The test was conducted to obtain the mean average precision (mAP), recall rate (Recall) and frame rate (Frame Per Second) during detection. When the target occupies more than half of the entire image, the detection effect of the network is not good due to the limitation of the actual effective receptive field. The training process is as follows: 1) Build a YOLOv4 network model and use the Initialization function to initialize the weight parameters of each layer of the neural network; YOLOv4 consists of four parts: (1) Input: refers to the original sample data input into the network; (2) BackBone network: refers to the convolutional neural network structure that performs feature extraction operations; (3) Neck: fuses the image features extracted by the backbone network and passes the fused features to the prediction layer; (4) Head: predicts the target object of interest in the image and generates a visual prediction box and target category; After downloading the standard YOLOv4 network, use the make command to compile the YOLOv4 network to form an executable file darknet; edit the kitti7.data file in the cfg folder for the road target dataset KITTI-7classes, and change the class, train, valid, and names strings to the directory and parameters of the corresponding dataset. In this way, the parameters required for the Input part of the standard YOLOv4 network are edited. After setting epoch in the command line for training execution, load kitti7.data according to the experimental dataset and load yolov4.cfg at the same time, and the program can start training; the program will use the Initialization function to initialize the weight parameters of each layer of the neural network when it is running; 2) Input image data from the Input part, pass through the Backbone part, and finally output feature maps of two scales, and use the classifier to output the prediction box Pb 1 and classification probability CP x ; The image data is input from the Input part, passed through the Backbone part, and finally output the feature maps of two scales. The feature maps of two different scales are sent to the Neck part composed of the Feature Pyramid Network (FPN), and the fused features are passed to the prediction layer. At the same time, the Head part completes the classification of the target and outputs the prediction box Pb 1 and classification probability CP x , where x is the index of each category; 3) Perform IoU and NMS post-processing on these data and convert the predicted box Pb 2 Compare with the real frame Gtb, and use the Adam algorithm to update the weights of each layer of the neural network; The prediction box Pb generated by the Backbone network 1 The number is too large, and there are a large number of detection boxes for the same object in the picture, resulting in redundant detection results; the Head part of YOLOv4 will simultaneously complete the prediction box and its corresponding classification probability; IoU and NMS post-processing are performed on these data to obtain processed data; the IoU and NMS used here are the CIoU_loss and NMS of the standard YOLOv4; after these post-processing, the prediction box Pb of the target of interest can be obtained 2 The corresponding classification probability CP x ; At the same time, the Adam algorithm is used to update the weights of each layer of the neural network using the loss obtained in the post-processing process; 4) Loop through steps 2) and 3) and continue iterating until the epoch value specified in the command is reached, stop training, and output a file Q recording the weights and offsets of each layer. 1 ; Using Q 1 The obtained weights and offsets are used to detect the test set, and the mAP, Recall and frame rate FPS during detection are calculated; According to the accuracy requirements, the iteration threshold epoch=100 is set. When the number of iterations is less than the threshold, the Adam algorithm is used to update the weights of each layer of the network until the threshold epoch=100 stops training, calculates mAP and Recall, and outputs a file Q recording the weights and offsets of each layer. 1 ; Step 3: Design a Dense-YOLOv4 network model; design two cross-layer fusion modules in the model, namely the Dense-SPP module and the Dense-feature fusion module; take YOLOv4 as the basic network and introduce the above two modules into the YOLOv4 model; On the skeleton network, CSPDarknet-53 is used as the feature extraction network. On the feature fusion network, the original SPP module is replaced with a Dense-SPP module. At the same time, the single-channel five-layer convolution module on the Path Aggregation Network (PAN) structure that fuses the output feature map of the previous scale is replaced with the designed Dense-feature fusion module. In terms of the network detector, the original three-scale detection is adopted. The input size of the network is 640×640×3, and the sizes of the feature maps of the final detection layer are 20×20, 40×40, and 80×80, respectively, to detect large, medium and small objects; (a) Design a Dense-SPP module, introduce a cross-layer connection module based on the original SPP module, so that the feature map is divided into two branches, one of which is convolved and pooled by the original SPP module, and the other is convolved with 1×1×512 single convolution, and then the output feature maps of the two branches are concatenated; the number of convolution kernels of the Dense-SPP module is 11264, while the number of convolution kernels of the original SPP module is 20480, which reduces the number of parameters by 45%; the designed Dense-SPP module also uses 5 CBL modules, but the difference is that a cross-layer connection is added to integrate the previous convolution information and retain more original information; secondly, the network is hierarchical, that is, for the same task, different samples can be detected with different types of features; the shallow network extracts simple features, including texture features, while different samples require features of different complexity for judgment. Without the Concat operation, the network does not save the features extracted by the previous shallow network; After adding the Concat operation, it is equivalent to splicing the feature information of the first layer of the module on the output, making it valid; (b) Design a Dense-feature fusion module. Entering the Dense-feature fusion module, the features are divided into two branches, one of which undergoes four convolutions with kernel sizes of 1×1×256, 3×3×256, 1×1×256, and 3×3×256, respectively. The other feature map undergoes a single convolution of 1×1×256, and then the features of the two branches are concat-joined. Compared with the single-channel quintuple convolution in the original YOLOv4, the Dense-feature fusion module has less computational complexity. The number of convolution kernel parameters is 5376, while the number of convolution kernel parameters of the single-channel quintuple convolution is 9984, which reduces the computational complexity by 40%. Since each convolution wastes some information: the inhibitory effect of the activation function, the randomness of the convolution kernel parameters, , this cross-layer connection is equivalent to directly taking the previously processed information and processing it together now, which has the effect of reducing loss; and the Concat splicing operation is equivalent to splicing the feature information of the first layer of the module on the output, realizing feature reuse; the output of each layer of the dense cross-layer connection module will establish an input-output relationship with all the subsequent layers, and the input of each layer is the accumulation of all the previous layers. This mode can retain the simple features of the shallow layer of the network to the deep layer of the network, and fuse them with high semantic features to achieve feature reuse; this model can reduce the number of network parameters; the subsequent network layers obtain the gradient of the loss function and the original input signal, so that the network contains implicit deep supervision, which makes it easy to train deeper networks and has a regularization effect, alleviating the gradient disappearance problem during training; Step 4: Compare the test results of the model performance in step 2 and step 3, including model detection accuracy, model detection speed, model detection recall rate, and model weight file size, and check the images in the data sets actually detected in step 2 and step 3 to analyze the test results.

Citation Information

Patent Citations

  • Improved YOLOv4 vehicle pedestrian detection algorithm based on activation function

    CN114694104A