An Improved YOLOv4 Vehicle and Pedestrian Detection Algorithm Based on Activation Function

By introducing the FMish activation function in the YOLOv4 network, the problems of slow detection speed and insufficient accuracy of the YOLOv4 network are solved, and more efficient vehicle and pedestrian detection is achieved, which improves the detection performance on the KITTI dataset.

CN114694104BActive Publication Date: 2025-07-22XIDIAN UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210007093.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-05
Publication Date
2025-07-22
Estimated Expiration
2042-01-05

AI Technical Summary

Technical Problem

The existing YOLOv4 network has problems in vehicle pedestrian detection, large memory usage and insufficient detection accuracy, especially in specific categories of the KITTI data set.

Method used

Based on the network structure of Dense-YOLOv4 and Dense-YOLOv4-Small, the FMish activation function was introduced to replace all activation functions as the FMish activation function. The gradient of the FMish activation function does not change at the zero point and is a small negative gradient, avoiding gradient saturation and explosion problems, and improving training stability.

Benefits of technology

It improves detection accuracy and speed, improves the detection effect of the model, especially the mAP and Recall indicators on the KITTI-7 Classes dataset, reducing the memory usage of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694104B_ABST
    Figure CN114694104B_ABST
Patent Text Reader

Abstract

The present invention proposes an improved YOLOv4 vehicle and pedestrian detection algorithm based on activation functions; constructs an FMish activation function whose gradient does not mutate at zero points to ensure information flow and avoid the "gradient disappearance"; based on the Dense-YOLOv4 and Dense-YOLOv4-Small network structures, the present invention replaces all activation functions with FMish activation functions; the present invention eliminates "Misc" and "Dontcare" from the KITTI road target dataset to obtain the KITTI-7classes road target dataset, trains three models on the KITTI-7classes dataset, and compares the detection speed and detection performance; the network model based on the FMish activation function not only avoids the saturation problem but also avoids the "gradient explosion" problem, ensures the stability of the training process, and improves the detection effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of image recognition, and relates to an improved YOLOv4 vehicle and pedestrian detection algorithm based on an activation function, which shows good detection performance on a general standard data set. Background Art

[0002] With the continuous development of computer technology and the continuous improvement of computing power, computer vision and object detection therein have become popular directions in recent years. Object detection can be used to identify and locate specific objects, and has broad development prospects in driving assistance systems, military warning systems, etc. Object detection technologies include traditional object detection technologies and object detection technologies based on deep learning. The latter has become the mainstream algorithm in the current object detection field because it is superior to the former in terms of performance and complexity. In order to more efficiently manage traffic roads and maintain social stability, it is necessary to detect targets such as pedestrians and vehicles on the road. The detection task of vehicles and pedestrians occupies an important position in the field of autonomous driving. Intelligent vehicle and pedestrian recognition can assist traffic police in effective management and traffic flow control, and can also predict the next traffic condition in time to prevent traffic congestion.

[0003] Based on the YOLOv4 network and the KITTI road target data set, the present invention constructs a vehicle and pedestrian detection algorithm with higher performance. Taking YOLOv4 as the basic network and drawing on the idea of DenseNet, a Dense-SPP module and a Dense-feature fusion module are designed, called Dense-YOLOv4, which can effectively perform multi-scale pooling on high-level features to increase the receptive field and more fully fuse the features of the high level of the network, while also reducing the computational amount of the network.

[0004] The data set used in the present invention is the KITTI road target data set. In order to make the model more lightweight while basically maintaining the detection accuracy, a Dense-YOLOv4-Small network model is designed, and an FMish activation function is constructed. Its gradient does not mutate at zero, but is a very small negative gradient, thus ensuring information flow. The YOLOv4, Dense-FMish-YOLOv4, and Dense-FMish-YOLOv4-Small models are trained on the KITTI road target data set, and the performance of the three models in terms of detection speed, mAP, and Recall metrics is compared. The FMish activation function can not only avoid the saturation problem, but also the function is relatively gentle, avoiding the problem of "gradient explosion", which can ensure the stability of the training process and improve the detection effect. Summary of the Invention

[0005] In view of the above problems, the purpose of the present invention is to provide an improved YOLOv4 vehicle and pedestrian detection algorithm based on activation function for the YOLOv4 network structure.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] An improved YOLOv4 vehicle and pedestrian detection algorithm based on activation function is proposed. On the basis of Dense-YOLOv4 and Dense-YOLOv4-Small network structures, an FMish activation function is constructed. Its gradient does not change suddenly at the zero point, but is a very small negative gradient. All activation functions are replaced by the FMish activation function, called Dense-FMish-YOLOv4 and Dense-FMish-YOLOv4-Small algorithms. The FMish activation function can not only avoid the saturation problem, but also has a relatively smooth function, avoiding the problem of "gradient explosion", which can ensure the stability of the training process and improve the detection effect.

[0008] The vehicle pedestrian detection algorithm comprises the following steps:

[0009] Step 1: Download the KITTI road target dataset, a general dataset in the current target detection field, remove the "Misc" and "Dontcare" data in the original KITTI dataset, and create the KITTI-7Classes road target dataset. Using this dataset can ensure that the algorithm detection effect is consistent with the general dataset disclosed in this field, and construct the road target dataset used in the present invention; divide the test set, validation set and training set in a ratio of 6:2:2;

[0010] The KITTI dataset is currently the largest dataset for autonomous driving scenarios. KITTI contains real image data collected from various road scenes. The KITTI dataset contains nine categories, namely Car, Van, Truck, Pedestrian, Person (sitting), Cyclist, Tram, Misc and Dontcare. Since there are two categories in KITTI, namely "Misc" and "Dontcare", which are "disorganized" and "don't care" respectively, these two categories are meaningless, and since these two categories have no specific target features, the objects that may be contained in the "Misc" class in different pictures are different. The present invention removes "Misc" and "Dontcare" from the original KITTI dataset to form the KITTI-7Classes dataset. The present invention will be trained and tested on KITTI-7Classes.

[0011] Step 2: Use the standard YOLOv4 network to train, identify, and locate vehicles and pedestrians; use the standard YOLOv4 network to train on the road target dataset based on Step 1. Download the standard YOLOv4 network and compile it. The download address of the standard YOLOv4 network is: https: / / github.com / AlexeyAB / darknet; for the road target data kitti-7classes, change the training set, validation set, and test set directories in the kitti7.data file in the cfg folder to the address of the downloaded dataset, specify the number of classes and class names, set the number of iterations (epoch) to 100 according to the accuracy requirements in the command line during training, load kitti7.data according to the experimental dataset this time, and load yolov4.cfg at the same time, then the program can start training; save the weight file Q1 of each layer during the training process as the weight input file for detection after training; use the weight file Q1 for testing to obtain the Mean Average Precision (mAP), Recall, and Frame Per Second (FPS) during detection.

[0012] 1) Construct the YOLOv4 network model and use the Initialization function to initialize the weight parameters of each layer of the neural network.

[0013] YOLOv4 consists of four parts, namely: (1) Input input end: refers to the original sample data input into the network; (2) Backbone network: refers to the convolutional neural network structure for feature extraction operations; (3) Neck neck: fuses the image features extracted by the backbone network and transfers the fused features to the prediction layer; (4) Head head: predicts the target objects of interest in the image and generates visual prediction boxes and target classes.

[0014] After downloading the standard YOLOv4 network, use the make command to compile the YOLOv4 network to form an executable file darknet; edit the kitti7.data file in the cfg folder for the road target dataset KITTI-7classes, and change the strings class, train, valid, and names to the corresponding dataset directories and parameters, so that the parameters required for the Input part of the standard YOLOv4 network are edited. After setting epoch in the command line during training, load kitti7.data according to the experimental dataset this time, and load yolov4.cfg at the same time, then the program can start training; the program will use the Initialization function to initialize the weight parameters of each layer of the neural network during operation.

[0015] 2) Input the image data. After passing through the Backbone part, finally output feature maps of two scales, and use the classifier to output the prediction box Pb1 and the classification probability CP x ;

[0016] Input the image data. After passing through the Backbone part, finally output feature maps of two scales. Send the feature maps of the two different scales into the Neck part composed of the Feature Pyramid Network (FPN), and transfer the fused features to the prediction layer. At the same time, the Head part completes the classification of the target and outputs the prediction box Pb1 and the classification probability CP x , where x is the index of each classification;

[0017] 3) Perform post-processing of IoU and NMS on these data, compare the prediction box Pb2 with the ground truth box Gtb, and use the Adam algorithm to update the weights of each layer of the neural network;

[0018] The number of prediction boxes Pb1 generated by the Backbone network is too large, and there are a large number of detection boxes for the same object in the image, resulting in redundant detection results; the Head part of YOLOv4 will simultaneously complete the prediction box and its corresponding classification probability; perform post-processing of IoU and NMS on these data to obtain the processed data; the IoU and NMS used here are the CIoU_loss and NMS of the standard YOLOv4; after these post-processings, the prediction box Pb2 of the target of interest and its corresponding classification probability CP can be obtained x ; At the same time, use the Adam algorithm to update the weights of each layer of the neural network using the loss obtained during the post-processing;

[0019] 4) Loop and execute steps 2) and 3) and continue to iterate until the epoch value set in the command, stop training, and output the file Q1 that records the weights and offsets of each layer; use the weights and offsets obtained from Q1 to detect the test set, and calculate the mAP, Recall, and the frame rate FPS during detection;

[0020] The present invention sets the iteration threshold epoch = 100 according to the accuracy requirement. When the number of iterations is less than the threshold, use the Adam algorithm to update the weights of each layer of the network until the threshold epoch = 100 to stop training, calculate the mAP and Recall, and output the file Q1 that records the weights and offsets of each layer;

[0021] YOLOv4 has good real-time performance. The model detection speed and the size of the model weight file are also very important evaluation indicators. The detection speed varies depending on the hardware configuration. In all the experiments of the present invention, the same hardware platform is used. The standard of the detection speed is the number of pictures detected per second. The detection of vehicle and pedestrian targets based on YOLOv4 shows that the model detection speed is not high and the memory occupancy is large. In order to further improve the detection speed and detection accuracy, the dense-FMish-YOLOv4 and Dense-FMish-YOLOv4-Small models based on the FMish activation function are designed;

[0022] Step 3: Design the FMish activation function so that the gradient of the function does not mutate at the zero point, but is a very small negative gradient, avoiding the saturation problem. Moreover, in the part where x>0, its gradient is slightly smaller than that of Mish. Compared with Mish, the FMish function is relatively gentle, which can ensure the stability of the training process;

[0023] The present invention designs the FMish activation function. The Mish activation function and the FMish formula designed by the present invention are as follows:

[0024] y Mish = x·tanh(ln(1 + e x ))

[0025] where x is the matrix transmitted by the Batch Normalization (BN) layer;

[0026] Based on the Dense-YOLOv4 and Dense-YOLOv4-Small network structures, the present invention introduces the FMish activation function and replaces all the activation functions with the FMish activation function, which are called the Dense-FMish-YOLOv4 and Dense-FMish-YOLOv4-Small algorithms;

[0027] Step 4: Compare the detection results of the model performances in Step 2 and Step 3, including the model detection accuracy, the model detection speed, the model detection recall rate, and the size of the model weight file, and view the images in the actual detection dataset in Step 2 and Step 3 to analyze the detection results;

[0028] Based on the Dense-YOLOv4 and Dense-YOLOv4-Small network structures, the present invention introduces the FMish activation function and replaces all the activation functions with the FMish activation function. The FMish activation function can not only avoid the saturation problem, but also the function is relatively gentle, avoiding the problem of "gradient explosion", which can ensure the stability of the training process and improve the detection effect. Description of the Drawings

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0030] Figure 1 is the flowchart of the method of the present invention;

[0031] Figure 2 is the flowchart of training using YOLOv4;

[0032] Figure 3 is the comparison chart of the FMish and Mish activation functions of the present invention;

[0033] Figure 4 is the structure diagram of the Dense-FMish-YOLOv4 model;

[0034] Figure 5 is the structure diagram of the Dense-FMish-YOLOv4-Small model;

[0035] Figure 6 is the bar chart of the comparison of the detection performance of the three models; 6(a) Bar chart of mAP comparison, 6(b) Bar chart of recall rate comparison;

[0036] Figure 7 is the comparison chart of the detection results of YOLOv4 and Dense-FMish-YOLOv4; 7(a) Detection result diagram of YOLOv4 for Picture A, 7(b) Detection result diagram of Dense-FMish-YOLOv4 for Picture A;

[0037] Figure 8 is the comparison chart of the detection results of YOLOv4 and Dense-FMish-YOLOv4-Small; 8(a) Detection result diagram of YOLOv4 for Picture B, 8(b) Detection result diagram of Dense-FMish-YOLOv4-Small for Picture B;

[0038] Figure 9 is the comparison of the detection speed performance of the three models;

[0039] Figure 10 is the mAP performance analysis of the three models; Specific Embodiments

[0040] To make the above and other objects, features, and advantages of the present invention more obvious, the following specifically presents the embodiments of the present invention and, in conjunction with the accompanying drawings, makes a detailed description as follows:

[0041] Figure 1 This is the specific flowchart of the method, which can be divided into four steps:

[0042] Step 1: Download the KITTI road object dataset, which is a general dataset in the current object detection field. Exclude the two types of data, "Misc" and "Dontcare", from the original KITTI dataset to create the KITTI-7Classes road object dataset. Using this dataset can ensure that the detection effect of the algorithm is consistent with the general dataset publicly available in this field, and thus construct the road object dataset used in the present invention. Divide the test set, validation set, and training set in a ratio of 6:2:2;

[0043] The KITTI dataset is currently the largest dataset in the field of autonomous driving scenarios; the KITTI dataset contains real image data collected from various road scenarios; the KITTI dataset contains a total of nine categories, namely Car, Van, Truck, Pedestrian, Person(sitting), Cyclist, Tram, Misc, and Dontcare; since there are two categories in KITTI, "Misc" and "Dontcare", which are the "disorganized" category and the "unconcerned" category respectively, and these two categories are meaningless, and since these two categories do not have specific target features, the objects included in the "Misc" category may be different in different pictures. The present invention excludes "Misc" and "Dontcare" from the original KITTI dataset to form the KITTI-7Classes dataset, and the present invention will train and test on the KITTI-7Classes dataset;

[0044] Step 2: Use the standard YOLOv4 network to train, identify, and locate vehicles and pedestrians; use the standard YOLOv4 network to train the road target dataset based on Step 1. Download the standard YOLOv4 network and compile it. The download address of the standard YOLOv4 network is: https: / / github.com / AlexeyAB / darknet. Change the training set, validation set, and test set directories in the kitti7.data file in the cfg folder for the road target data kitti-7classes to the address of the downloaded dataset, specify the number of classes and class names, set the number of iterations (epoch) to 100 according to the accuracy requirements in the command line during training, load kitti7.data according to the experimental dataset this time, and load yolov4.cfg at the same time, then the program can start training; save the weight file Q1 of each layer during the training process as the weight input file for detection after training; use the weight file Q1 for testing to obtain the Mean Average Precision (mAP), Recall, and Frame Per Second (FPS) during detection.

[0045] Refer to Figure 2 , the training process can be divided into four steps:

[0046] 1) Build the YOLOv4 network model and use the Initialization function to initialize the weight parameters of each layer of the neural network;

[0047] YOLOv4 consists of four parts, namely: (1) Input input end: refers to the original sample data input into the network; (2) BackBone network: refers to the convolutional neural network structure for feature extraction operations; (3) Neck neck: fuses the image features extracted by the backbone network and transfers the fused features to the prediction layer; (4) Head head: predicts the target objects of interest in the image and generates visual prediction boxes and target classes;

[0048] After downloading the standard YOLOv4 network, use the make command to compile the YOLOv4 network to form an executable file darknet; edit the kitti7.data file in the cfg folder for the road target dataset KITTI-7classes, and change the strings of class, train, valid, and names to the directories and parameters of the corresponding dataset. In this way, the parameters required for the Input part of the standard YOLOv4 network are edited. After setting the epoch in the command line for training, load kitti7.data according to the experimental dataset of this time, and at the same time load yolov4.cfg, and the program can start training; when the program runs, it will use the Initialization function to initialize the weight parameters of each layer of the neural network;

[0049] 2) Input the image data from the Input part, pass through the Backbone part, finally output feature maps of two scales, and use the classifier to output the prediction box Pb1 and the classification probability CP x ;

[0050] Input the image data from the Input part, pass through the Backbone part, finally output feature maps of two scales, send the feature maps of the two different scales into the Neck part composed of the Feature Pyramid Network (FPN), and transfer the fused features to the prediction layer. At the same time, the Head part completes the classification of the target and outputs the prediction box Pb1 and the classification probability CP x , where x is the index of each classification;

[0051] 3) Perform post-processing of IoU and NMS on these data, compare the prediction box Pb2 with the ground truth box Gtb, and use the Adam algorithm to update the weight values of each layer of the neural network;

[0052] The number of prediction boxes Pb1 generated by the Backbone network is too large, and there are a large number of detection boxes for the same object in the image, resulting in redundant detection results; the Head part of YOLOv4 will simultaneously complete the prediction box and its corresponding classification probability; perform post-processing of IoU and NMS on these data to obtain the processed data; the IoU and NMS used here are the CIoU_loss and NMS of the standard YOLOv4; after these post-processings, the prediction box Pb2 of the target of interest and its corresponding classification probability CP can be obtained x ; At the same time, use the Adam algorithm to update the weight values of each layer of the neural network using the loss obtained during the post-processing;

[0053] 4) Loop through steps 2) and 3) and continue iterating until the epoch value set in the command is reached, stop training, and output the file Q1 that records the weights and offsets of each layer; use the weights and offsets obtained in Q1 to test the test set, and calculate the mAP, Recall, and FPS during testing;

[0054] The present invention sets the iteration threshold epoch=100 according to the accuracy requirement. When the number of iterations is less than the threshold, the Adam algorithm is used to update the weights of each layer of the network until the threshold epoch=100 stops training, calculates mAP and Recall, and outputs a file Q1 that records the weights and offsets of each layer;

[0055] The most basic network performance evaluation indicators are divided into four categories, namely TP (True Positives): positive samples are correctly identified as positive samples, that is, dogs are correctly identified as dogs; TN (True Negatives): negative samples are correctly identified as negative samples, that is, cats are correctly identified as cats; FP (False Positives): negative samples are incorrectly identified as positive samples, that is, cats are incorrectly identified as dogs; FN (False Negatives): positive samples are incorrectly identified as negative samples, that is, dogs are incorrectly identified as cats; Accuracy represents the ratio of the number of correctly predicted samples to the total number of samples, which is used to evaluate the overall accuracy of the algorithm model. The calculation method is Precision is the ratio of the number of correctly identified samples to the total number of identified samples. The calculation method is: The recall rate is the proportion of samples correctly identified as positive examples to all positive samples. The calculation method is: An algorithm model with good performance should maintain a high recall rate while ensuring a high accuracy rate. The Precision-Recall (PR) curve is used to show the trade-off between the accuracy and recall rate of the algorithm model. AP refers to the area enclosed by the PR curve plotted by the accuracy and recall rate obtained at a certain threshold and the horizontal and vertical axes, which measures the detection performance of the model in each category. mAP refers to the average AP of multiple target categories, which is used to measure the detection performance of the algorithm model on all tested categories. If there are N categories, the calculation method of mAP is The present invention mainly uses the model overall evaluation index mAP and Recall as the main evaluation indicators;

[0056] YOLOv4 has good real-time performance. The model detection speed and the size of the model weight file are also very important evaluation indicators. The detection speed varies depending on the hardware configuration. In all the experiments of the present invention, the same hardware platform is used. The standard of the detection speed is the number of pictures detected per second. The detection of vehicle and pedestrian targets based on YOLOv4 shows that the model detection speed is not high and the memory occupancy is large. In order to further improve the detection speed and detection accuracy, the dense-FMish-YOLOv4 and Dense-FMish-YOLOv4-Small models based on the FMish activation function are designed;

[0057] Step 3: Design the FMish activation function so that the gradient of the function does not mutate at the zero point but is a very small negative gradient, avoiding the saturation problem. Moreover, in the part where x>0, its gradient is slightly smaller than that of Mish. Compared with Mish, the FMish function is relatively gentle, which can ensure the stability of the training process;

[0058] Refer to Figure 3 : The present invention designs the FMish activation function. The Mish activation function and the FMish formula designed by the present invention are as follows:

[0059] y Mish = x·tanh(ln(1 + e x )),

[0060] where x is the matrix transmitted by the Batch Normalization (BN) layer;

[0061] Refer to Figure 4 and Figure 5 : Based on the Dense-YOLOv4 and Dense-YOLOv4-Small network structures, the present invention introduces the FMish activation function and replaces all the activation functions with the FMish activation function, which are called the Dense-FMish-YOLOv4 and Dense-FMish-YOLOv4-Small algorithms;

[0062] Step 4: Compare the detection results of the model performances in Step 2 and Step 3, including the model detection accuracy, the model detection speed, the model detection recall rate, and the size of the model weight file, and view the images in the actual detection dataset in Step 2 and Step 3 to analyze the detection results;

[0063] Based on the Dense-YOLOv4 and Dense-YOLOv4-Small network structures, the present invention introduces the FMish activation function and replaces all activation functions with the FMish activation function. The FMish activation function can not only avoid the saturation problem, but also is relatively gentle, avoiding the problem of "gradient explosion", which can ensure the stability of the training process and improve the detection effect.

[0064] The present invention constructs the FMish activation function, whose gradient does not mutate at zero point, but is a very small negative gradient, thus ensuring the information flow. The FMish activation function can not only avoid the saturation problem, but also is relatively gentle, avoiding the problem of "gradient explosion", which can ensure the stability of the training process and improve the detection effect.

[0065] The following further describes the invention with reference to simulation examples.

[0066] Simulation example:

[0067] The present invention uses the original YOLOv4 as a comparison sample, and both the training dataset and the test dataset come from the general dataset KITTI dataset to verify the universality of the algorithm for different datasets.

[0068] Figure 9 The AP values of seven classes, namely Car, Van, Truck, Pedestrian, Person(sitting), Cyclist, and Tram, in the KITTI-7classes dataset for three network models, YOLOv4, Dense-FMish-YOLOv4, and Dense-FMish-YOLOv4-Small, are given. From Figure 9 It can be seen that the test accuracies of the seven classes given by the Dense-FMish-YOLOv4 method are all better than those of the YOLOv4 method. The model Dense-FMish-YOLOv4-Small after a large number of pruning is very close to the performance of YOLOv4, and the calculation speed is greatly improved.

[0069] Performance comparison of the model after introducing the FMish activation function Figure 6 As shown, referring to Figure 6: The mAP value of the original YOLOv4 model is 89.1%, and the Recall is 89.5%. After introducing the FMish activation function, the mAP of the Dense-FMish-YOLOv4 model is increased by 1.4% to reach 90.5%, and the Recall is increased by 1.6% to reach 91.1%. Therefore, the FMish activation function proposed in this section can effectively improve the detection effect and accuracy of the network model. After a large number of pruning operations, the model Dense-FMish-YOLOv4-Small, due to the introduction of the FMish activation function, has an mAP value of 88.5%, and its performance is very close to that of the YOLOv4 model. Its Recall value reaches 90.2%, which is 0.7% higher than that of the original YOLOv4 model.

[0070] Figure 7 The comparison chart of the actual detection effects of the Dense-FMish-YOLOv4 algorithm model and the original YOLOv4 model is given. The same picture is respectively detected on the two models of the original YOLOv4 and Dense-FMish-YOLOv4. The detection confidence levels of the two cars on the left and right of the original YOLOv4 are 95% and 74% respectively, and the detection confidence levels of the two cars on the left and right of the Dense-FMish-YOLOv4 are 96% and 88% respectively, with an increase of 1% and 14% respectively. This shows that the FMish activation function designed in the present invention can not only avoid the saturation problem, but also the function is relatively gentle, avoiding the problem of "gradient explosion", which can ensure the stability of the training process and improve the detection effect.

[0071] Figure 8 The comparison chart of the actual detection effects of Dense-FMish-YOLOv4-Small and the original YOLOv4 model is given. The same picture is respectively detected on the two models of the original YOLOv4 and Dense-FMish-YOLOv4-Small. In Figure 8 (a), the original YOLOv4 model misidentifies a "Tram" tram in the picture as two, resulting in a misdetection problem. And in Figure 8In Figure (b), the Dense-FMish-YOLOv4-Small model can be recognized normally without any misdetection problems. This is because the Dense cross-layer fusion module designed in the present invention can fuse the information of the previous convolution, perform cross-layer fusion of the extracted features in the network, make the network more hierarchical, and improve the detection accuracy and effect. Moreover, for the Dense-YOLOv4-Small network after pruning the Dense-YOLOv4 network, the redundant calculations are removed, and the effective calculations are retained. The network detection speed is increased, but the detection accuracy and confidence do not decrease. The FMish activation function designed in this invention can also avoid the "gradient explosion", make the training process more stable, and improve the detection accuracy and effect.

[0072] Figure 10 The comparison relationships of the processing speeds, total parameter numbers, and memory usages of the three network models of YOLOv4, Dense-FMish-YOLOv4, and Dense-FMish-YOLOv4-Small on the KITTI-7classes dataset are given. It can be seen from the table that the processing speeds of YOLOv4 and Dense-FMish-YOLOv4 are not much different. The processing speed of Dense-FMish-YOLOv4 is slightly faster than that of YOLOv4 and the memory usage is slightly less than that of YOLOv4. The processing speed and memory usage of Dense-FMish-YOLOv4-Small are significantly better than those of YOLOv4.

[0073] In summary, the simulation results show that compared with the original YOLOv algorithm model, the performances of the Dense-FMish-YOLOv4 and Dense-FMish-YOLOv4-Small proposed in the present invention based on the FMish activation function have been significantly improved. The AP values of the Dense-FMish-YOLOv4 in 7 classes have been significantly improved. The number of residual structures reduced in the Dense-FMish-YOLOv4-Small eliminates the redundant calculations of the network without causing a significant decrease in the overall performance of the network. The model detection speed has been greatly improved, and the memory usage has been greatly reduced.

Claims

1. An improved YOLOv4 vehicle and pedestrian detection method based on activation function, for the detection and localization of road target datasets based on the Kitti-7classes general dataset, characterized in that: Use the KITTI road target dataset for vehicle and pedestrian detection. KITTI contains real image data collected from various road scenes. The KITTI dataset contains a total of nine categories, namely Car, Van, Truck, Pedestrian, Person sitting, Cyclist, Tram, Misc, and Dontcare. Since there are two categories in KITTI, "Misc" and "Dontcare", which are the "chaotic" category and the "unconcerned" category respectively, these two categories are meaningless. And because these two categories have no specific target features, the objects included in the "Misc" category are different in different pictures. Remove the "Misc" and "Dontcare" from the original KITTI dataset to form the KITTI-7Classes dataset, and train and test on the KITTI-7Classes. The vehicle and pedestrian detection algorithm includes the following steps: Step 1: Download the KITTI road target dataset, which is a general dataset in the current object detection field, and create the KITTI-7Classes road target dataset. Using this dataset can ensure that the detection effect of the algorithm is consistent with the general datasets publicly available in this field, and construct the used road target dataset; divide the test set, validation set, and training set in a ratio of 6:2:

2. Step 2: Use the standard YOLOv4 network to train, identify, and locate vehicles and pedestrians; use the standard YOLOv4 network to train on the road target dataset based on Step 1, download the standard YOLOv4 network and compile it; change the training set, validation set, and test set directories in the kitti7.data file in the cfg folder for the road target data kitti-7classes to the address of the downloaded dataset, specify the number of categories and category names, set the number of iterations to 100 according to the accuracy requirements in the command line during training, load kitti7.data according to the experimental dataset this time, and at the same time load yolov4.cfg, and the program can start training; save the weight file Q1 of each layer during the training process as the weight input file for detection after training ends. Use the weight file Q1 for testing to obtain the mean average precision mAP, recall Recall, and the frame rate FPS during detection. 1) Build a YOLOv4 network model and use the Initialization function to initialize the weight parameters of each layer of the neural network. YOLOv4 consists of four parts, namely: (1) Input: refers to the original sample data input into the network; (2) BackBone network: refers to the convolutional neural network structure for feature extraction operations; (3) Neck: fuses the image features extracted by the backbone network and transfers the fused features to the prediction layer; (4) Head: predicts the target objects of interest in the image and generates visual prediction boxes and target categories; After downloading the standard YOLOv4 network, use the make command to compile the YOLOv4 network to form an executable file darknet; edit the kitti7.data file in the cfg folder for the road target dataset KITTI-7classes, and change the strings of class, train, valid, and names to the directories and parameters of the corresponding dataset. In this way, the parameters required for the Input part of the standard YOLOv4 network are edited. After setting the epoch in the command line for training, load kitti7.data according to the experimental dataset this time, and at the same time load yolov4.cfg, and the program can start training; when the program runs, it will use the Initialization function to initialize the weight parameters of each layer of the neural network; 2) Input the image data from the Input section. After passing through the Backbone section, finally output feature maps of two scales, and use the classifier to output the prediction box Pb1 and the classification probability CP x ; Input the image data. After passing through the Backbone part, finally output feature maps of two scales. Send the feature maps of the two different scales into the Neck part composed of a Feature Pyramid Network, and transfer the fused features to the prediction layer. At the same time, the Head part completes the classification of the target and outputs the prediction box Pb1 and the classification probability CP x , where x is the index of each classification; 3) Perform post-processing of IoU and NMS on these data, compare the predicted box Pb2 with the ground truth box Gtb, and use the Adam algorithm to update the weight of each layer of the neural network; The number of predicted bounding boxes Pb1 generated by the Backbone network is too large, and there are a large number of detection bounding boxes for the same object in the image, resulting in redundant detection results; the Head part of YOLOv4 will simultaneously complete the prediction of the bounding box and its corresponding classification probability; after performing IoU and NMS post-processing on these data, the processed data is obtained; the IoU and NMS used here are the CIoU_loss and NMS of the standard TOLOv4; after these post-processings, the predicted bounding box Pb2 of the target of interest and its corresponding classification probability CP can be obtained x ; at the same time, the Adam algorithm is used to update the weights of each layer of the neural network using the loss obtained during the post-processing 4) Loop through steps 2) and 3) and continue to iterate until the epoch value set in the command, stop training, and output the file Q1 that records the weights and offsets of each layer; use the weights and offsets obtained from Q1 to detect the test set, and calculate the mAP, Recall, and the frame rate FPS during detection; Set the iteration threshold epoch = 100 according to the accuracy requirement. When the number of iterations is less than the threshold, use the Adam algorithm to update the weights of each layer of the network until the threshold epoch = 100 to stop training, calculate the mAP and Recall, and output the file Q1 that records the weights and offsets of each layer; YOLOv4 has good real-time performance. The model detection speed and the size of the model weight file are also very important evaluation indicators; the detection speed varies depending on the hardware configuration. In all experiments, the same hardware platform is used. The standard for the detection speed is the number of pictures detected per second. The detection of vehicle and pedestrian targets based on YOLOv4 shows that the model detection speed is not high and the memory occupancy is large. In order to further improve the detection speed and detection accuracy, the dense-FMish-YOLOv4 and Dense-FMish-YOLOv4-Small models based on the FMish activation function are designed; Step 3: Design the FMish activation function so that the gradient of the function does not mutate at zero, which can ensure the stability of the training process; The FMish activation function is designed. The Mish activation function and the designed FMish formula are as follows: where x is the matrix passed by the batch normalization layer; Based on the Dense-YOLOv4 and Dense-YOLOv4-Small network structures, the FMish activation function is introduced, and all activation functions are replaced with the FMish activation function, which are called the Dense-FMish-YOLOv4 and Dense-FMish-YOLOv4-Small algorithms; Step 4: Compare the detection results of the model performance in Step 2 and Step 3, including the model detection accuracy, model detection speed, model detection recall rate, and model weight file size, and view the images in the actual detection dataset in Step 2 and Step 3 to analyze the detection results.

Citation Information

Patent Citations

  • Cross-layer fusion improved YOLOv4 road target recognition algorithm

    CN114565896A