A YOLOv4 target detection algorithm with improved loss function

By introducing the improved loss function SCIoU in YOLOv4, the problem of large regression loss in CIoU_Loss when processing small targets is solved, and the detection performance is improved and the stability of model training is achieved.

CN114463718BActive Publication Date: 2025-05-23BAODING RUIGEXING TECHNOLOGY SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210007089.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-05
Publication Date
2025-05-23
Estimated Expiration
2042-01-05

AI Technical Summary

Technical Problem

When CIoU_Loss loss function deals with small targets, the regression loss is also large when the target aspect ratio changes greatly, which is not conducive to the stability of model training.

Method used

A loss function SCIoU based on CIoU_Loss is proposed, and it is embedded in YOLOv4 by the improved aspect ratio metric νs and the arctan function is replaced by the Sigmoid function prototype.

Benefits of technology

The detection performance of YOLOv4 algorithm is improved, the values ​​of mAP and Recall are improved, and no more computational volume is introduced, and the real-time performance is not affected. It is suitable for replacing the loss function of other object detection algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114463718B_ABST
    Figure CN114463718B_ABST
Patent Text Reader

Abstract

The present invention proposes a YOLOv4 object detection algorithm with an improved loss function, which improves the position regression loss function CIoU_Loss in the standard YOLOv4 model, proposes a new loss function SCIoU, embeds it into YOLOv4, and obtains performance improvement; first, download the general datasets tt100k and LISA in the current object detection field and perform data augmentation; secondly, use the standard YOLOv4 network to train the two augmented general datasets and detect their performance; then, for the position regression loss function CIoU_Loss in the standard YOLOv4 model, propose an improved loss function SCIoU and embed it into the YOLOv4 model for training; finally, compare with the standard YOLOv4 algorithm and analyze the test results; the YOLOv4 algorithm improved based on SCIoU proposed by the present invention includes improving the aspect ratio measurement index v s , and replacing the arctan function with the Sigmoid function; embedding it into YOLOv4, obtaining performance improvement, this model does not introduce more computational complexity, and the real-time performance is not affected. The improved YOLOv4 algorithm of the present invention has good robustness and can be used for performance improvement of multiple datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of image recognition, and in particular relates to a YOLOv4 target detection algorithm with an improved loss function, which exhibits good detection performance on a general standard dataset. Background Art

[0002] With the continuous development of computer technology, computer vision and target detection have become popular directions. Target detection can be used to identify and locate specific objects, and has broad development prospects in driving assistance systems and military early warning systems. Target detection technology includes traditional target detection technology and target detection technology based on deep learning. The latter has become the mainstream algorithm in the current target detection field because it is superior to the former in terms of performance and complexity.

[0003] Object detection technology based on deep learning is mainly divided into two methods: one-stage and two-stage. In the first stage of the two-stage method, candidate regions are defined for the input image, and in the second stage, convolutional neural networks are used to classify the candidate regions. Typical algorithms include R-CNN and Fast-R-CNN. This algorithm has high accuracy, but because two sub-networks are used to complete a single object detection task, the training cost and detection cost are high, and the speed is slow. The one-stage method divides the input image into a fixed number of patches, and each patch has a fixed number of anchor boxes. The position and classification label of the anchor box are output at the same time. Typical algorithms include SSD512 and YOLOv4. Although the one-stage method is slightly less accurate than the two-stage method, it only uses one network to complete the detection work, with low training cost and detection cost, and fast speed, which is suitable for scenarios that require fast response.

[0004] The loss function is an important part of the target detection function. It is a function used to measure the difference between the output prediction value and the true value. By calculating the loss function value, the error between the prediction result and the true result generated during the model training process is obtained to guide the learning of network parameters and achieve the purpose of optimizing the algorithm model.

[0005] Usually, target detection has two major tasks: classification and position regression. Therefore, the loss function of target detection is generally composed of two parts: classification loss and position regression loss. For classification problems, the commonly used loss functions in the algorithm model are cross entropy loss (CE_Loss) and Softmax_Loss; for regression problems, the development process of position regression loss in the algorithm model is: Smooth·L1_Loss, IoU_Loss, GIoU_Loss, DIoU_Loss, CIoU_Loss. IoU_Loss uses the intersection over union (IoU) as the loss function, but the gradient cannot be returned when the two boxes do not intersect, and when different areas of the two boxes overlap, the degree of overlap cannot be distinguished; in response to the above shortcomings, GIoU_Loss was proposed, which introduced the minimum area of ​​the two boxes, but when the real box completely contains the predicted box, it will degenerate into IoU_Loss, slowing down the convergence; in order to further improve the regression accuracy, DIoU_Loss was proposed, which introduced the Euclidean distance between the center points of the two boxes; on this basis, the aspect ratio of the predicted box and the real box was taken into account, and CIoU_Loss was proposed. However, when the CIoU_Loss loss function is used to process small targets, the regression loss is also large when the target aspect ratio changes greatly, which is not conducive to the stability of model training. The present invention proposes a loss function SCIoU (Sigmoid Complete Intersection over Union) based on CIoU_Loss and embeds it into YOLOv4 to improve performance. At the same time, the model does not introduce more calculation amount, and the real-time performance is not affected. It can also replace the loss function of other target detection algorithms and has good applicability. Summary of the invention

[0006] The method of the present invention proposes a YOLOv4 target detection algorithm with an improved loss function, and improves the detection performance of the YOLOv4 algorithm by embedding an improved new loss function SCIoU.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] Step 1: Download the tt100k dataset and LISA dataset, which are common datasets in the current target detection field. Using these two datasets can ensure that the detection effect of the algorithm is consistent with the common datasets disclosed in this field and verify the actual effect of the algorithm; enhance the downloaded data, including flipping, cropping, adding noise, and rotating operations. The data generated after enhancement can not only increase the number of images contained in the dataset, but also because the enhanced images are more complex than the original images in the dataset, the style and size of the images are changed while retaining the feature points of the original images, and the blurriness of the images is increased, making the enhanced images more diverse and closer to the actual situation, which can improve the robustness of the trained network; the download address of the tt100k dataset is: http: / / cg.cs.tsinghua.edu.cn / traffic-sign / ; the download address of the LISA dataset is: http: / / cvrr.ucsd.edu / LISA / lisa-traffic-sign-dataset.html;

[0009] The full name of tt100k is Tsinghua-Tencent 100K, which is a general dataset of road traffic signs that can be used to identify provided by the Tsinghua-Tencent Internet Innovation Technology Joint Laboratory. The resolution of the images in the TT100K dataset is 2048×2048, and there are 221 sign categories, which are roughly divided into three categories: warning signs, prohibition signs, and instruction signs. The dataset covers traffic sign images under different weather conditions and different lighting conditions, of which the training set contains 6105 images and the test set contains 3071 images. Due to the large resolution of the original images, the original images were cropped in the experiment of this paper, and the scale of the cropped images was 608×608. Due to the serious imbalance of the amount of data between each category in the dataset, this experiment only selected 45 types of traffic signs with a large amount of labeled data for recognition, and divided the test set, validation set and training set in a ratio of 6:2:2, and flipped, cropped, denoised, and rotated each image.

[0010] LISA stands for Laboratory for Intelligent & Safe Automobiles. It is a general dataset for identifying road traffic signs provided by the LISA laboratory in the United States. By driving a vehicle to shoot a video, a certain segment with a traffic sign is extracted from the video, and then up to 30 frames are extracted based on this segment, and each frame of the video image is annotated. The annotation of each traffic sign contains four parts of information: Tag, Position, Occluded, and On side rode. The process of collecting pictures is extracted from the video. The vehicle has a certain speed during driving and is not stationary, so it is blurred, which also makes the traffic sign recognition algorithm based on this dataset more applicable to real scenes. The LISA dataset in the United States contains 47 categories, but the number of annotations between categories is seriously unbalanced. Therefore, in order to ensure data availability, this experiment will select four categories with a large number of annotations for training and testing. The test set, validation set and training set are divided in a ratio of 6:2:2, and each image is flipped, cropped, denoised, and rotated.

[0011] Step 2: Use the standard YOLOv4 network to train and detect traffic signs; Use the standard YOLOv4 network to train the two traffic sign data sets based on step 1 respectively, download the standard YOLOv4 network and compile it, the standard YOLOv4 network download address: https: / / github.com / AlexeyAB / darknet, change the training set, validation set, and test set directories in the tt100k.data and LISA.data files in the cfg folder to the addresses of the downloaded data sets for the two traffic sign data sets tt100k and LISA respectively, and specify the number of categories and category names; Set epoch = 20000 according to the accuracy requirements, load tt100k.data or LISA.data according to this experimental data set, and load yolov4.cfg at the same time, and the program can start training; Use the loss function CIoU_Loss of the standard YOLOv4 network during training; Save the weight files Q of each layer during training 1 , as the weight input file for detection after training; using the weight file Q 1 Tests were conducted to obtain mAP and Recall. The obtained mAP, Recall and loss during training were analyzed. It was found that when the CIoU_Loss loss function was processing small targets, the regression loss was also large when the target aspect ratio changed greatly, which was not conducive to the stability of model training.

[0012] 1) Build a YOLOv4 network model and use the Initialization function to initialize the weight parameters of each layer of the neural network;

[0013] YOLOv4 is composed of four connected parts: (1) Input: refers to the original sample data input into the network; (2) BackBone: refers to the convolutional neural network structure that performs feature extraction operations; (3) Neck: fuses the image features extracted by the backbone network and passes the fused features to the prediction layer; (4) Head: predicts the target object of interest in the image and generates a visual prediction box and target category;

[0014] After downloading the standard YOLOv4 network, use the make command to compile the YOLOv4 network to form an executable file darknet; edit the tt100k.data and LISA.data files in the cfg folder for the two traffic sign datasets tt100k and LISA respectively, and change the class, train, valid, and names strings to the enhanced directories and parameters of the corresponding datasets. In this way, the parameters required for the Input part of the standard YOLOv4 network are edited. After setting the epoch, load tt100k.data or LISA.data according to the experimental dataset, and load yolov4.cfg at the same time, and the program can start training; the program will use the Initialization function to initialize the weight parameters of each layer of the neural network when it is running;

[0015] 2) Input image data from the Input part, pass through the Backbone part, and finally output feature maps of three scales, and use the classifier to output the prediction box Pb 2 and classification probability CP x ;

[0016] The image data is input from the Input part, passes through the Backbone part, and finally outputs feature maps of three scales. The feature maps of three different scales are sent to the Neck part composed of SPP and PANet, and the fused features are passed to the prediction layer. At the same time, the Head part completes the classification of the target and outputs the prediction box Pb 1 and classification probability CP x , where x is the index of each category; the prediction box Pb generated at this time 1 The number is too large, and there are a large number of detection frames for the same object in the picture, resulting in redundant detection results. After using the NMS algorithm to filter the prediction frames, the prediction frame Pb of the target of interest can be obtained. 2 The corresponding classification probability CP x ;

[0017] After the data enters the BackBone network from the Input, information extraction continues. The backbone network in the YOLOv4 network structure has 53 convolutional layers and outputs feature maps of three different scales. The feature maps of three different scales are sent to the Neck part composed of SPP and PANet to fuse and extract features, and the fused features are passed to the prediction layer. YOLOv4 is a one-stage target detection algorithm, so the Head part will simultaneously complete the prediction box and its corresponding classification probability;

[0018] YOLOv4 is a one-stage target detection algorithm. The Head part will simultaneously generate prediction boxes and their corresponding classification probabilities. However, the number of prediction boxes generated at this time is too large. There are a large number of detection boxes for the same object in the image, resulting in redundant detection results. It is necessary to perform non-maximum suppression on the redundant detection boxes so that each object in the image retains one prediction box. After the prediction boxes are screened by the NMS algorithm, the prediction boxes of the target of interest and their corresponding classification probabilities can be obtained.

[0019] 3) Compare the predicted box with the real box and calculate the IoU loss L CIoU , confidence loss L C , category loss L P , and add the three together to calculate the loss error Loss all =L CIoU +L C +L P And use the Adam algorithm to update the weights of each layer of the neural network;

[0020] The prediction box Pb obtained in 2) 2 Compared with the real box Gtb in the data set, the loss error is obtained. The loss error includes three parts, namely, IoU loss L CIoU , confidence loss L C , category loss L P , we will focus on the IoU loss L IoU ;

[0021] IoU loss L IoU After development, the more practical IoU loss can be summarized into the following four stages:

[0022] IoU_Loss, GIoU_Loss, DIoU_Loss, CIoU_Loss; IoU_Loss uses the intersection-over-union ratio (IoU) as the loss function. The disadvantage is that when the two boxes do not intersect or the two boxes have overlapping areas of different regions, they cannot be distinguished, resulting in gradient return errors; GIoU_Loss improves IoU_Loss. The disadvantage is that when the real box completely contains the predicted box, the IoU value between the two boxes is the same as the GIoU value, and the relative relationship cannot be distinguished; DIoU_Loss uses the Euclidean distance between the two center points to improve the shortcomings of GIoU_Loss; GIoU_Loss takes into account the aspect ratio of the predicted box and the real box, focusing on the regression problem of the predicted box under different aspect ratios when the real box completely wraps the predicted box, and proposes the concept of CIoU, which is expressed as:

[0023]

[0024] Where α is the weight function, and v is an indicator used to measure the similarity of aspect ratios;

[0025] The gradient of CIoU_Loss is similar to DIoU_Loss, but the gradient of v is also considered, and its expression is:

[0026]

[0027] When the length and width are in [0,1], w 2 +h 2 The value of is usually very small, which will cause gradient explosion, so in implementation Replace with 1;

[0028] The standard YOLOv4 uses CIoU_Loss in the loss function part, and the comprehensive expression of its loss function is:

[0029] Loss all =L COoI +L C +L P

[0030] Where L CIoU The expression is:

[0031]

[0032] L C The expression is:

[0033]

[0034] where λ cls is the normalized weight factor of confidence, λ cRepresents the weight coefficient of confidence loss to total loss; is the category gain factor;

[0035] L P The expression is:

[0036]

[0037] in Represents the weight coefficient of classification loss to total loss;

[0038] s 2 Indicates that the input image is divided into s×s cells, and B represents the total number of Anchor boxes contained in each cell. Indicates whether the jth Anchor box of the i-th cell is responsible for this object, that is, the cell contains the center point of the target object, and the jth Anchor box of the cell has the largest IoU value with the real box. Otherwise, 0; Indicates that the jth anchor box of the i-th cell is not responsible for the target; L SCIoU represents the SCIoU loss, C i represents the confidence of the i-th bounding box predicted by the network, and When the cell contains the center point of the target object, Pr(object) = 1, otherwise it is 0; represents the confidence of the i-th bounding box where the target is located; P i represents the probability that the i-th bounding box predicted by the network belongs to each category, Indicates the probability that the i-th bounding box where the target is located belongs to each category;

[0039] 4) Loop through steps 2) and 3) and continue iterating until the epoch value is reached, stop training, calculate mAP and Recall, and output a file Q that records the weights and offsets of each layer. 1 ;

[0040] The present invention sets the iteration threshold epoch=20000 according to the accuracy requirement. When the number of iterations is less than the threshold, the Adam algorithm is used to update the weights of each layer of the network until the threshold epoch=20000 stops training, calculates mAP and Recall, and outputs a file Q recording the weights and offsets of each layer. 1 ;

[0041] The most basic network performance evaluation indicators are divided into four categories, namely TP (True Positives): positive samples are correctly identified as positive samples; TN (True Negatives): negative samples are correctly identified as negative samples; FP (False Positives): negative samples are incorrectly identified as positive samples; FN (False Negatives): positive samples are incorrectly identified as negative samples; Accuracy represents the ratio of the number of correctly predicted samples to the total number of samples, which is used to evaluate the overall accuracy of the algorithm model. The calculation method is: Precision is the ratio of the number of correctly identified samples to the total number of identified samples. The calculation method is: The recall rate is the ratio of samples correctly identified as positive examples to all positive examples. The calculation method is: If an algorithm model has good performance, it should meet the following conditions: while ensuring a high accuracy rate, the recall rate should also be maintained at a high level; in order to more vividly represent this condition, the Precision-Recall (PR) curve is used to show the trade-off between the accuracy rate and the recall rate of the algorithm model; AP refers to the area enclosed by the PR curve graph drawn by the accuracy rate and the recall rate obtained at a certain threshold and the horizontal and vertical axes, which measures the performance of the model in each category. mAP refers to the average AP of multiple target categories, which is used to measure the detection performance of the algorithm model on all tested categories. If there are N categories, the calculation method of mAP is The present invention mainly uses the model overall evaluation index mAP and Recall as the main evaluation indicators;

[0042] Step 3. In order to solve the problem that when the CIoU_Loss loss function processes small targets, the regression loss is also large when the target aspect ratio changes greatly. A loss function SCIoU (Sigmoid Complete Intersection over Union) based on CIoU_Loss is proposed. By using the improved aspect ratio measurement index ν s The arctan function is replaced by the Sigmoid function prototype and embedded into YOLOv4, which improves the performance. At the same time, the model does not introduce more computational complexity, and the real-time performance is not affected. It can also replace the loss function of other target detection algorithms and has good applicability. The YOLOv4 network with the replaced loss function SCIoU is trained using the two data sets in step 1 to obtain the weight file Q 2; Using the weight file Q 2 Test and get mAP and Recall;

[0043] Based on CIoU, this paper proposes an improved aspect ratio measurement index v s , whose expression is:

[0044]

[0045] The arctan function is replaced by the Sigmoid function prototype, and the aspect ratio w / h is used as the independent variable of the function. It can be seen that the range of the arctan function is [0,π / 2), and the range of the Sigmoid function is [0.5,1). When the aspect ratio w / h is small, the change of Sigmoid is smaller than that of arctan as the w / h value changes. Corresponding to the expression of the aspect ratio metric v in the loss function, the aspect ratio of the predicted box and the true box is mapped and normalized. Sigmoid makes the regression loss change smaller. That is, this paper proposes the SCIoU loss function, which is conducive to the stability of model training.

[0046] The expression of SCIoU as a loss function is:

[0047]

[0048] Likewise s The gradient expression is:

[0049]

[0050] Similarly, in the experimental case, in order to prevent the gradient from exploding, The value of is 1;

[0051] The total loss in step 2 is Loss all =L CIoU +L C +L P , L COoU Replaced by L in the present invention SCIoU , and embed it into YOLOv4, and train it in the same way as in step 2, iterate to the epoch value and use the Adam algorithm to update the weights, calculate mAP and Recall, and save the weight file Q 2 ;

[0052] Step 4: Compare the mAP and Recall obtained in step 3 with those obtained in step 2, set different IoU_thresh settings for retraining, compare the mAP and Recall performance of step 2 and step 3 under different settings, and analyze the test results;

[0053] In the above steps, the value range of i is 0 to the number of divided cells s×s, and the value range of j is 0 to the number of divided anchor boxes B.

[0054] This paper proposes a loss function SCIoU (Sigmoid Complete Intersection over Union) based on CIoU_Loss and an improved aspect ratio metric v s The arctan function is replaced by the Sigmoid function prototype, and the loss function of the standard YOLOv4 is replaced to form a new neural network algorithm; this algorithm can improve the accuracy without affecting the real-time performance of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0056] Figure 1 is a flow chart of the method of the present invention;

[0057] Figure 2 This is the YOLOv4 network model structure diagram;

[0058] Figure 3 This is a schematic diagram of IoU_Loss;

[0059] Figure 4 This is a schematic diagram of GIoU_Loss;

[0060] Figure 5 This is a schematic diagram of DIoU_Loss;

[0061] Figure 6 This is a flowchart for training using YOLOv4;

[0062] Figure 7 This is a comparison chart of Sigmoid and arctan function images;

[0063] Figure 8 This is a comparison chart of some detection results of the original YOLOv4 and improved YOLOv4 models;

[0064] Fig. 9 It is the overall performance of the original YOLOv4 and improved YOLOv4 models on the tt100k validation dataset;

[0065] Fig.10It is the total loss function graph during model training;

[0066] Fig.11 It is the overall performance of the original YOLOv4 and improved YOLOv4 models on the LISA validation dataset. DETAILED DESCRIPTION

[0067] In order to make the above and other purposes, features and advantages of the present invention more obvious, the embodiments of the present invention are specifically cited below, and the accompanying drawings are used to provide a detailed description as follows:

[0068] Figure 1 The specific flow chart of this method can be divided into four steps:

[0069] Step 1: Download the tt100k dataset and LISA dataset, which are common datasets in the current target detection field. Using these two datasets can ensure that the detection effect of the algorithm is consistent with the common datasets publicly available in this field, and verify the actual effect of the algorithm. Enhance the downloaded data, including flipping, cropping, adding noise, and rotating operations; the data generated after enhancement can not only increase the number of images contained in the dataset, but also because the enhanced images are more complex than the original images in the dataset, the style and size of the images are changed while retaining the feature points of the original images, and the blur of the images is increased, making the enhanced images more diverse and closer to the actual situation, which can improve the robustness of the trained network; the download address of the tt100k dataset is: http: / / cg.cs.tsinghua.edu.cn / traffic-sign / ; the download address of the LISA dataset is: http: / / cvrr.ucsd.edu / LISA / lisa-traffic-sign-dataset.html;

[0070] The full name of tt100k is Tsinghua-Tencent 100K, which is a general dataset of road traffic signs that can be used to identify provided by the Tsinghua-Tencent Internet Innovation Technology Joint Laboratory. The resolution of the images in the TT100K dataset is 2048×2048, and there are 221 sign categories, which are roughly divided into three categories: warning signs, prohibition signs, and instruction signs. The dataset covers traffic sign images under different weather conditions and different lighting conditions, of which the training set contains 6105 images and the test set contains 3071 images. Due to the large resolution of the original images, the original images were cropped in the experiment of this paper, and the scale of the cropped images was 608×608. Due to the serious imbalance of the amount of data between each category in the dataset, this experiment only selected 45 types of traffic signs with a large amount of labeled data for recognition, and divided the test set, validation set and training set in a ratio of 6:2:2, and flipped, cropped, denoised, and rotated each image.

[0071] LISA stands for Laboratory for Intelligent & Safe Automobiles. It is a general dataset for identifying road traffic signs provided by the LISA laboratory in the United States. By driving a vehicle to shoot a video, a certain segment with a traffic sign is extracted from the video, and then up to 30 frames are extracted based on this segment, and each frame of the video image is annotated. The annotation of each traffic sign contains four parts of information: Tag, Position, Occluded, and On side rode. The process of collecting pictures is extracted from the video. The vehicle has a certain speed during driving and is not stationary, so it is blurred, which also makes the traffic sign recognition algorithm based on this dataset more applicable to real scenes. The LISA dataset in the United States contains 47 categories, but the number of annotations between categories is seriously unbalanced. Therefore, in order to ensure data availability, this experiment will select four categories with a large number of annotations for training and testing. The test set, validation set and training set are divided in a ratio of 6:2:2, and each image is flipped, cropped, denoised, and rotated.

[0072] Step 2: Use the standard YOLOv4 network to train and detect traffic signs; Use the standard YOLOv4 network to train the two traffic sign data sets based on step 1 respectively, download the standard YOLOv4 network and compile it. The standard YOLOv4 network download address is: https: / / github.com / AlexeyAB / darknet) For the two traffic sign data sets tt100k and LISA, change the training set, validation set, and test set directories in the tt100k.data and LISA.data files in the cfg folder to the address of the downloaded data set, and specify the number of categories and category names; Set epoch = 20000 according to the accuracy requirements, load tt100k.data or LSA.data according to this experimental data set, and load yolov4.cfg at the same time, and the program can start training; Use the loss function GIoU_Loss of the standard YOLOv4 network during training; Save the weight file Q of each layer during training 1 , as the weight input file for detection after training; using the weight file Q 1 Tests were conducted to obtain mAP and Recall. The obtained mAP, Recall and loss during training were analyzed. It was found that when the CIoU_Loss loss function was processing small targets, the regression loss was also large when the target aspect ratio changed greatly, which was not conducive to the stability of model training.

[0073] 1) Build a YOLOv4 network model and use the Initialization function to initialize the weight parameters of each layer of the neural network;

[0074] YOLOv4 is composed of four connected parts: (1) Input: refers to the original sample data input into the network; (2) BackBone: refers to the convolutional neural network structure that performs feature extraction operations; (3) Neck: fuses the image features extracted by the backbone network and passes the fused features to the prediction layer; (4) Head: predicts the target object of interest in the image and generates a visual prediction box and target category;

[0075] After downloading the standard YOLOv4 network, use the make command to compile the YOLOv4 network to form an executable file darknet; edit the tt100k.data and LISA.data files in the cfg folder for the two traffic sign datasets tt100k and LISA respectively, and change the class, train, valid, and names strings to the enhanced directories and parameters of the corresponding datasets. In this way, the parameters required for the Input part of the standard YOLOv4 network are edited. After setting the epoch, load tt100k.data or LISA.data according to the experimental dataset, and load yolov4.cfg at the same time, and the program can start training; the program will use the Initialization function to initialize the weight parameters of each layer of the neural network when it is running;

[0076] 2) Input image data from the Input part, pass through the Backbone part, and finally output feature maps of three scales, and use the classifier to output the prediction box Pb 2 and classification probability CP x ;

[0077] The image data is input from the Input part, passes through the Backbone part, and finally outputs feature maps of three scales. The feature maps of three different scales are sent to the Neck part composed of SPP and PANet, and the fused features are passed to the prediction layer. At the same time, the Head part completes the classification of the target and outputs the prediction box Pb 1 and classification probability CP x , where x is the index of each category; the prediction box Pb generated at this time 1 The number is too large, and there are a large number of detection frames for the same object in the picture, resulting in redundant detection results. After using the NMS algorithm to filter the prediction frames, the prediction frame Pb of the target of interest can be obtained. 2 The corresponding classification probability CP x ;

[0078] Reference Figure 2 :After the data enters the BackBone network from the Input, information extraction continues. The backbone network in the YOLOv4 network structure has 53 convolutional layers and outputs feature maps of three different scales. The feature maps of three different scales are sent to the Neck part composed of SPP and PANet to fuse and extract the feature maps, and the fused features are passed to the prediction layer. YOLOv4 is a one-stage target detection algorithm, so the Head part will simultaneously complete the prediction box and its corresponding classification probability;

[0079] YOLOv4 is a one-stage target detection algorithm. The Head part will simultaneously generate prediction boxes and their corresponding classification probabilities. However, the number of prediction boxes generated at this time is too large. There are a large number of detection boxes for the same object in the image, resulting in redundant detection results. It is necessary to perform non-maximum suppression on the redundant detection boxes so that each object in the image retains one prediction box. After the prediction boxes are screened by the NMS algorithm, the prediction boxes of the target of interest and their corresponding classification probabilities can be obtained.

[0080] 3) Compare the predicted box with the real box and calculate the IoU loss L CIoU , confidence loss L C , category loss L P , and add the three together to calculate the loss error Loss all =L CIoU +L C +L P And use the Adam algorithm to update the weights of each layer of the neural network;

[0081] The prediction box Pb obtained in 2) 2 Compared with the real box Gtb in the data set, the loss error is obtained. The loss error includes three parts, namely, IoU loss L CIoU , confidence loss L C , category loss L P , we will focus on the IoU loss L IoU ;

[0082] IoU loss L IoU After development, the more practical IoU loss can be summarized into the following four stages: IoU_Loss, GIoU_Loss, DIoU_Loss, CIoU_Loss;

[0083] Reference Figure 3 : IoU_Loss uses the intersection over union (IoU) as the loss function. The disadvantage is that when the two boxes do not intersect or have different overlapping areas, they cannot be distinguished, resulting in gradient return errors;

[0084] Reference Figure 4 : GIoU_Loss improves IoU_Loss. The disadvantage is that when the real box completely contains the predicted box, the IoU value between the two boxes is the same as the GIoU value, and the relative relationship cannot be distinguished;

[0085] Reference Figure 5 : DIoU_Loss uses the Euclidean distance between two center points to improve the shortcomings of GIoU_Loss;

[0086] CIoU_Loss takes into account the aspect ratio of the predicted box and the true box, focusing on the regression problem of the predicted box under different aspect ratios when the true box completely wraps the predicted box, and proposes the concept of CIoU, which is expressed as:

[0087]

[0088] Where α is the weight function, and v is an indicator used to measure the similarity of aspect ratios;

[0089] The gradient of CIoU_Loss is similar to DIoU_Loss, but the gradient of v is also considered, and its expression is:

[0090]

[0091] When the length and width are in [0,1], w 2 +h 2 The value of is usually very small, which will cause gradient explosion, so in implementation Replace with 1;

[0092] YOLOv4 uses CIoU_Loss in the loss function part, and the comprehensive expression of its loss function is:

[0093] Loss all =L CIoU +L C +L P

[0094] Where L CIoU The expression is:

[0095]

[0096] L C The expression is:

[0097]

[0098] where λ cls is the normalized weight factor of confidence, λ c Represents the weight coefficient of confidence loss to total loss; is the category gain factor;

[0099] L P The expression is:

[0100]

[0101] in Represents the weight coefficient of classification loss to total loss;

[0102] s 2 Indicates that the input image is divided into s×s cells, and B represents the total number of Anchor boxes contained in each cell. Indicates whether the jth Anchor box of the i-th cell is responsible for this object, that is, the cell contains the center point of the target object, and the jth Anchor box of the cell has the largest IoU value with the real box. Otherwise, 0; Indicates that the jth anchor box of the i-th cell is not responsible for the target; L SCIoU represents the SCIoU loss, C i represents the confidence of the i-th cell bounding box predicted by the network, and When the cell contains the center point of the target object, Pr(object) = 1, otherwise it is 0; represents the confidence of the bounding box of the i-th cell where the target is located; P i represents the probability that the i-th cell bounding box predicted by the network belongs to each category, Indicates the probability that the bounding box of the i-th cell where the target is located belongs to each category;

[0103] 4) Loop through steps 2) and 3) and continue iterating until the epoch value is reached, stop training, calculate mAP and Recall, and output a file Q that records the weights and offsets of each layer. 1 ;

[0104] The present invention sets the iteration threshold epoch=20000 according to the accuracy requirement. When the number of iterations is less than the threshold, the Adam algorithm is used to update the weights of each layer of the network until the threshold epoch=20000 stops training, calculates mAP and Recall, and outputs a file Q recording the weights and offsets of each layer. 1 ;

[0105] The most basic network performance evaluation indicators are divided into four categories, namely TP (True Positives): positive samples are correctly identified as positive samples, that is, dogs are correctly identified as dogs; TN (True Negatives): negative samples are correctly identified as negative samples, that is, cats are correctly identified as cats; FP (False Positives): negative samples are incorrectly identified as positive samples, that is, cats are incorrectly identified as dogs; FN (False Negatives): positive samples are incorrectly identified as negative samples, that is, dogs are incorrectly identified as cats; Accuracy represents the ratio of the number of correctly predicted samples to the total number of samples, which is used to evaluate the overall accuracy of the algorithm model. The calculation method is Precision is the ratio of the number of correctly identified samples to the total number of identified samples. The calculation method is: The recall rate is the ratio of samples correctly identified as positive examples to all positive examples. The calculation method is: If an algorithm model has good performance, it should meet the following conditions: while ensuring a high accuracy rate, the recall rate should also be maintained at a high level; in order to more vividly represent this condition, the Precision-Recall (PR) curve is used to show the trade-off between the accuracy rate and the recall rate of the algorithm model; AP refers to the area enclosed by the PR curve graph drawn by the accuracy rate and the recall rate obtained at a certain threshold and the horizontal and vertical axes, which measures the performance of the model in each category. mAP refers to the average AP of multiple target categories, which is used to measure the detection performance of the algorithm model on all tested categories. If there are N categories, the calculation method of mAP is The present invention mainly uses the model overall evaluation index mAP and Recall as the main evaluation indicators;

[0106] Step 3. In order to solve the problem that when the CIoU_Loss loss function processes small targets, the regression loss is also large when the target aspect ratio changes greatly. A loss function SCIoU (Sigmoid Complete Intersection over Union) based on CIoU_Loss is proposed. By using the improved aspect ratio measurement index ν s The arctan function is replaced by the Sigmoid function prototype and embedded into YOLOv4, which improves the performance. At the same time, the model does not introduce more computational complexity, and the real-time performance is not affected. It can also replace the loss function of other target detection algorithms and has good applicability. The YOLOv4 network with the replaced loss function SCIoU is trained using the two data sets in step 1 to obtain the weight file Q 2; Using the weight file Q 2 Test and get mAP and Recall;

[0107] Based on CIoU, this paper proposes an improved aspect ratio measurement index v s , whose expression is:

[0108]

[0109] Reference Figure 7 :The arctan function is replaced by the Sigmoid function prototype, and the aspect ratio w / h is used as the independent variable of the function. It can be seen that the range of the arctan function is [0,π / 2), and the range of the Sigmoid function is [0.5,1). When the aspect ratio w / h is small, the change of Sigmoid is smaller than that of arctan as the w / h value changes. Corresponding to the expression of the aspect ratio metric v in the loss function, the aspect ratio of the predicted box and the true box is normalized after mapping. Sigmoid makes the regression loss change smaller. That is, this paper proposes the SCIoU loss function, which is conducive to the stability of model training.

[0110] The expression of SCIoU as a loss function is:

[0111]

[0112] Likewise s The gradient expression is:

[0113]

[0114] Similarly, in the experimental case, in order to prevent the gradient from exploding, The value of is 1;

[0115] The total loss in step 2 is Loss all =L CIoU +L C +L P , L CIoU Replaced by L in the present invention SCIoI , and embed it into YOLOv4, and train it in the same way as in step 2, iterate to the epoch value and use the Adam algorithm to update the weights, calculate mAP and Recall, and save the weight file Q 2 ;

[0116] Step 4: Compare the mAP and Recall obtained in step 3 with those obtained in step 2, set different IoU_thresh settings for retraining, compare the mAP and Recall performance of step 2 and step 3 under different settings, and analyze the test results;

[0117] In the above steps, the value range of i is 0 to the number of divided cells s×s, and the value range of j is 0 to the number of divided anchor boxes B;

[0118] The present invention proposes a loss function SCIoU (Sigmoid Complete Intersection over Union) based on CIoU_Loss, and replaces the CIoU_Loss used in the standard YOLOv4 to form a new neural network algorithm; compared with the standard YOLOv4, the improved YOLOv4 of the present invention can improve the algorithm performance without affecting the real-time performance of the algorithm, so that the mAP and Recall of the algorithm can be improved under different data sets, and in practical applications, the recognition rate can be improved under the premise of the same detection time.

[0119] The invention is further described below in conjunction with a simulation example.

[0120] Simulation example:

[0121] The present invention uses the original YOLOv4 as a comparison, and the training data set and the test data set are both from the general data sets tt100k and LISA to verify the universality of the algorithm to different data sets.

[0122] Figure 8Figure 2 is the detection effect diagram of some test images in the dataset using the original YOLOv4 model and the improved YOLOv4 model, where (a) and (b) are detection images with a smaller tilt angle, (c) and (d) are detection images with a slightly larger tilt angle, (e) and (f) are detection images with a larger tilt angle, and (a), (c) and (e) are detection results of the original YOLOv4 model, (b), (d) and (f) are detection results of the improved YOLOv4 model. The detection results show that when the tilt angle is small, both the original YOLOv4 model and the improved YOLOv4 model correctly identify the target object, as shown in Figures (a) and (b), the detected target and the confidence are pne:100% and i5:99% respectively; when the tilt angle increases, the information displayed by the target decreases, as shown in Figures (c) and (d), where the detected target and the confidence are w57:60% in Figure (c) and w57:100% in Figure (d). It can be seen that compared with the detection result of the improved YOLOv4 model in Figure (d), the original YOLOv4 model in Figure (c) correctly identifies the target object. The traffic sign target can be identified, but the target confidence is low, and the position of the prediction box is not as accurate as in Figure (d), with some differences. When the tilt angle is larger, as shown in Figures (e) and (f), the detection result of the original YOLOv4 model in Figure (e) is not ideal, and the category pn cannot be correctly identified. However, the detection result of the improved YOLOv4 model in Figure (f) shows that the detected target and the confidence are pn:99% respectively. It can be seen that the improved YOLOv4 model can still maintain good detection performance when the tilt angle is large, that is, the degree of overcoming the influencing factor of the tilt angle is higher than that of the original YOLOv4 model, and the performance of the algorithm model is better.

[0123] Fig. 9 It is the overall performance of the original YOLOv4 model and the improved YOLOv4 model after using the improved loss function SCIoU of the present invention on the tt100k verification data set. The mAP value and the Recall value are slightly improved. When the IoU threshold is set to 0.5, the mAP value reaches 88.23%, and the Recall value reaches 86.82%, which are 0.12% and 0.31% higher than the mAP value and Recall value of the original YOLOv4 model, respectively.

[0124] Reference Fig.10The loss function result graph is the image obtained by discarding the first 500 Batch training samples and taking a loss value for every 200 Batch afterwards, where Batch represents the number of samples sent to the network model for training each time. It can be seen that although the mAP value of the improved YOLOv4 model is not significantly improved compared to the original YOLOv4 model, the loss value is maintained at a relatively low level in the later stage of training, the fluctuation range of the loss during training is reduced, the model training is more stable, and it is conducive to the rapid convergence of the model.

[0125] When the IoU threshold is set to 0.75, the mAP value reaches 80.07%, and the Recall value reaches 80.42%, which are 0.9% and 0.41% higher than the mAP value and Recall value of the original YOLOv4 model, respectively. This shows that based on the SCIoU loss function, the detection performance is more significantly improved when the IoU threshold is set larger. This is because when the predicted box is completely wrapped by the real box, for predicted boxes with different aspect ratios, the SCIoU loss function enables the training process to complete the gradient backpropagation without degenerating into IoU loss, thereby making the model more accurate in target positioning during training and achieving better detection performance.

[0126] Fig.11 The overall performance of the original YOLOv4 model and the improved YOLOv4 model after using the improved loss function SCIoU of the present invention on the LISA validation data set, the experimental results show that the mAP value of the original YOLOv4 model is 99.15%, and the mAP value of the improved YOLOv4 model using the SCIoU loss function reaches 99.38%, which is 0.23% higher than the original YOLOv4 model; while analyzing the performance when the IoU threshold is 0.75, the improved YOLOv4 model is 0.53% higher than the original YOLOv4 model. As analyzed above, the training model in the improved YOLOv4 model has a higher accuracy in target positioning, thereby achieving a better detection effect.

[0127] In summary, the simulation results show that compared with the original YOLOv4 model, the YOLOv4 model after improving the loss function SCIoU in the present invention takes into account the situation where the different aspect ratios of the predicted box completely wrap the predicted box and the impact of the different aspect ratios of the predicted box on the regression result. On this basis, the algorithm of the present invention proposes a new aspect ratio measurement index v s , improve the stability of model training, and the detection effect is better than the original YOLOv4 model using CIoU. At the same time, the YOLOv4 model after this improved loss function SCIoU has universality and has improved detection performance on both tt100k and LISA datasets.

Claims

1. A YOLOv4 target detection method with improved loss function, It is characterized in that The following steps are involved: Step 1: Download the tt100k dataset and LISA dataset, which are common datasets in the field of current target detection. Using these two datasets can ensure that the detection effect of the algorithm is consistent with the common datasets disclosed in the field and verify the actual effect of the algorithm. Enhance the downloaded data, including flipping, cropping, adding noise, and rotating operations. The data generated after enhancement can not only increase the number of images included in the dataset, but also because the enhanced images are more complex than the original images in the dataset, the style and size of the images are changed while retaining the feature points of the original images, and the blur of the images is increased, making the enhanced images more diverse and closer to the actual situation, which can improve the robustness of the trained network. Step 2: Use the standard YOLOv4 network to train and detect traffic signs; Use the standard YOLOv4 network to train the two traffic sign data sets based on step 1 respectively, download the standard YOLOv4 network and compile it, and change the training set, validation set, and test set directories in the tt100k.data and LISA.data files in the cfg folder for the two traffic sign data sets tt100k and LISA to the addresses of the downloaded data sets, and specify the number of categories and category names; Set epoch = 20000 according to the accuracy requirements, load tt100k.data or LISA.data according to this experimental data set, and load yolov4.cfg at the same time, and the program can start training; Use the loss function CIoU_Loss of the standard YOLOv4 network during training; Save the weight file Q of each layer during training 1 , as the weight input file for detection after training; Using the weight file Q 1 Tests were performed to obtain mAP and Recall. The obtained mAP, Recall and loss during training were analyzed. It was found that when the CIoU_Loss loss function processed small targets, the regression loss also increased when the target aspect ratio changed more, which was not conducive to the stability of model training. Step 3: To solve the problem that when the CIoU_Loss loss function processes small targets, the regression loss also increases when the target aspect ratio changes. A loss function Sigmoid Complete Intersectionover Union (SCIoU) based on CIoU_Loss is proposed. The improved aspect ratio metric ν s The arctan function is replaced by the Sigmoid function prototype and embedded into YOLOv4, which improves the performance. At the same time, the model does not introduce more computational complexity, and the real-time performance is not affected. It can also replace the loss function of other target detection algorithms and has good applicability. The YOLOv4 network with the replaced loss function SCIoU is trained using the two data sets in step 1 to obtain the weight file Q 2 ; Using the weight file Q 2 Test and get mAP and Recall; Based on CIoU, an improved aspect ratio metric ν is proposed s , whose expression is: The arctan function is replaced by the Sigmoid function prototype, and the aspect ratio w / h is used as the independent variable of the function. It can be seen that the range of the arctan function is [0,π / 2), and the range of the Sigmoid function is [0.5,1). When the aspect ratio w / h becomes smaller, the change of Sigmoid is smaller than that of arctan as the w / h value changes. Corresponding to the expression of the aspect ratio metric v in the loss function, the aspect ratio of the predicted box and the real box is mapped and normalized. Sigmoid makes the regression loss change smaller. That is, this paper proposes the SCIoU loss function, which is conducive to the stability of model training. The expression of SCIoU as a loss function is: Similarly v s The gradient expression is as follows: Similarly, in the experimental case, in order to prevent the gradient from exploding, The value of is 1; The total loss in step 2 is Loss all =L CIoU +L C +L P , L CIoU Replace with L in SCIoU , and embed it into YOLOv4, and train it in the same way as in step 2, iterate to the epoch value and use the Adam algorithm to update the weights, calculate mAP and Recall, and save the weight file Q 2 ; Step 4: Compare the mAP and Recall obtained in step 3 with those obtained in step 2, set different IoU_thresh settings for retraining, compare the mAP and Recall performance of step 2 and step 3 under different settings, and analyze the test results.

2. According to the YOLOv4 target detection method with improved loss function described in claim 1, step 1, download the current target detection field general data set tt100k data set and LISA data set; the full name of tt100k is Tsinghua-Tencent100K, which is a general data set of road traffic signs that can be used to identify provided by Tsinghua-Tencent Internet Innovation Technology Joint Laboratory; the resolution of the image in the TT100K data set is 2048×2048, and there are 221 sign categories, which are divided into three categories: warning signs, prohibition signs and instruction signs; the data set covers different weather conditions. The training set contains 6105 images and the test set contains 3071 images. Due to the large resolution of the original images, the original images were cropped in this experiment, and the scale of the cropped images was 608×608. Due to the serious imbalance in the amount of data between categories in the data set, this experiment only selected 45 types of traffic signs with a large amount of labeled data for recognition, and divided the test set, validation set and training set in a ratio of 6:2:2, and flipped, cropped, denoised and rotated each image. The full name of LISA is Laboratory for Intelligent&SafeAutomobiles is a general dataset for identifying road traffic signs provided by the LISA laboratory in the United States. It extracts a certain segment with traffic signs from the video by driving a vehicle, and then extracts up to 30 frames based on this segment, and annotates each frame of the video. The annotation of each traffic sign contains four parts of information: Tag, Position, Occluded, and On side rode. The process of collecting pictures is extracted from the video. The vehicle has a certain speed during driving, not static, so it is blurred, which also makes the traffic sign recognition algorithm based on this dataset more applicable to real scenes. The LISA dataset in the United States contains 47 categories, but the number of annotations between categories is seriously unbalanced. Therefore, in order to ensure data availability, this experiment will select four categories with a large number of annotations for training and testing. The test set, validation set and training set are divided in a ratio of 6:2:2, and each image is flipped, cropped, denoised, and rotated.

3. According to the YOLOv4 target detection method with improved loss function described in claim 1, step 2, use standard YOLOv4 network to train and detect traffic signs; use standard YOLOv4 network to train two traffic sign data sets based on step 1 respectively, download standard YOLOv4 network and compile, change the training set, verification set and test set directories in tt100k.data and LISA.data files in cfg folder to the address of downloading data set respectively for two traffic sign data sets tt100k and LISA, and specify the number of categories and category names; set epoch=20000 according to accuracy requirements, load tt100k.data or LISA.data according to this experimental data set, and load yolov4.cfg at the same time, and the program can start training; use the loss function CIoU_Loss of standard YOLOv4 network during training; save the weight file Q of each layer during training 1 , as the weight input file for detection after training ; Using the weight file Q 1 Tests were performed to obtain mAP and Recall. The obtained mAP, Recall and loss during training were analyzed. It was found that when the CIoU_Loss loss function processed small targets, the regression loss also increased when the target aspect ratio changed more, which was not conducive to the stability of model training. 1) Build a YOLOv4 network model and use the Initialization function to initialize the weight parameters of each layer of the neural network; YOLOv4 is composed of four connected parts: (1) Input: refers to the original sample data input into the network; (2) BackBone: refers to the convolutional neural network structure that performs feature extraction operations; (3) Neck: fuses the image features extracted by the backbone network and passes the fused features to the prediction layer; (4) Head: predicts the target object of interest in the image and generates a visual prediction box and target category; After downloading the standard YOLOv4 network, use the make command to compile the YOLOv4 network to form an executable file darknet; edit the tt100k.data and LISA.data files in the cfg folder for the two traffic sign datasets tt100k and LISA respectively, and change the class, train, valid, and names strings to the enhanced directories and parameters of the corresponding datasets. In this way, the parameters required for the Input part of the standard YOLOv4 network are edited. After setting the epoch, load tt100k.data or LISA.data according to the experimental dataset, and load yolov4.cfg at the same time, and the program can start training; the program will use the Initialization function to initialize the weight parameters of each layer of the neural network when it is running; 2) Input image data from the Input part, pass through the Backbone part, and finally output feature maps of three scales, and use the classifier to output the prediction box Pb 2 and classification probability CP x ; The image data is input from the Input part, passes through the Backbone part, and finally outputs feature maps of three scales. The feature maps of three different scales are sent to the Neck part composed of SPP and PANet, and the fused features are passed to the prediction layer. At the same time, the Head part completes the classification of the target and outputs the prediction box Pb 1 and classification probability CP x , where x is the index of each category; the prediction box Pb generated at this time 1 The number is too large, and there are a large number of detection frames for the same object in the picture, resulting in redundant detection results. After using the NMS algorithm to filter the prediction frames, the prediction frame Pb of the target of interest can be obtained. 2 The corresponding classification probability CP x ; After the data enters the BackBone network from the Input, information extraction continues. The backbone network in the YOLOv4 network structure has 53 convolutional layers and outputs feature maps of three different scales. The feature maps of three different scales are sent to the Neck part composed of SPP and PANet to fuse and extract features, and the fused features are passed to the prediction layer. YOLOv4 is a one-stage target detection algorithm, so the Head part will simultaneously complete the prediction box and its corresponding classification probability; YOLOv4 is a one-stage target detection algorithm. The Head part will simultaneously generate prediction boxes and their corresponding classification probabilities. However, the number of prediction boxes generated at this time is too large. There are a large number of detection boxes for the same object in the image, resulting in redundant detection results. It is necessary to perform non-maximum suppression on the redundant detection boxes so that each object in the image retains one prediction box. After the prediction boxes are screened by the NMS algorithm, the prediction boxes of the target of interest and their corresponding classification probabilities can be obtained. 3) Compare the predicted box with the real box and calculate the IoU loss L CIoU , confidence loss L C , category loss L P , and add the three together to calculate the loss error Loss all =L CIoU +L C +L P And use the Adam algorithm to update the weights of each layer of the neural network; The prediction box Pb obtained in 2) 2 Compared with the real box Gtb in the data set, the loss error is obtained. The loss error includes three parts, namely, IoU loss L CIoU , confidence loss L C , category loss L P , we will focus on the IoU loss L IoU ; IoU loss L IoU After development, the more practical IoU loss can be summarized into the following four stages: IoU_Loss, GIoU_Loss, DIoU_Loss, CIoU_Loss; IoU_Loss uses the intersection-over-union ratio IoU as the loss function. The disadvantage is that when the two boxes do not intersect or the two boxes have overlapping areas of different regions, they cannot be distinguished, resulting in gradient return errors; GIoU_Loss improves IoU_Loss. The disadvantage is that when the real box completely contains the predicted box, the IoU value between the two boxes is the same as the GIoU value, and the relative relationship cannot be distinguished; DIoU_Loss uses the Euclidean distance between the two center points to improve the shortcomings of GIoU_Loss; CIoU_Loss takes into account the aspect ratio of the predicted box and the real box, focusing on the regression problem of the predicted box under different aspect ratios when the real box completely wraps the predicted box, and proposes the concept of CIoU, which is expressed as: Where α is the weight function, and v is an indicator used to measure the similarity of aspect ratios; The gradient of CIoU_Loss is similar to DIoU_Loss, but the gradient of v is also considered, and its expression is: When the length and width are in [0,1], This will cause gradient explosion, so when implementing Replace with 1; The standard YOLOv4 uses CIoU_Loss in the loss function part, and the comprehensive expression of its loss function is: Loss all =L CIoU +L C +L P Where L CIoU The expression is: L C The expression is: where λ cls is the normalized weight factor of confidence, λ c Represents the weight coefficient of confidence loss to total loss; L P The expression for: in Represents the weight coefficient of classification loss to total loss; s 2 Indicates that the input image is divided into s×s cells, and B represents the total number of Anchor boxes contained in each cell. Indicates whether the jth Anchor box of the i-th cell is responsible for this object, that is, the cell contains the center point of the target object, and the jth Anchor box of the cell has the largest IoU value with the real box. Otherwise, 0; Indicates that the jth anchor box of the i-th cell is not responsible for the target; L SCIoU represents the SCIoU loss, C i represents the confidence of the i-th bounding box predicted by the network, and When the cell contains the center point of the target object, Pr(object) = 1, otherwise it is 0; represents the confidence of the i-th bounding box where the target is located; P i represents the probability that the i-th bounding box predicted by the network belongs to each category, Indicates the probability that the i-th bounding box where the target is located belongs to each category; 4) Loop through steps 2) and 3) and continue iterating until the epoch value is reached, stop training, calculate mAP and Recall, and output a file Q that records the weights and offsets of each layer. 1 ; According to the accuracy requirements, the iteration threshold epoch = 20000 is set. When the number of iterations is less than the threshold, the Adam algorithm is used to update the weights of each layer of the network until the threshold epoch = 20000 stops training, calculates mAP and Recall, and outputs a file Q recording the weights and offsets of each layer 1 ; The model's overall evaluation indicators mAP and Recall are used as evaluation indicators.