An object detection method based on an improved YOLOv4 network
By introducing the ECANet channel attention module and the new border regression loss function LCIOU in the YOLOv4 network, combined with the fusion confidence loss function, the problem of difficulty in taking into account the accuracy and speed of the existing object detection algorithm in complex scenarios is solved, and the target detection effect of high precision and high speed is achieved.
Patent Information
- Application Number
- CN202211487874.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-11-25
AI Technical Summary
The existing object detection algorithm is difficult to achieve high accuracy and high speed at the same time under the interference of complex application scenarios and occlusion factors, resulting in difficulty in applying it in a real-time detection environment.
Based on the improved object detection method of YOLOv4 network, the convergence speed and generalization performance of the model are enhanced by introducing the ECANet channel attention module in the backbone network, and the new border regression loss function LCIOU and the fusion confidence loss function are adopted.
It improves the accuracy and detection speed of target detection, is better than other advanced algorithms, is suitable for real-time detection environments, and at the same time reduces the number of parameters of the model and maintains the detection speed.
Smart Images

Figure CN115731392B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and digital processing, and particularly relates to an object detection method based on an improved YOLOv4 network. Background Art
[0002] Object detection is a popular direction in computer vision and digital image processing, and is widely applied in fields such as robot navigation, intelligent video surveillance, industrial inspection, and aerospace. Due to the different appearances, shapes, and postures of the objects to be detected, as well as the interference of factors such as illumination and occlusion during imaging, object detection has always been one of the most challenging problems in the field of machine vision. In early object detection algorithms, Felzenszwalb et al. proposed the deformable part model (DMP) algorithm. This algorithm extracts image features from the image, creates a corresponding excitation template for a certain object, and calculates in the original image to obtain an excitation effect diagram to determine the target position. However, due to the problems of unsatisfactory performance and large manual workload, the application of such algorithms is limited.
[0003] In recent years, with the vigorous development of deep learning, object detection algorithms based on deep learning have developed rapidly, and have been greatly improved in terms of accuracy and speed, gradually replacing traditional methods and being applied in various fields. Due to complex application scenarios and the interference of factors such as occlusion, object detection algorithms based on deep learning have higher requirements for accuracy. Many scholars have continuously deepened the convolutional layer to extract more object information and increased the network structure to improve accuracy. However, a large amount of computational complexity will reduce the detection speed, making it difficult to be applied to scenarios with high real-time requirements, and also having higher requirements for hardware devices, increasing the application cost.
[0004] It is difficult to simultaneously obtain satisfactory results for the speed and accuracy of object detection algorithms. Two-stage object detection algorithms have a deeper network structure and a large amount of computation, resulting in higher detection accuracy, but also bringing problems such as a large model and slow speed, resulting in poor real-time performance and being difficult to be applied to real-time detection environments. Single-stage detection algorithms have a more streamlined network, faster detection speed and high real-time performance, but lower accuracy. Summary of the Invention
[0005] Due to the influence of complex application scenarios, various factors interfering during imaging, different postures and sizes of targets, etc., many expected targets cannot be detected. Some scholars have continuously deepened the structural depth of the algorithm, resulting in a bloated network structure and reduced detection speed, which is not conducive to being applied to various fields in the real world. There are also researchers who continuously add information of the ground truth box and the predicted box to the regression loss function during the algorithm analysis stage to improve accuracy. However, the current function not only lacks important bounding box information but also requires a longer training time for the algorithm model. To solve the above problems, the present invention provides a target detection method based on an improved YOLOv4 network, including the following steps:
[0006] S1. Construct an improved YOLOv4 network based on the existing YOLOv4 network. The improved YOLOv4 network includes an improved backbone network, a neck network, and a head network; an ECANet channel attention module is added to the improved backbone network;
[0007] S2. Obtain an image dataset to train the improved YOLOv4 network, and perform iterative training using a new bounding box regression loss function, a fused confidence loss function, and a classification loss function to adjust the network parameters;
[0008] S3. Use the trained improved YOLOv4 network for target detection to obtain the target detection result.
[0009] Further, based on the backbone network of the existing YOLOv4 network, an improved backbone network is formed by adding an ECANet channel attention module after each of the CBM layer, the CSP1 layer, and the CSP2 layer.
[0010] Further, the diagonal distance of the intersection rectangle of the ground truth box and the predicted box and the diagonal distance of the minimum bounding rectangle are used as additional bounding box information, and an additional information penalty term is constructed based on the additional bounding box information. The additional information penalty term is introduced into the CIOU function to obtain a new regression loss function, where the CIOU function is expressed as:
[0011]
[0012]
[0013]
[0014] Among them, IOU represents the intersection over union of the bounding boxes of the ground truth box and the predicted box, w gt and h gt respectively represent the width and height of the ground truth box, w and h respectively represent the width and height of the predicted box, b represents the predicted box, b gt represents the ground truth box, p 2 (b, b gt) represents the Euclidean distance between the center points of the ground truth box and the predicted box, c represents the diagonal distance of the minimum bounding rectangle of the ground truth box and the predicted box, α is a parameter function used to balance the aspect ratio, and v is a function used to evaluate the proportional consistency between the ground truth box and the predicted box.
[0015] Furthermore, normalize the additional bounding box information, and construct the basic form of the additional information penalty term based on the normalized additional bounding box information. To ensure that the gradient direction of the additional information penalty term is the same as that of the CIOU function, the constructed basic form is expressed as:
[0016]
[0017] Then, combine the norm advantage of the Huber function to construct the final additional information penalty term, which is expressed as:
[0018]
[0019] Among them, r represents the diagonal distance of the intersection rectangle of the ground truth box and the predicted box, c represents the diagonal distance of the minimum bounding rectangle of the ground truth box and the predicted box, λ represents the proportion of the additional information penalty term in the new bounding box regression function, and δ represents the error when the model is stable.
[0020] Furthermore, introduce the additional information penalty term into the CIOU function to obtain the new bounding box regression loss function, which is expressed as:
[0021]
[0022] Among them, R δ (diff) represents the additional information penalty term.
[0023] Furthermore, the fused confidence loss function is expressed as:
[0024]
[0025] Among them, S 2 is the number of grids divided for each image, and B is the prior box of each grid generated by the network; and indicate whether there is an object in the j-th prior box of the i-th grid of the image. If there is an object in the prior box, at this time On the contrary, if there is no object in the prior box, then represents the weight of the loss calculation without an object, indicates whether there is an object in the i-th grid. If there is an object, then Otherwise, it is 0; represents the probability of predicting the existence of an object, and the value range is [0,1].
[0026] Advantages of the present invention:
[0027] The present invention proposes an object detection method based on an improved YOLOv4 network, with improvements mainly in three aspects. First, a new bounding box regression loss function, L-Norm Complete Intersection over Union (LCIOU), is proposed to enhance the convergence speed and generalization performance of the model. Second, a fused confidence loss function is proposed, which uses both binary cross-entropy function and mean squared error function, and this function can more accurately calculate the confidence loss when there is or is not an object in the image. Third, a channel attention mechanism, Efficient Channel Attention (ECANet), is introduced into the backbone feature extraction network. Different from previous introduction methods, it enhances the expression of effective information, suppresses the extraction of invalid information, and does not increase additional parameters. The results show that the algorithm proposed in this patent is superior to other advanced algorithms in terms of accuracy and detection speed. Description of the Drawings
[0028] Figure 1 It is the flowchart of the method of the present invention;
[0029] Figure 2 It is the schematic diagram of the CSP structure of the present invention;
[0030] Figure 3 It is the schematic diagram of the improved YOLOv4 network structure of the present invention;
[0031] Figure 4 It is the schematic diagram of the improved backbone network structure of the present invention
[0032] Figure 5 It is the schematic diagram of the ECANet structure of the present invention;
[0033] Figure 6 It is the schematic diagram of the additional bounding box information of the present invention;
[0034] Figure 7 It is the function image of the additional information penalty term of the present invention:
[0035] Figure 8 It is the reciprocal function graph of the additional information penalty term of the present invention;
[0036] Figure 9 It is the comparison diagram of the detection effects of the present invention. Detailed Embodiments
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] The present invention provides an object detection method based on an improved YOLOv4 network, as Figure 1 shown, including the following steps:
[0039] S1. Construct an improved YOLOv4 network based on the existing YOLOv4 network. The improved YOLOv4 network includes an improved backbone network, a neck network, and a head network; an ECANet channel attention module is added to the improved backbone network;
[0040] S2. Obtain an image dataset to train the improved YOLOv4 network, and perform iterative training using a new bounding box regression loss function and a fused confidence loss function to adjust the network parameters;
[0041] S3. Use the trained improved YOLOv4 network for object detection to obtain object detection results.
[0042] In the field of object detection, researchers have proposed various object detection model architectures, and various studies have also been conducted on the backbone network structure of object detection models. Among them, most object detection models often adopt the YOLOv4 network structure. The backbone feature extraction network of the YOLOv4 network uses the CSPDarkNet53 network, and the CSPDarkNet53 network introduces the CSP structure on the basis of the DarkNet53 network. The CSP structure enhances the learning ability of the convolutional neural network, removes the computational bottleneck, significantly reduces the use of video memory, and at the same time speeds up the inference speed of the network.
[0043] The CSP structure divides each layer of feature map into two parts, and its form is as Figure 2 shown. In the Base layer, the feature map is divided into part1 and part2. Part2 is stacked with multiple residual blocks, and the stacked result is fused with part1 and then output to the next layer. The form of dividing the feature into two parts by the CSP structure also separates the gradient flow, allowing the gradient flow to propagate in different network paths, alleviating the problems of gradient disappearance and gradient explosion, and at the same time improving the model learning ability.
[0044] Figure 2The purpose of the Partial transition layers is to maximize the difference in gradient combination. By using the means of gradient flow truncation, it can avoid different layers from learning duplicate gradient information, effectively reducing duplicate gradient learning and greatly improving the learning ability of the network. The CSPDarknet53 network adds the CSP structure to each large residual block of the Darknet53 network, obtaining better performance. However, a large amount of invalid information such as background information is still contained in the three output feature layers of CSPDarknet53, which will interfere with the regression of prediction boxes, resulting in missed detections and false detections. To suppress the extraction of invalid information, an ECANet channel attention module is introduced into the backbone network to enhance the expression of valid information.
[0045] In one embodiment, the present invention improves on the network structure of the existing YOLOv4 network to obtain the structure of the improved YOLOv4 network, as Figure 3 shown, including an improved backbone (BackBorn) network, a neck (Neck) network, and a head (Head) network. The head network is also the predict network.
[0046] Specifically, the backbone network structure of the existing YOLOv4 network is a sequentially connected CBM layer, CSP1 layer, CSP2 layer, CSP8 layer, CSP8 layer, and CSP4 layer. To enhance the expression of target information and at the same time suppress the extraction of invalid information, in this embodiment, an ECANet channel attention module is added after each of the CBM layer, CSP1 layer, and CSP2 layer, thus forming an improved backbone network as Figure 4 shown.
[0047] The ECANet channel attention module is a lightweight attention mechanism that avoids dimensionality reduction. It is a local cross-channel interaction module without dimensionality reduction proposed on the basis of SENet. This module can bring obvious performance improvement by adding only a small number of parameters. Its network structure is as Figure 5 shown. First, a feature map with dimensions H×W×C is input. The global average pooling GAP is used to compress the spatial features of this feature map to obtain a feature map of 1×1×C. Channel feature learning is performed on the compressed feature map through a 1×1 convolution to learn the importance between different channels. At this time, the output dimension is still 1×1×C. Finally, the 1×1×C feature map after convolution is multiplied with the original input H×W×C feature map channel by channel to obtain a feature map with channel attention.
[0048] The ECANet network structure is simple. After integrating it into the backbone network, the improved backbone network can extract more feature information by adding only a small number of parameters, achieving the purpose of increasing the accuracy of the model.
[0049] In the object detection algorithm, the calculation of the training loss using the YOLOv4 network is generally divided into three parts, namely the regression loss, the confidence loss, and the classification loss.
[0050] In one embodiment, in order to help more prediction boxes accurately regress, improve the detection effect and convergence speed, the present invention adds a new penalty term to the original regression loss to obtain a new bounding box regression loss function. Specifically, the CIOU function is currently the bounding box regression loss function with the best effect and wide application. It fully considers important bounding box information such as the overlapping area, the distance between the center points of the ground truth box and the predicted box, and the aspect ratio, and alleviates problems such as the gradient disappearance when the two boxes do not overlap and divergence during the training process. The CIOU function is expressed as:
[0051]
[0052] Among them, IOU represents the intersection over union of the bounding boxes of the ground truth box and the predicted box, b represents the predicted box, and b gt represents the ground truth box, and p 2 (b, b gt ) represents the Euclidean distance between the center points of the ground truth box and the predicted box, c represents the diagonal distance of the smallest enclosing rectangle of the ground truth box and the predicted box, α is a parameter function for balancing the aspect ratio column, and v is a function for evaluating the proportional consistency between the ground truth box and the predicted box, where:
[0053]
[0054]
[0055] w gt and h gt respectively represent the width and height of the ground truth box, and w and h respectively represent the width and height of the predicted box. CIOU is an improvement based on IOU. CIOU improves the performance of the original function by adding a penalty term for bounding box-related information. To improve the regression accuracy of the model, in this embodiment, an additional information penalty term containing additional bounding box information is added to CIOU.
[0056] Specifically, the design of the additional information penalty term mainly considers two aspects. First, it is necessary to introduce appropriate bounding box information into the additional information penalty term, such as the center distance, aspect ratio, etc.; second, it is to select the appropriate form of the additional bounding box information in the additional information penalty term, and this form should have scale invariance. Based on the above considerations, the additional bounding box information selected in this embodiment is the diagonal distance of the intersection rectangle and the diagonal distance of the smallest enclosing rectangle between the ground truth box and the predicted box. As Figure 6As shown in the figure, where r is the diagonal distance of the intersection rectangle between the ground truth box and the predicted box, and c is the diagonal distance of the minimum bounding rectangle of the ground truth box and the predicted box. Since the distance information is sensitive to scale, it is necessary to normalize the distance information to make the subsequent additional information penalty term scale-invariant. At the same time, considering that the gradient direction is the same as that of CIOU, the basic form of the additional information penalty term is first determined as follows:
[0057]
[0058] After determining the basic form of the additional information penalty term, it is necessary to select a suitable form to improve the performance of CIOU. Since the basic form of the additional information penalty term already has the properties of scale invariance and the same direction gradient as CIOU, the final form of the additional information penalty term only needs to consider smoothness.
[0059] Furthermore, the derivative of the L1 norm loss function is a constant, that is, no matter what the input value is, the L1 norm loss function can provide a stable gradient. At the same time, its robust performance can handle outliers well. The derivative of the L2 norm loss function is a linear function, that is, the value of the derivative increases as the error increases. When the model error is small, the L2 loss can stably reduce the gradient value as the error decreases, and it is easier to stably converge to the vicinity of the extreme point, enabling the model to obtain higher accuracy. The Huber function not only inherits the advantages of the L1 and L2 norms but also abandons some of their disadvantages and has good smoothness. Therefore, in this embodiment, the final form of the additional information penalty term is determined based on the Huber function, which is expressed as:
[0060]
[0061] Where r represents the diagonal distance of the intersection rectangle between the ground truth box and the predicted box, c represents the diagonal distance of the minimum bounding rectangle of the ground truth box and the predicted box, λ represents the proportion of the additional information penalty term in the new bounding box regression function, and δ represents the error when the model is stable.
[0062] Based on the CIOU function and the additional information penalty term, a new bounding box regression loss function is obtained, which is expressed as:
[0063]
[0064] In this embodiment, the final form of the additional information penalty term consists of two parts. is the penalty term based on the L1 norm form. is the penalty term based on the L2 norm form. The function image of this additional information penalty term is as shown in Figure 7As shown, when the error is large in the early stage of training, a sufficiently large penalty is given. When the error is small, the penalty can be quickly reduced. At the same time, no matter how large the error generated by the model is in the early stage of training, the additional information penalty term can provide a stable gradient to continue training. In the later stage of training, the additional information penalty term can provide a changing gradient to make the model more likely to converge to the extreme point. By adjusting the value of λ, the additional information penalty term proposed in this embodiment exhibits the advantages of L1 and L2 to varying degrees in the regression function. The main contribution of the additional information penalty term is to introduce additional bounding box information into the regression function, helping more prediction boxes continuously approach the true box, improving the bounding box regression accuracy, and thus achieving the purpose of improving the model accuracy. Experiments have shown that multiple evaluation indicators of the model with loss functions such as CIOU and DIOU with this penalty term added have been improved.
[0065] Specifically, from Figure 8 it can be seen that the penalty term obtains the properties of L1 and L2 norms to different degrees through different values of λ. In view of problems such as gradient explosion that may occur in the early stage of model training, the value of λ should be reduced to suppress the effect of the L1-form penalty term.
[0066] In one embodiment, in order to more accurately calculate the confidence loss when there is or is not a target in the image, the present invention provides a fused confidence loss function that simultaneously uses the binary cross-entropy function and the mean squared error function.
[0067] Confidence represents the possibility of whether there is a target predicted by each grid in the image. The loss function includes two parts: the loss calculation when there is a target and the loss calculation when there is no target. The YOLOv4 confidence loss function adopts the form:
[0068]
[0069] where S 2 is the number of grids divided for each image, and B is the prior box of each grid generated by the network. and indicate whether there is a target in the j-th prior box of the i-th grid in the image. If there is a target in the prior box, at this time On the contrary, if there is no target in the prior box, then They determine whether the corresponding partial loss of the loss function participates in the calculation. λ noobj represents the weight of the loss calculation when there is no target. indicates whether there is a target in the i-th grid. If there is a target, then Otherwise, it is 0. represents the probability of predicting the existence of a target, and its value is in [0, 1].
[0070] In actual situations, only a small part of an image has targets. The YOLOv4 confidence loss function adds weights to the loss of areas without targets to reduce the contribution of this part of the loss value. Otherwise, it may cause the network to tend to predict that the cell does not contain an object, thus affecting the model's accuracy. Since most grids in an image do not contain objects, the calculation of the loss for this part is particularly important.
[0071] However, the BCE function has defects when used to calculate the two parts of the loss. Because most of the image content is background and most samples are negative samples, this will amplify the loss and cause the training result to be biased towards negative samples. Although weights are given to limit this situation, it still has an impact on the final result. If the MSE is used to calculate this part of the loss, the calculated value of the MSE is [0,1], and the model will not be biased towards the presence or absence of an object. At the same time, experiments show that the MSE has a performance not lower than that of the BCE function. Therefore, the MSE is used to calculate this part of the loss. So, the mean squared error function is used instead of the binary cross-entropy loss function to calculate the confidence loss of areas without targets. The formula for the new confidence loss function is as follows:
[0072]
[0073] In one embodiment, the present invention follows the classification loss function used in the YOLOv4 network, which is expressed as:
[0074]
[0075] Among them, c represents the classification probability, and p i (c) represents the classification probability of the predicted box, represents the classification probability of the ground truth box. When the j-th ground truth box in the i-th grid is responsible for a certain real target, the predicted box generated by this ground truth box will calculate the classification loss. In one embodiment, to evaluate the method proposed by the present invention, this embodiment conducts experiments on the Pascal VOC 2007 dataset. To ensure the comparability of the experimental results, all the training set and test set images of the experiments are exactly the same, and the experimental environment remains the same all the time. The experimental equipment is shown in the following table:
[0076] Table 1 Experimental environment configuration
[0077]
[0078] This embodiment selects the following indicators to evaluate the performance improvement of the method proposed by the present invention for YOLOv4.
[0079] The mean average precision (mAP), that is, the average of the accuracies of all classes, is expressed as:
[0080]
[0081] Among them, N represents the number of all categories, and AP is the average accuracy of a certain category. Its expression is:
[0082]
[0083] Among them, P is the accuracy of a certain category of samples, and R represents the recall rate. The expressions are as follows:
[0084]
[0085]
[0086] Among them, TP represents the number of samples correctly classified as positive samples, FP is the number of samples predicted as positive samples but actually negative samples; FN refers to the number of samples predicted as negative samples but actually positive samples.
[0087] F1-Score, also known as the F1 score, is a measure for classification problems and is often used as the final metric for multi-classification problems. It is the harmonic mean of precision and recall.
[0088]
[0089] Among them, recall k represents the recall rate, and precision k represents the accuracy rate.
[0090] The average misdetection rate (LAMR) represents the proportion of the test set in the dataset where the target is not detected. The larger the LAMR, the more objects are missed; the smaller the LAMR, the fewer objects are missed, which also indicates better model performance. The expression for the miss rate is:
[0091]
[0092] Specifically, in order to analyze the model performance of the LCIOU loss function (new bounding box regression loss function), the improved confidence function (fused confidence loss function), and the backbone network with the ECANet channel attention mechanism in this paper, combined experiments of different modules were established. The experimental dataset is PASCAL VOC2007, where 80% of the images are used for training and the rest are used for testing. The benchmark network for the experiment is YOLOv4, using the improved CSPDarkNet53 (improved backbone network) as the backbone extraction network, PANet and SPP as the neck networks, and the prediction network of YOLOv4 as the prediction network. The regression loss function is LCIOU. During training, the "leave-one-out" cross-validation method is used for the validation method. The input image size is 416×416, the batch size is 4, and the number of epochs is 50.
[0093] Table 2 Initial values of hyperparameters
[0094]
[0095] As shown in Table 3, various indicators of the three improved methods proposed in this paper under different combination forms prove the effectiveness of the method proposed in this paper. AP50 in the table represents the average value of AP for all categories when IoU is 0.5, and AP75 is the average value of AP for all categories when IoU is 0.75. AP represents the average value of AP when the value of IoU ranges from 0.5 to 0.95 with a step of 0.05. The advantages of LCIOU are mainly reflected in AP50, AP75, AP and Precision, with an increase of more than 2%. The confidence loss function mainly improves the accuracy of the model by affecting AP50 and recall rate, and their growth rates are 1.34% and 3.14% respectively. AP and AP75 increase by 1.07% and 0.63% respectively. After introducing the attention mechanism module, the accuracy of the model is mainly improved in terms of recall, with an increase of 1.9%.
[0096] Table 3 Network performance of different modules
[0097]
[0098] Since the improvement of the model accuracy by different modules is not independent, when different modules are combined, the increase in various evaluation indicators is not a simple addition of the increases of individual modules. Even some indicators have an increase lower than that of a single module, but the overall situation is reasonable. For example, when the LCIOU module is included, since the LCIOU module has a relatively large increase in AP and Precision, the combined module still has a relatively large increase in AP75. Similarly, when the fused confidence function is included, the increase in AP50 is relatively large. The model used in this paper uses the three methods simultaneously, and at this time, the convergence accuracy of the model has the largest improvement, and the increase in various indicators is the largest. The increases in AP50, AP75, AP and Precision are 1.01%, 2.44%, 3.00% and 2.72% respectively.
[0099] Such as visualization Figure 9 As shown, there are cases of missed detection and misdetection in YOLOv4, which mainly occur when the target is small or the target is occluded and not fully displayed. After using the three methods, the model obtains a more accurate training direction and extracts more feature information at the same time, and the detection effect reaches the best. Compared with the original model, it can not only detect difficult-to-detect targets, reduce the misdetection and missed detection rates, but also has a higher confidence for the detected targets.
[0100] The FPS indicates how many images the target model can detect per second and is used to evaluate the detection speed of the model. The larger the FPS, the faster the network detection speed. Among the three methods in this paper, LCIOU and the improved confidence function do not introduce additional parameters on the basis of the original model. Only the ECANet module introduces additional parameters. Therefore, the parameters increased by the model in this paper compared with YOLOv4 all come from the ECANet module. As shown in Table 4, compared with YOLOv4, the number of parameters increased by the model in this paper is only 9, which is attributed to the lightweight structure of the ECANet module. Although the image detection speed of the model in this paper is not as fast as that of the original model, the gap is only 0.62 ms. Such a gap will not affect the detection speed of the model. That is, the model in this paper can have good improvements in various evaluation indicators without reducing the detection speed of the model, improving the model accuracy.
[0101] Table 4 Comparison of model test time, FPS, and number of parameters
[0102]
[0103] To further analyze the improvement of the penalty term proposed in this paper on the model performance, this paper not only improves CIOU, but also improves DIOU and GIOU. This paper conducts experiments on the YOLOv4 model to compare the performance improvement of the penalty term proposed in this paper on the three loss functions. As shown in Table 5, the loss functions after adding the penalty term have increased in various evaluation indicators. Generally, the less border information carried by the loss function itself, the greater the performance improvement, which is attributed to the additional border information carried by the penalty term.
[0104] Table 5 Performance comparison of improved bounding box regression functions
[0105]
[0106] Table 6 shows the performance of several attention mechanisms on YOLOv4. Each module is used at the same position in YOLOv4, that is, in the backbone feature extraction network. It can be seen from the table that only the ECANet channel attention module improves the accuracy of the model, but this does not prove that other modules cannot improve the accuracy. An attention mechanism is added to the front part of the feature extraction network to enhance the expression of target information, so as to obtain better results in the subsequent feature extraction process. Adding the ECANet channel attention module can achieve the expected effect. At the same time, combined with the advantage mentioned above that it adds very few parameters and does not affect the detection speed of the model, this is the reason for this paper to choose the ECANet channel attention module.
[0107] Table 6 Performance of several attention mechanisms on YOLOV4
[0108]
[0109] This paper also compares the model accuracies of using the BCE and MSE functions in combination with using each function alone, as shown in Table 7. Obviously, in YOLOV4, the performance of the model using the two functions in combination is better, and multiple evaluation metrics have been significantly improved. Although the performance of the model using MSE alone is not as good as that using BCE, the accuracy of the targeted use of the MSE model has been improved well.
[0110] Table 7 Performance of Different Confidence Functions
[0111]
[0112] To compare the performance of the model in this paper, several current well-known object detection algorithms are selected, and the performance metrics of the models are compared in the same experimental environment to verify the performance of the model in this paper.
[0113] Table 8 Performance of Different Models
[0114]
[0115] As can be seen from Table 8, in the experiments with three image sizes, the model in this paper has the best performance in various evaluation metrics. Compared with YOLOv4, which has the second-best performance in the evaluation metrics, the performance of the model in this paper still has a good improvement. It is worth mentioning that while achieving a higher precision rate, the detection speed of the model in this paper is not lower than that of YOLOv4, that is, the model in this paper is also a model with a high balance between precision rate and speed.
[0116] In this invention, unless otherwise clearly specified and defined, terms such as "installation", "setting", "connection", "fixation", "rotation", etc. shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the internal communication of two components or the interaction relationship between two components. Unless otherwise clearly defined, for those of ordinary skill in the art, the specific meanings of the above terms in this invention can be understood according to specific circumstances.
[0117] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A target detection method based on an improved YOLOv4 network, characterized in that, it includes the following steps: S1. Build an improved YOLOv4 network based on the existing YOLOv4 network. The improved YOLOv4 network includes an improved backbone network, a neck network, and a head network; an ECANet channel attention module is added to the improved backbone network; S2. Obtain an image dataset to train the improved YOLOv4 network, and use a new bounding box regression loss function, a fused confidence loss function, and a classification loss function for iterative training to adjust the network parameters; Use the diagonal distance of the intersection rectangle between the ground truth box and the predicted box and the diagonal distance of the minimum bounding rectangle as additional bounding box information, construct an additional information penalty term based on the additional bounding box information, and introduce the additional information penalty term into the CIOU function to obtain a new bounding box regression loss function, where the CIOU function is expressed as: Among them, IOU represents the intersection over union of the bounding boxes of the ground truth box and the predicted box, w gt and h gt represent the width and height of the ground truth box respectively, w and h represent the width and height of the predicted box respectively, b represents the predicted box, b gt represents the ground truth box, p 2 (b, b gt ) represents the Euclidean distance between the centers of the ground truth box and the predicted box, c represents the diagonal distance of the minimum bounding rectangle of the ground truth box and the predicted box, α is a parameter function for balancing the aspect ratio, and v is a function for evaluating the proportional consistency between the ground truth box and the predicted box; Normalize the additional bounding box information, and construct a basic form of the additional information penalty term based on the normalized additional bounding box information. To ensure that the gradient direction of the additional information penalty term is the same as that of the CIOU function, the constructed basic form is expressed as: Then combine the norm advantage of the Huber function to construct the final additional information penalty term, expressed as: where, r represents the diagonal distance of the intersection rectangle between the ground truth box and the predicted box, c represents the diagonal distance of the minimum bounding rectangle between the ground truth box and the predicted box, λ represents the proportion of the additional information penalty term in the new bounding box regression function, and δ represents the error when the model is stable; Introduce the additional information penalty term into the CIOU function to obtain a new bounding box regression loss function, expressed as: Among them, R δ (diff) represents the extra information penalty term; S3. Use the trained improved YOLOv4 network for target detection to obtain the target detection result.
2. The target detection method based on an improved YOLOv4 network according to claim 1, characterized in that, Based on the backbone network of the existing YOLOv4 network, an improved backbone network is constructed by adding an ECANet channel attention module after each of the CBM layer, the CSP1 layer, and the CSP2 layer.
3. The target detection method based on an improved YOLOv4 network according to claim 1, characterized in that, The fused confidence loss function is expressed as: Among them, S 2 is the number of grids divided for each image, and B is the prior box of each grid generated by the network; and indicate whether there is an object in the j-th prior box of the i-th grid of the image. If there is an object in the prior box, at this time On the contrary, if there is no object in the prior box, then λ noobj represents the weight of the no-object loss calculation, indicates whether there is an object in the i-th grid. If there is an object, then otherwise it is 0; represents the probability of predicting the existence of an object, and the value range is [0, 1].