Traffic element multi-target detection method based on deep learning

By integrating the CBAM attention module and SwinT module in the YOLOv10 model, the backbone network is optimized, and the problem of traffic element detection accuracy and speed in complex scenarios is solved, more efficient traffic element detection is achieved, and the risk of traffic accidents is reduced.

CN120298741APending Publication Date: 2025-07-11ANHUI POLYTECHNIC UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510102599.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In a complex and changeable traffic environment, existing single-stage object detection algorithms such as YOLOv1 are difficult to meet the actual application needs in terms of detection accuracy and speed, especially in complex scenarios, which increases the risk of traffic accidents.

Method used

Build a multi-objective detection model for traffic elements based on YOLOv10. By integrating CBAM attention module and SwinT module, optimize the backbone network, enhance attention to the target area and anti-complex background capabilities, and improve detection accuracy.

Benefits of technology

It improves the traffic factor detection capability in complex scenarios, reduces the occurrence and impact of traffic accidents, and improves the safety of traffic travel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298741A_ABST
    Figure CN120298741A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic element multi-target detection method based on deep learning, and the method comprises the steps: firstly setting a plurality of traffic element types, collecting traffic road image data, carrying out the data labeling of a plurality of traffic elements in a traffic road image through employing a target bounding box labeling principle, and constructing a traffic data set and a label data set; then, a traffic element multi-target detection model is constructed based on YOLOv10, the traffic element multi-target detection model comprises a backbone network and a head prediction network, and the backbone network integrates a CBAM attention module and a Swi nT module; according to the invention, the traffic element detection capability in a complex scene is effectively improved, and early warning is given out in time, so that the occurrence of traffic accidents and the influence caused by the traffic accidents are reduced to the maximum extent, and the safety of traffic travel is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic data processing, and specifically to a multi-object detection method for traffic elements based on deep learning. Background Technique

[0002] Currently, with the continuous increase in the demand for traffic travel in modern society, the number of vehicles on the road is also increasing, and traffic safety issues have thus become the focus of public attention. Among the many causes of traffic accidents, the inability to observe traffic elements such as vehicles and pedestrians in a timely and effective manner is one of the important factors leading to traffic accidents and even casualties. Especially in complex and changeable traffic environments such as rainy days, snowy days, foggy days, and low light conditions, the observation ability of the human eye is greatly limited, making it difficult to quickly and accurately capture the dynamic information on the road. This undoubtedly increases the risk of traffic accidents and makes prevention particularly difficult. In addition, after a traffic accident occurs, the complex and changeable traffic environment often poses a great challenge to the determination of accident liability. For the above phenomena, how to avoid or minimize the harm and impact of traffic accidents can be achieved by detecting and identifying traffic elements in the pictures of videos, images, cameras, etc. through a monitoring system, timely identifying the positions of vehicles on the road and the behaviors of pedestrians, and providing real-time alerts and assistance to reduce the occurrence and impact of traffic accidents.

[0003] In the research field of traffic element detection, the rapid progress of deep learning technology has given rise to a series of new algorithms and technologies aimed at improving the accuracy and real-time performance of traffic element detection. In recent years, the research focus in this field has been on optimizing algorithm performance and enhancing the practicality of detection systems, especially under the strict requirements of simultaneously meeting high real-time performance and high accuracy. In this context, single-stage object detection algorithms have become a research hotspot, and SSD and YOLO are the classic representatives of single-stage object detection algorithms. Although the SSD algorithm performs well in terms of accuracy, its detection speed is relatively slow and it is difficult to meet the requirements of actual application scenarios. In contrast, since the first proposal of YOLOv1 in 2015, the YOLO series of algorithms have been continuously iterated and improved, gradually showing stronger competitiveness. Especially the YOLOv10 algorithm, by introducing a consistent dual assignment strategy without NMS training, effectively improves the efficiency and performance of the model during training and inference. In addition, YOLOv10 also adopts an overall efficiency-accuracy-driven model architecture design strategy, including lightweight classification heads, spatial-channel decoupled downsampling, and large kernel convolution and other technologies, further improving the efficiency and accuracy of the model. However, there is still room for improvement in the recognition accuracy of YOLOv10 in complex scenarios. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a multi-object detection method for traffic elements based on deep learning, construct a multi-object detection model for traffic elements based on YOLOv10, and through further optimization and adjustment of the YOLOv10 model, effectively improve the traffic element detection ability in complex scenarios, issue early warnings in a timely manner, thereby minimizing the occurrence of traffic accidents and their impacts, and improving the safety of traffic travel.

[0005] The technical solution of the present invention is as follows:

[0006] A multi-object detection method for traffic elements based on deep learning specifically includes the following steps:

[0007] (1) Set multiple traffic element categories, collect traffic road image data, perform data annotation on multiple traffic elements in the traffic road image using the target bounding box annotation principle, and thereby construct a traffic data set and a label data set for the traffic data set;

[0008] (2) Construct a multi-object detection model for traffic elements based on YOLOv10. The multi-object detection model for traffic elements includes a backbone network and a head prediction network. The backbone network includes three CBS modules, two groups of first combination modules, two groups of second combination modules, a SwinT module, and an SPPF module. Each group of first combination modules consists of a C2F module and a CBAM attention module, and each group of second combination modules consists of an SCDown module and a C2FCIB module. The processing process of the backbone network is shown in the following formula (1):

[0009]

[0010] In formula (1), F0 is the input of the backbone network, that is, the traffic data set, F1 is the intermediate feature map output after the processing of the second group of first combination modules of the backbone network, F2 is the intermediate feature map output after the processing of the first group of second combination modules of the backbone network, and F3 is the output of the backbone network;

[0011] The head prediction network adopts the head prediction network of the YOLOv10 model, which is used to generate bounding box predictions and corresponding class probabilities, and outputs prediction maps of multiple different scales to achieve the purpose of detecting multi-object traffic elements in traffic road images;

[0012] (3) Input the constructed traffic data set and label data set into the multi-object detection model for traffic elements for training to obtain a trained multi-object detection model for traffic elements for multi-object detection of traffic elements.

[0013] The processing process of the CBAM attention module is as follows: First, the input feature map O I enters the channel attention module, and the input feature map O IAfter global average pooling and global max pooling respectively, two feature maps are obtained. After the two feature maps are respectively fed into a multi-layer perceptron (MLP) with two fully-connected layers, element-wise addition and activation by the Sigmoid activation function are performed, and then it is combined with the input feature map O I to perform an element-wise multiplication operation, and the channel attention module outputs the feature map O CA , and the feature map O CA is used as the input of the spatial attention module. The feature map O CA After global average pooling and global max pooling respectively, the two output feature maps are concatenated based on channels, and then through a 7×7 convolution operation, the dimension is reduced to 1 channel. After activation by the Sigmoid activation function, a channel attention feature map is generated. Then, the channel attention feature map and the feature map O CA are subjected to an element-wise multiplication operation to obtain the output feature map of the CBAM attention module. The specific processing process is shown in the following formulas (2) and (3):

[0014] O CA = Sigmoid(MLP(AvgPool(O I )) + MLP(MAxPool(O I )))) ⊙ O I (2);

[0015] O CBAM = Sigmoid(Conv 7*7 (concat(AvgPool(O CA ), MaxPool(O CA )))) ⊙ O CA (3);

[0016] In formulas (2) and (3), AvgPool represents global average pooling, MaxPool represents global max pooling, MLP represents a multi-layer perceptron, Sigmoid represents the Sigmoid activation function, concat represents concatenation, and Conv 7*7 represents a 7×7 convolution operation, and ⊙ represents an element-wise multiplication operation.

[0017] The described SwinT module consists of a multi-layer perceptron, layer normalization, a window multi-head self-attention layer, and a sliding window multi-head self-attention layer. The processing process of the SwinT module is as follows: First, the input feature map Y l-1 After layer normalization processing, the attention within the window is calculated through the window multi-head self-attention layer, and then the feature map output by the window multi-head self-attention layer and the input feature map Y l-1 Perform a residual connection to fuse features at different levels and output the feature map after one residual operation Then, the feature map After layer normalization is performed again, a non-linear transformation is carried out through a multi-layer perceptron MLP, and then the feature map output after the non-linear transformation of the multi-layer perceptron MLP and the feature map are subjected to residual connection to output the feature map Y after the second residual l , the feature map Y l After layer normalization is performed again, the attention within the window is calculated through a sliding window multi-head self-attention layer, and then the feature map output by the sliding window multi-head self-attention layer and the feature map Y l are subjected to residual connection to output the feature map after the third residual Finally, the feature map After layer normalization is performed again, a second non-linear transformation is carried out through the multi-layer perceptron MLP again, and then the feature map output after the second non-linear transformation of the multi-layer perceptron MLP and the feature map are subjected to residual connection to output the feature map Y after the fourth residual l+1 , the feature map Y l+1 is the output of the SwinT module. The specific processing process is shown in the following formula (4):

[0018]

[0019] In formula (4), LN represents layer normalization, W_MSA represents the window multi-head self-attention layer, MLP represents the multi-layer perceptron, and SW_MSA represents the sliding window multi-head self-attention layer.

[0020] 4. A multi-object detection method for traffic elements based on deep learning according to claim 3, characterized in that: the processing process of the CBS module is as follows in formula (5):

[0021] CBS = SiLu(BN(Conv(O CBS ))) (5);

[0022] In formula (5), CBS represents the output of the CBS module, O CBS represents the input of the CBS module, Conv represents the convolution operation, BN represents batch normalization, and SiLu represents the SiLu activation function.

[0023] The C2F module first generates an intermediate feature map through convolution processing The generated intermediate feature map is split into two parts through the Split module and It is passed to the Bottleneck bottleneck layer for one-time processing to obtain a one-time output result. Then, the one-time output result is passed through the Bottleneck bottleneck layer again for secondary processing. Finally, the secondary processing result of the Bottleneck bottleneck layer, and After being concatenated and convolved, the feature map C2F is output. The specific processing process is shown in the following formula (6):

[0024]

[0025] In formula (6), C2F represents the output of the C2F module, P C2F represents the input of the C2F module, Conv represents the convolution operation, Bottleneck represents the Bottleneck bottleneck layer, and the Bottleneck bottleneck layer consists of two convolutional kernels with a size of 3 and a stride of 1 and a residual structure; concat represents concatenation.

[0026] The described SCDown module consists of two convolutional kernels. The size of the first convolutional kernel is 3 and the stride is 1, and the size of the second convolutional kernel is 3 and the stride is 2. The specific processing process is shown in the following formula (7):

[0027] SCDown = Conv(Conv(O SCDown )) (7);

[0028] In formula (7), O SCDown represents the input of the SCDown module, SCDown represents the output of the SCDown module, and Conv represents the convolution;

[0029] The described C2FCIB module first generates an intermediate feature map through convolution processing The generated intermediate feature map is split into two parts through the Split module and It is passed to the CIB module for one-time processing to obtain a one-time output result. Then, the one-time output result is passed through the CIB module again for secondary processing. Finally, the secondary processing result of the CIB module, and After being concatenated and convolved, the feature map C2FCIB is output. The specific processing process is shown in the following formula (8):

[0030]

[0031] In formula (8), C2FCIB represents the output of the C2FCIB module, O C2FCIBdenotes the input of the C2FCIB module, Conv denotes the convolution operation, and concat denotes concatenation; CIB denotes the CIB module, which is composed of two convolutional kernels and three depthwise separable convolutions, as shown in the following formula (9):

[0032] CIB = DWConv(Conv(DWConv(Conv(DWConv(O CIB ))))) (9);

[0033] In formula (9), O CIB denotes the input of the CIB module, CIB denotes the output of the CIB module, Conv denotes convolution, and DWConv denotes depthwise separable convolution.

[0034] The described head prediction network is composed of a PSA module, two UP upsamplings, four Concat concatenation modules, three C2FCIB modules, a C2F module, a CBS module, an SCDown module, and three Detect modules. The specific processing process is as follows in formula (10):

[0035]

[0036] In formula (10), O1 denotes the output of the PSA module, O2 denotes the output of one-time processing of the C2FCIB module, and O L 、O M 、O S are the feature maps of three scales output by the head prediction network respectively. After passing through the corresponding Detect modules respectively, three prediction maps of different scales are output.

[0037] Advantages of the present invention:

[0038] (1), The present invention integrates the CBAM attention module in the backbone network of YOLOv10, enhancing the attention of the traffic element multi-object detection model to the target area, enabling the traffic element multi-object detection model to adaptively learn the dependency relationship between different scale features, thereby improving the detection accuracy;

[0039] (2), In order to prevent interference from complex backgrounds, the present invention adds a SwinT (Swin Transformer) module to the backbone network of YOLOv10. While improving the anti-complex background ability of the backbone network, it enhances the ability of the backbone network to extract detailed information, and improves problems such as false detection and missed detection of small road targets. Brief Description of the Drawings

[0040] Figure 1 is the flowchart of the present invention.

[0041] Figure 2It is the framework structure diagram of the multi-object detection model for traffic elements of the present invention.

[0042] Figure 3 It is the framework structure diagram of the SwinT module of the present invention.

[0043] Figure 4 It is the comparison diagram of the object detection results of different algorithms for vehicles in the embodiment of the present invention.

[0044] Figure 5 It is the comparison of the multi-object detection results of different algorithms in complex scenarios in the embodiment of the present invention. Detailed implementation manners

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0046] See Figure 1 , a multi-object detection method for traffic elements based on deep learning, specifically including the following steps:

[0047] (1). Set multiple traffic element categories, including vehicles, pedestrians, traffic lights, traffic signs, etc. Extract and collect images from traffic videos to obtain traffic road image data. Use the labelImg annotation software to annotate the data of multiple traffic elements in the traffic road images, and thus construct a traffic data set and a label data set for the traffic data set. Then, perform object detection classification of the above traffic elements on the label data set after the traffic element bounding box annotation;

[0048] Establish an object detection sample data set, and divide the data set into a training set, a test set, and a validation set, where 80% is divided into the training set, 10% is divided into the test set, and 10% is divided into the validation set;

[0049] (2). See Figure 2 , build a multi-object detection model for traffic elements based on YOLOv10. The multi-object detection model for traffic elements includes a backbone network and a head prediction network. The backbone network includes three CBS modules, two groups of first combination modules, two groups of second combination modules, a SwinT module, and an SPPF module. Each group of first combination modules is composed of a C2F module and a CBAM attention module. Each group of second combination modules is composed of an SCDown module and a C2FCIB module. The processing process of the backbone network is shown in the following formula (1):

[0050]

[0051] In formula (1), F0 is the input of the backbone network, i.e., the traffic dataset, F1 is the intermediate feature map output after processing by the first combined module of the second group of the backbone network, F2 is the intermediate feature map output after processing by the second combined module of the first group of the backbone network, and F3 is the output of the backbone network;

[0052] Among them, the processing process of the CBAM attention module is as follows: First, the input feature map O I enters the channel attention module. The input feature map O I After passing through global average pooling and global max pooling respectively, two feature maps are obtained. After the two feature maps are respectively fed into a multi-layer perceptron MLP with two fully connected layers, then after element-wise addition and activation by the Sigmoid activation function, it is multiplied element-wise with the input feature map O I to perform an elementwise multiplication operation. The channel attention module outputs the feature map O CA , and the feature map O CA is used as the input of the spatial attention module. The feature map O CA After passing through global average pooling and global max pooling respectively, the two output feature maps are concatenated based on the channel, and then through a 7*7 convolution operation, the dimension is reduced to 1 channel. After activation by the Sigmoid activation function, a channel attention feature map is generated. Then, the channel attention feature map is multiplied element-wise with the feature map O CA to obtain the output feature map of the CBAM attention module. The specific processing process is shown in the following formulas (2) and (3):

[0053] O CA = Sigmoid(MLP(AvgPool(O I )) + MLP(MaxPool(O I ))) ⊙ O I (2);

[0054] O CBAM = Sigmoid(Conv 7*7 (Concat(AvgPool(O CA ), MaxPool(O CA )))) ⊙ O CA (3);

[0055] In formulas (2) and (3), AvgPool represents global average pooling, MaxPool represents global max pooling, MLP represents a multi-layer perceptron, Sigmoid represents the Sigmoid activation function, concat represents concatenation, and Conv 7*7Denotes a 7*7 convolution operation, and ⊙ denotes an elementwise multiplication operation;

[0056] See Figure 3 , The SwinT module consists of a multi-layer perceptron, layer normalization, window multi-head self-attention layer, and sliding window multi-head self-attention layer. The processing process of the SwinT module is as follows: First, the input feature map Y l-1 After layer normalization processing, calculate the attention within the window through the window multi-head self-attention layer, and then combine the feature map output by the window multi-head self-attention layer and the input feature map Y l-1 Perform a residual connection to fuse features at different levels and output the feature map after one residual Then normalize the feature map After layer normalization again, perform a non-linear transformation once through the multi-layer perceptron MLP, and then combine the feature map output after the non-linear transformation of the multi-layer perceptron MLP once and the feature map Perform a residual connection to output the feature map Y after two residuals l , Feature map Y l After layer normalization again, calculate the attention within the window through the sliding window multi-head self-attention layer, and then combine the feature map output by the sliding window multi-head self-attention layer and the feature map Y l Perform a residual connection to output the feature map after three residuals Finally, the feature map After layer normalization again, perform a second non-linear transformation through the multi-layer perceptron MLP again, and then combine the feature map output after the second non-linear transformation of the multi-layer perceptron MLP and the feature map Perform a residual connection to output the feature map Y after four residuals l+1 , Feature map Y l+1 Is the output of the SwinT module. The specific processing process is shown in the following formula (4):

[0057]

[0058] In formula (4), LN represents layer normalization, W_MSA represents the window multi-head self-attention layer, MLP represents the multi-layer perceptron, and SW_MSA represents the sliding window multi-head self-attention layer;

[0059] Among them, the processing process of the CBS module is shown in the following formula (5):

[0060] CBS = SiLu(BN(Conv(O CBS ))) (5);

[0061] In formula (5), CBS represents the output of the CBS module, and O CBSdenotes the input of the CBS module, Conv denotes the convolution operation, BN denotes batch normalization, and SiLu denotes the SiLu activation function;

[0062] Among them, the C2F module first generates an intermediate feature map through convolution processing The generated intermediate feature map is split into two parts by the Split module and They are passed to the Bottleneck bottleneck layer for one-time processing to obtain a one-time output result. Then, the one-time output result is passed through the Bottleneck bottleneck layer for secondary processing again. Finally, the secondary processing result of the Bottleneck bottleneck layer, and After concatenation and convolution processing, the output feature map C2F is obtained. The specific processing process is shown in the following formula (6):

[0063]

[0064] In formula (6), C2F represents the output of the C2F module, O C2F represents the input of the C2F module, Conv represents the convolution operation, Bottleneck represents the Bottleneck bottleneck layer, and the Bottleneck bottleneck layer consists of two convolutional kernels with a size of 3 and a stride of 1 and a residual structure; concat represents concatenation;

[0065] Among them, the SCDown module consists of two convolutional kernels. The size of the first convolutional kernel is 3 and the stride is 1, and the size of the second convolutional kernel is 3 and the stride is 2. The specific processing process is shown in the following formula (7):

[0066] SCDown = Conv(Conv(O SCDown )) (7);

[0067] In formula (7), O SCDown represents the input of the SCDown module, SCDown represents the output of the SCDown module, and Conv represents the convolution;

[0068] The C2FCIB module first generates an intermediate feature map through convolution processing The generated intermediate feature map is split into two parts by the Split module and They are passed to the CIB module for one-time processing to obtain a one-time output result. Then, the one-time output result is passed through the CIB module for secondary processing again. Finally, the secondary processing result of the CIB module, and After concatenation and convolution processing, the output feature map C2FCIB is obtained. The specific processing process is shown in the following formula (8):

[0069]

[0070] In Equation (8), C2FCIB represents the output of the C2FCIB module, O C2FCIB represents the input of the C2FCIB module, Conv represents the convolution operation, and concat represents the concatenation; CIB represents the CIB module, and the CIB module is composed of two convolutional kernels and three depthwise separable convolutions, as shown in the following Equation (9):

[0071] CIB = DWConv(Conv(DWConv(Conv(DWConv(O CIB ))))) (9);

[0072] In Equation (9), O CIB represents the input of the CIB module, CIB represents the output of the CIB module, Conv represents the convolution, and DWConv represents the depthwise separable convolution

[0073] The head prediction network adopts the head prediction network of the YOLOv10 model, which is used to generate bounding box predictions and corresponding class probabilities, outputs prediction maps of multiple different scales, and achieves the purpose of detecting multiple traffic elements in traffic road images;

[0074] The head prediction network is composed of a PSA module, two UP upsamplings, four Concat concatenation modules, three C2FCIB modules, one C2F module, one CBS module, one SCDown module, and three Detect modules. The specific processing process is as follows in Equation (10):

[0075]

[0076] In Equation (10), O1 represents the output of the PSA module, O2 represents the output of one-time processing of the C2FCIB module, O L 、O M 、O S are the feature maps of three scales output by the head prediction network respectively. After the feature maps of the three scales pass through the corresponding Detect modules respectively, three prediction maps of different scales are output;

[0077] (3) Input the training set into the multi-object detection model of traffic elements for training to obtain a trained multi-object detection model of traffic elements for multi-object detection of traffic elements. On the test set, test the trained multi-object detection model of traffic elements to verify its accuracy and efficiency in detecting traffic elements, including traffic element categories and confidence levels.

[0078] Performance analysis:

[0079] 1. Conduct traffic element multi-object detection experiments on the traffic element multi-object detection model (Swin-YOLOv10) of the present invention and the existing YOLOv10 model. The results of the traffic element multi-object detection experiment are shown in Table 1 below.

[0080] Table 1

[0081] Network model Pedestrian AP Vehicle AP Traffic light AP Traffic sign AP mAP FPS YOLOv10 0.663 0.822 0.589 0.663 0.684 47 Swin-YOLOv10 0.694 0.831 0.614 0.682 0.705 42

[0082] In Table 1, AP is Average Precision, that is, the average precision, which comprehensively considers the performance of Precision and Recall across the entire score range; mAP is mean average precision, that is, the mean average precision, representing the average value of AP for all classes; FPS is frames per second, that is, the number of frames transmitted per second, representing the detection speed.

[0083] As can be seen from Table 1, compared with YOLOv10, Swin-YOLOv10 has a greater advantage in the detection of traffic elements. Although the speed has decreased, the mAP has increased by 3%, that is, the detection accuracy has been greatly improved.

[0084] 2. Conduct traffic element multi-object detection experiments on Swin-YOLOv10 of the present invention and other model algorithms (YOLOv3, YOLOv5s, YOLOv5m, YOLOv7, Swin_YOLOv5_To). Among them, YOLOv3, YOLOv5s, YOLOv5m, and YOLOv7 are other versions of the YOLO series, and Swin_YOLOv5_To is an improved algorithm that replaces the backbone network in YOLOv5 with the SwinTransformer structure. The results of the traffic element multi-object detection experiment are shown in Table 2 below.

[0085]

[0086]

[0087] As can be seen from Table 2, in terms of the mAP value, Swin-YOLOv10 has increased by 2.4% compared with the sub-optimal YOLOv7 algorithm. For better illustration, Figure 4 Add two dotted lines to each target detection map of Figure 4It can be seen that during the detection process for vehicles, compared with YOLOv7, the detection results of Swin-YOLOv10 have a higher confidence level. Compared with YOLOv3, the detection of small targets is more accurate. Compared with YOLOv5s, YOLOv5m, and Swin_YOLOv5_To, there will be no misdetection of small targets (the blank area on the right side of the image is misdetected as a vehicle).

[0088] It is known from Figure 5 that during the detection process for multiple target types in complex scenarios, compared with YOLOv7, the detection results of Swin-YOLOv10 have a higher confidence level. Compared with YOLOv3, YOLOv5s, YOLOv5m, and Swin_YOLOv5_To, the detection of small targets is more accurate, reducing misdetection and missed detection.

[0089] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-object detection method for traffic elements based on deep learning, characterized in that: Specifically, it includes the following steps: (1) Set multiple traffic element categories, collect traffic road image data, perform data annotation on multiple traffic elements in the traffic road image using the target bounding box annotation principle, and construct a traffic dataset and a label dataset for the traffic dataset based on this; (2) Build a multi-object detection model for traffic elements based on YOLOv10. The multi-object detection model for traffic elements includes a backbone network and a head prediction network. The backbone network includes three CBS modules, two groups of first combination modules, two groups of second combination modules, a SwinT module, and an SPPF module. Each group of first combination modules consists of a C2F module and a CBAM attention module. Each group of second combination modules consists of an SCDown module and a C2FCIB module. The processing process of the backbone network is shown in the following formula (1): In formula (1), F0 is the input of the backbone network, that is, the traffic dataset, F1 is the intermediate feature map output after being processed by the second group of first combination modules of the backbone network, F2 is the intermediate feature map output after being processed by the first group of second combination modules of the backbone network, and F3 is the output of the backbone network; The head prediction network adopts the head prediction network of the YOLOv10 model, which is used to generate bounding box predictions and corresponding class probabilities, and outputs prediction maps of multiple different scales to achieve the purpose of detecting multi-object traffic elements in the traffic road image; (3) Input the constructed traffic dataset and label dataset into the multi-object detection model for traffic elements for training to obtain a trained multi-object detection model for traffic elements for multi-object detection of traffic elements.

2. The multi-object detection method for traffic elements based on deep learning according to claim 1, wherein: The processing process of the CBAM attention module is as follows: First, the input feature map O I enters the channel attention module. The input feature map O I After global average pooling and global max pooling respectively, two feature maps are obtained. After the two feature maps are sent into a multi-layer perceptron MLP with two fully connected layers respectively, element-wise addition and Sigmoid activation function activation are performed, and then element-wise multiplication operation is performed with the input feature map O I The channel attention module outputs the feature map O CA , and the feature map O CA is used as the input of the spatial attention module. The feature map O CA After global average pooling and global max pooling respectively, the output two feature maps are concatenated based on the channel, then through a 7×7 convolution operation, the dimension is reduced to 1 channel, and after Sigmoid activation function activation, a channel attention feature map is generated. Then the channel attention feature map and the feature map O CA perform element-wise multiplication operation to obtain the output feature map of the CBAM attention module. The specific processing process is shown in the following formulas (2) and (3): O CA = Sigmoid(MLP(AvgPool(O I ) + MLP(MaxPool(O I ))) ⊙ O I (2); O CBAM = Sigmoid(Conv 7*7 (concat(AvgPool(O CA ), MaxPool(O CA )))) ⊙ O CA (3); In equations (2) and (3), AvgPool represents global average pooling, MaxPool represents global max pooling, MLP represents multi-layer perceptron, Sigmoid represents the Sigmoid activation function, concat represents concatenation, and Conv 7×7 represents a 7×7 convolution operation, and ⊙ represents an elementwise multiplication operation.

3. The multi-object detection method for traffic elements based on deep learning according to claim 2, wherein: The described SwinT module consists of a multi-layer perceptron, layer normalization, window multi-head self-attention layer, and sliding window multi-head self-attention layer. The processing process of the SwinT module is as follows: First, for the input feature map Y l-1 After performing layer normalization processing, calculate the attention within the window through the window multi-head self-attention layer, and then combine the feature map output by the window multi-head self-attention layer and the input feature map Y l-1 Perform a residual connection to fuse features at different levels and output the feature map after the first residual connection Then, for the feature map After performing layer normalization again, perform a non-linear transformation once through the multi-layer perceptron MLP, and then combine the feature map output after the non-linear transformation of the multi-layer perceptron MLP once and the feature map Perform a residual connection to output the feature map Y after the second residual connection l , for the feature map Y l After performing layer normalization again, calculate the attention within the window through the sliding window multi-head self-attention layer, and then combine the feature map output by the sliding window multi-head self-attention layer and the feature map Y l Perform a residual connection to output the feature map after the third residual connection Finally, for the feature map After performing layer normalization again, perform a non-linear transformation twice through the multi-layer perceptron MLP, and then combine the feature map output after the non-linear transformation of the multi-layer perceptron MLP twice and the feature map Perform a residual connection to output the feature map Y after the fourth residual connection l+1 , for the feature map Y l+1 That is the output of the SwinT module. The specific processing process is shown in the following formula (4): In formula (4), LN represents layer normalization, W_MSA represents the window multi-head self-attention layer, MLP represents the multi-layer perceptron, and SW_MSA represents the sliding window multi-head self-attention layer.

4. A multi-object detection method for traffic elements based on deep learning according to claim 3, characterized in that: The processing process of the CBS module is shown in the following formula (5): CDS = SiLu(BN(Conv(O CBS ))) (5); In formula (5), CBS represents the output of the CBS module, and O CBS represents the input of the CBS module, Conv represents the convolution operation, BN represents batch normalization, and SiLu represents the SiLu activation function.

5. The multi-object detection method for traffic elements based on deep learning according to claim 4, characterized in that: The described C2F module first generates an intermediate feature map through convolution processing The generated intermediate feature map is split into two parts by the Split module and It is passed to the Bottleneck bottleneck layer for one-time processing to obtain a one-time output result. Then, the one-time output result is passed through the Bottleneck bottleneck layer again for secondary processing. Finally, after concatenating and convolving the secondary processing result of the Bottleneck bottleneck layer, and The feature map C2F is output after splicing and convolution processing. The specific processing process is shown in the following formula (6): In Equation (6), C2F represents the output of the C2F module, O C2F represents the input of the C2F module, Conv represents the convolution operation, Bottleneck represents the Bottleneck bottleneck layer. The Bottleneck bottleneck layer consists of two convolutional kernels with a size of 3 and a stride of 1 and a residual structure; concat represents concatenation.

6. The multi-object detection method for traffic elements based on deep learning according to claim 5, characterized in that: The SCDown module consists of two convolutional kernels. The size of the first convolutional kernel is 3 and the stride is 1. The size of the second convolutional kernel is 3 and the stride is 2. The specific processing process is shown in the following formula (7): SCDown = Conv(Conv(O SCDown )) (7); In Equation (7), O SCDown represents the input of the SCDown module, SCDown represents the output of the SCDown module, and Conv represents convolution; The described C2FCIB module first generates an intermediate feature map through convolution processing The generated intermediate feature map is split into two parts by the Split module and It is passed to the CIB module for one-time processing to obtain a one-time output result. Then, the one-time output result is passed through the CIB module again for secondary processing. Finally, the secondary processing result of the CIB module, and After concatenation and convolution processing, the feature map C2FCIB is output. The specific processing process is shown in the following formula (8): In Equation (8), C2FCIB represents the output of the C2FCIB module, O C2FCIB represents the input of the C2FCIB module, Conv represents the convolution operation, and concat represents the concatenation; CIB represents the CIB module, which is composed of two convolutional kernels and three depthwise separable convolutions, as shown in the following Equation (9): CIB = DWConv(Conv(DWConv(Conv(DWConv(O CIB ))))) (9); In Equation (9), O CIB represents the input of the CIB module, CIB represents the output of the CIB module, Conv represents convolution, and DWConv represents depthwise separable convolution.

7. A multi-object detection method for traffic elements based on deep learning according to claim 6, characterized in that: The head prediction network is composed of a PSA module, two UP upsampling modules, four Concat splicing modules, three C2FCIB modules, a C2F module, a CBS module, an SCDown module, and three Detect modules. The specific processing process is shown in the following formula (10): In Equation (10), O1 represents the output of the PSA module, O2 represents the output of the first - stage processing of the C2FCIB module, and O L , O M , O S are the feature maps of three scales output by the head prediction network respectively. After passing through the corresponding Detect modules respectively, three prediction maps of different scales are output.