Traffic target detection method and system fusing multi-scale features
Through the traffic object detection method that integrates multi-scale features, the multi-scale feature extractor and three-way semantic fusion module are used, combined with the hollow convolution and spatial enhancement perception mechanism, the problems of degradation of detection accuracy and missed detection in complex traffic scenarios are solved, and efficient and accurate traffic object detection is achieved, which meets the needs of high real-time and high precision, and improves the adaptability and universality of the model.
Patent Information
- Application Number
- CN202510182917.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-17
AI Technical Summary
The existing traffic target detection methods have reduced detection accuracy in complex traffic scenarios, and the problems of false detection and missed detection are prominent. Especially when dealing with small traffic targets, there is information loss and noise interference, which is difficult to meet the needs of high real-time and high precision, and the model adaptability and universality are insufficient.
The traffic object detection method that integrates multi-scale features is adopted, through a 27-layer detection model structure, combined with a multi-scale feature extractor and a three-way semantic fusion module, a multi-path structure is built using hollow convolution and spatial enhancement perception mechanisms to improve feature diversity and spatial attention focus capabilities.
It significantly improves detection performance, solves the problems of degradation of detection accuracy, false detection and missed detection caused by the diverse target size, light changes and occlusion, improves the accuracy of small target detection and detection reliability in complex backgrounds, meets the needs of traffic scenes for high real-time and high accuracy, and improves the adaptability and universality of the model under different lighting, weather and cross-regional traffic sign design differences.
Smart Images

Figure CN120164172A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of deep learning and image processing in artificial intelligence, and more specifically, relates to a traffic target detection method and system that fuses multi-scale features. Background Art
[0002] With the continuous growth of the car ownership and the increasingly complex road environment, traffic target detection faces challenges from external interference factors such as the diversity of target shapes and sizes, light changes, and occlusion, significantly increasing the detection difficulty of targets such as vehicles, road signs, and pedestrians. However, quickly and accurately identifying traffic targets is crucial for improving driving safety, assisting path planning and decision-making, realizing vehicle following, and enhancing vehicle autonomy. Traffic target detection not only provides key information for autonomous driving but also helps drivers drive safely, and has broad application prospects in fields such as autonomous driving, traffic monitoring, intelligent transportation systems, and urban planning.
[0003] Currently, traffic target detection is mainly divided into traditional feature extraction methods and deep learning-based methods: Traditional feature extraction methods rely on manually designed features (such as Scale-invariant feature transform (SIFT), Histogram of Oriented Gradient (HOG), etc.), and extract features through information such as angles, edges, and textures; Deep learning methods are divided into one-stage detection methods (such as the YOLO series, Single-shot detector (SSD), etc.) and two-stage detection methods (such as the Region-based Convolutional Neural Network (R-CNN) series). Among them, the one-stage method directly predicts the target category and bounding box from the image, and the two-stage method generates candidate boxes through the Region Proposal Network (RPN), and then performs classification and localization.
[0004] However, there are still some non-negligible defects in the above two existing traffic target detection methods:
[0005] (1) The one-stage detection method is more efficient, but in complex traffic scenarios, affected by factors such as the diversity of target shapes and sizes, light changes, and occlusion, the detection accuracy is significantly reduced, and the problems of false detection and missed detection are prominent;
[0006] (2) When the one-stage method processes small traffic targets, due to the small proportion and low resolution of the targets, information loss, noise interference, and detection box perturbation are likely to occur. Especially in the case of occlusion and complex backgrounds, the targets are easily ignored or misjudged;
[0007] (3) The two-stage detection method has high accuracy but insufficient real-time performance, making it difficult to simultaneously meet the dual requirements of high real-time performance and high accuracy in traffic scenarios, and unable to provide immediate decision-making support for autonomous driving and traffic monitoring systems.
[0008] (4) Traditional feature extraction methods rely on manually designed features and have poor robustness. Under different lighting conditions, weather conditions, and differences in traffic sign designs across regions, the adaptability and generality of the model are limited, making it difficult to handle complex traffic scenarios. Summary of the Invention
[0009] In view of the above-mentioned defects or improvement requirements of the prior art, the present invention provides a traffic target detection method and system that fuses multi-scale features. The purpose is to solve the technical problems of significantly reduced detection accuracy, false detection, and missed detection caused by factors such as the diversity of target shapes and sizes, light changes, and object occlusion in complex traffic scenarios, as well as the technical problems of information loss, noise interference, and low tolerance for detection box perturbations when detecting small traffic targets, especially in the case of occlusion and complex backgrounds, due to the small proportion and low resolution of small targets, which are easily overlooked or misjudged. It also solves the technical problems of being difficult to meet the dual requirements of high real-time performance and high accuracy in traffic scenarios, being difficult to support immediate decision-making for autonomous driving and traffic monitoring, and the lack of adaptability and generality of the model under different lighting conditions, weather conditions, and differences in traffic sign designs across regions, making it difficult to handle complex traffic scenarios.
[0010] To achieve the above object, according to one aspect of the present invention, a traffic target detection method that fuses multi-scale features is provided, including the following steps:
[0011] (1) Obtain a traffic image to be detected.
[0012] (2) Perform data preprocessing on the traffic image to be detected obtained in step (1) to obtain a preprocessed image.
[0013] (3) Input the preprocessed image obtained in step (2) into a pre-trained traffic target detection model that fuses multi-scale features to obtain a final detection result.
[0014] Preferably, step (2) is specifically as follows: First, adjust the size of the traffic image to be detected to 640×640×3; then, normalize the pixel values of the resized image from the range [0, 255] to the range [0, 1] to obtain a preprocessed image.
[0015] The detection result obtained in step (3) exists in the form of a detection box, and each detection box marks the predicted traffic target position and traffic target category.
[0016] Preferably, the traffic target detection model that fuses multi-scale features contains 27 layers, and its model structure is as follows:
[0017] The first layer takes an image with a dimension of 640×640×3 as the input. Perform the convolution normalization activation CBS operation on this image, and output a feature map with a dimension of 320×320×64.
[0018] The second layer takes the feature map output by the first layer as the input. Continue to perform the CBS operation on this feature map, and output a feature map with a dimension of 160×160×128.
[0019] The third layer takes the feature map output by the second layer as the input, and inputs this feature map into the first multi-scale feature extractor MFE module for feature extraction, and outputs a feature map with a dimension of 160×160×128.
[0020] The fourth layer takes the feature map output by the third layer as the input. First, perform a max pooling operation on this feature map to obtain a feature map with a dimension of 80×80×128; then, perform the CBS operation on this feature map with a dimension of 80×80×128 to output a feature map with a dimension of 80×80×256.
[0021] The fifth layer takes the feature map output by the fourth layer as the input. Input this feature map into the second MFE module for feature extraction to output a feature map with a dimension of 80×80×256.
[0022] The sixth layer takes the feature map output by the fifth layer as the input. First, perform a max pooling operation on this feature map to obtain a feature map with a dimension of 40×40×256; then, perform the CBS operation on this feature map with a dimension of 40×40×256 to output a feature map with a dimension of 40×40×512.
[0023] The seventh layer takes the feature map output by the sixth layer as the input. Input this feature map into the third MFE module for feature extraction to output a feature map with a dimension of 40×40×512.
[0024] The eighth layer takes the feature maps output by the third, fifth, and seventh layers as the input, and inputs these 3 feature maps into the three-way semantic fusion module TSF for feature fusion, and outputs a feature map with a dimension of 80×80×384.
[0025] The ninth layer takes the feature map output by the eighth layer as the input, and perform the CBS operation on this feature map, and output a feature map with a dimension of 40×40×256.
[0026] The tenth layer takes the feature maps output by the seventh and ninth layers as the input, and concatenates these 2 feature maps along the channel direction, and outputs a feature map with a dimension of 40×40×768.
[0027] Layer 11, with the input being the feature map output from Layer 10. Perform C2f operation on it, and output a feature map with a dimension of 40×40×512.
[0028] Layer 12, with the input being the feature map output from Layer 8. Perform upsampling operation on it, and output a feature map with a dimension of 160×160×384.
[0029] Layer 13, with the inputs being the feature maps output from Layer 3 and Layer 12. Concatenate these two feature maps along the channel direction, and output a feature map with a dimension of 160×160×512.
[0030] Layer 14, with the input being the feature map output from Layer 13. Perform C2f operation on it, and output a feature map with a dimension of 160×160×256.
[0031] Layer 15, with the inputs being the feature maps output from Layer 8, Layer 11, and Layer 14. Perform the second TSF operation on these three feature maps, and output a feature map with a dimension of 80×80×576.
[0032] Layer 16, with the input being the feature map output from Layer 15. Perform CBS operation on it, and output a feature map with a dimension of 40×40×256.
[0033] Layer 17, with the inputs being the feature maps output from Layer 9, Layer 11, and Layer 16. Concatenate these three feature maps along the channel direction, and output a feature map with a dimension of 40×40×1024.
[0034] Layer 18, with the input being the feature map output from Layer 17. Perform C2f operation on it, and output a feature map with a dimension of 40×40×512.
[0035] Layer 19, with the input being the feature map output from Layer 15. Perform upsampling operation on it, and output a feature map with a dimension of 160×160×576.
[0036] Layer 20, with the inputs being the feature maps output from Layer 12, Layer 14, and Layer 19. Concatenate these three feature maps along the channel direction, and output a feature map with a dimension of 160×160×1216.
[0037] Layer 21, with the input being the feature map output from Layer 20. Perform C2f operation on it, and output a feature map with a dimension of 160×160×256.
[0038] Layer 22, with the input being the feature map output from Layer 3. Perform CBS operation on it, and output a feature map with a dimension of 160×160×256.
[0039] The 23rd layer takes the feature map output from the 5th layer as input. A CBS operation is performed on it, and a feature map with a dimension of 80×80×576 is output.
[0040] The 24th layer takes the feature maps output from the 21st and 22nd layers as input. A Fusion operation is performed on these two feature maps, and a feature map with a dimension of 160×160×256 is output.
[0041] The 25th layer takes the feature maps output from the 15th and 23rd layers as input. A Fusion operation is performed on these two feature maps, and a feature map with a dimension of 80×80×576 is output.
[0042] The 26th layer takes the feature maps output from the 7th and 18th layers as input. A Fusion operation is performed on these two feature maps, and a feature map with a dimension of 40×40×512 is output.
[0043] The 27th layer takes the feature maps output from the 24th, 25th, and 26th layers as input. A Detect operation is performed on these three feature maps to obtain the final detection result.
[0044] Preferably, the MFE module includes two sub-modules: a three-way dilated convolution TDC sub-module and a spatial enhancement perception SAP sub-module;
[0045] The execution process of the MFE module is as follows: First, the feature map output from the 2nd layer is input into the TDC sub-module for processing; then, the feature map obtained by processing the TDC sub-module is added element-wise to the feature map output from the 2nd layer; then, the added feature map is input into the SAP sub-module for processing; finally, the feature map obtained by processing the SAP sub-module and the added feature map are added element-wise to obtain the feature map output by the MFE module, whose dimension is the same as that of the feature map output from the 2nd layer;
[0046] The process of the TDC sub-module processing the feature map output from the 2nd layer includes the following steps:
[0047] (a1) Perform a convolution operation on the feature map F (with a dimension of h×w×c) output from the second layer to obtain a feature map with a dimension of h×w×c, where the parameters of the convolution operation are: the number of output channels c, the convolution kernel size 3, the stride 1, and the padding value 1.
[0048] (a2) Divide the feature map output in step (a1) into two output feature maps F21 and F22 equally by channel, and their dimensions are both h×w×0.5c.
[0049] (a3) Perform a dilated convolution operation on the feature map F22 obtained in step (a2) to obtain a feature map with dimensions h×w×0.5c, where the parameters of the convolution operation are: the number of output channels is 0.5c, the convolution kernel size is 3, the stride is 1, the padding value is 2, and the dilation rate is 2.
[0050] (a4) Perform a dilated convolution operation on the feature map obtained in step (a3) to obtain a feature map with dimensions h×w×0.5c, where the parameters of the dilated convolution operation are: the number of output channels is 0.5c, the convolution kernel size is 3, the stride is 1, the padding value is 3, and the dilation rate is 3.
[0051] (a5) Perform maximum and average operations on the cross-channel corresponding pixel points of the feature map F22 obtained in step (a2) respectively to obtain two feature maps with dimensions h×w×1, and concatenate these two feature maps along the channel direction to obtain a feature map with dimensions h×w×2.
[0052] (a6) Perform maximum and average operations on the cross-channel corresponding pixel points of the feature map obtained in step (a3) respectively to obtain two feature maps with dimensions h×w×1, and concatenate these two feature maps along the channel direction to obtain a feature map with dimensions h×w×2.
[0053] (a7) Perform maximum and average operations on the cross-channel corresponding pixel points of the feature map obtained in step (a4) respectively to obtain two feature maps with dimensions h×w×1, and concatenate these two feature maps along the channel direction to obtain a feature map with dimensions h×w×2.
[0054] (a8) Concatenate the feature maps obtained in steps (a5), (a6) and (a7) along the channel direction to obtain a feature map with dimensions h×w×6.
[0055] (a9) Perform a convolution operation on the feature map obtained in step (a8) to obtain a feature map with dimensions h×w×3, where the parameters of the convolution operation are: the number of output channels is 3, the convolution kernel size is 7, the stride is 1, and the padding value is 3.
[0056] (a10) Perform a Sigmoid activation operation on the feature map obtained in step (a9) to obtain a feature map with dimensions h×w×3.
[0057] (a11) Divide the feature map obtained in step (a10) equally into three feature maps F111, F112 and F113 along the channel, and their dimensions are all h×w×1.
[0058] (a12) Multiply the feature map F22 obtained in step (a2) element-wise with the feature map F111 obtained in step (a11) to obtain a feature map with dimensions h×w×0.5c.
[0059] (a13) Element-wise multiply the feature map obtained in step (a3) with the feature map F112 obtained in step (a11) to obtain a feature map with dimensions h×w×0.5c.
[0060] (a14) Element-wise multiply the feature map obtained in step (a4) with the feature map F113 obtained in step (a11) to obtain a feature map with dimensions h×w×0.5c.
[0061] (a15) Element-wise add the feature maps obtained in steps (a12), (a13), and (a14) to obtain a feature map with dimensions h×w×0.5c.
[0062] (a16) Concatenate the feature map F21 obtained in step (a2) and the feature map obtained in step (a15) along the channel dimension, and perform the SiLU activation operation on the concatenated result to obtain a feature map with dimensions h×w×c.
[0063] (a17) Perform a convolution operation on the feature map output by the second layer to obtain a feature map with dimensions h×w×c, where the convolution operation parameters are: number of output channels c, convolution kernel size 1, stride value 1, and padding value 0.
[0064] (a18) Element-wise multiply the feature maps obtained in steps (a16) and (a17) to obtain a feature map with dimensions h×w×c.
[0065] (a19) Perform a convolution operation on the feature map obtained in step (a18) to obtain a feature map with dimensions h×w×c, where the convolution operation parameters are: number of output channels c, convolution kernel 1, stride 1, and padding value 0.
[0066] (a20) Element-wise add the feature map output by the second layer F and the feature map obtained in step (a19) to obtain a feature map with dimensions h×w×c, which is used as the output of the TDC sub-module;
[0067] The execution process of the SAP sub-module is as follows: Perform a convolution operation on the feature map (the convolution operation parameters are: number of output channels convolution kernel size 1, stride value 1, padding value 0) to obtain a feature map after the convolution operation with dimensions ; Subsequently, perform the SiLU activation operation on the feature map after the convolution operation to obtain the feature map after the activation operation; Then, perform a convolution operation on the feature map after the activation operation again (the convolution operation parameters are: number of output channels c, convolution kernel 1, stride 1, padding value 0) to obtain a feature map with dimensions h×w×c, which is used as the final output of the SAP sub-module.
[0068] Preferably, the process of the TSF module performing feature fusion on the feature maps output by the 3rd, 5th, and 7th layers specifically includes the following steps:
[0069] (b1) Perform an ADown operation on the feature map F output by the 3rd layer l (with dimensions h l ×w l ×dim l ) to obtain a feature map with dimensions h m ×w m ×0.5dim m .
[0070] (b2) Perform a CBS operation on the feature map F output by the 5th layer m (with dimensions h m ×w m ×dim m ) to obtain a feature map with dimensions h m ×w m ×0.5dim m .
[0071] (b3) Perform an upsampling operation on the feature map F output by the 7th layer h (with dimensions h h ×w h ×dim h ) and perform a CBS operation on the upsampled feature map to obtain a feature map with dimensions h m ×w m ×0.5dim m .
[0072] (b4) Concatenate the feature maps obtained in steps (b1), (b2), and (b3) along the channel dimension to obtain a feature map with dimensions h m ×w m ×1.5dim m .
[0073] (b5) Perform a depthwise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map with dimensions h m ×w m ×1.5dim m .
[0074] (b6) Perform a depthwise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map with dimensions h m ×w m ×1.5dim m .
[0075] (b7) Perform a depthwise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map with dimensions hm ×w m ×1.5dim m Feature map
[0076] (b8) Perform a depthwise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map with dimensions h m ×w m ×1.5dim m Feature map
[0077] (b9) Element-wise add the feature maps obtained in steps (b4), (b5), (b6), (b7), and (b8) to obtain a feature map with dimensions h m ×w m ×1.5dim m Feature map
[0078] (b10) Perform a CBS operation on the feature map obtained in step (b9) to obtain a feature map with dimensions h m ×w m ×1.5dim m Feature map
[0079] (b11) Element-wise add the feature maps obtained in step (b4) and step (b10) to obtain a feature map with dimensions h m ×w m ×1.5dim m Feature map, as the final output of the TFS sub-module
[0080] Preferably, the Detect operation performs forward propagation by dividing the input feature map into two branches:
[0081] Branch 1: The feature map passes through 2 CBS operations and 1 convolution operation in sequence. The parameters of the first 2 CBS operations are: kernel size 3, stride 1, padding value 1, and output channel number c t ; The parameters of the subsequent convolution operation are: kernel size 1, stride 1, padding value 1, and output channel number c t . This branch is used to calculate the bounding box regression loss of the target
[0082] Branch 2: The feature map passes through 2 CBS operations and 1 convolution operation in sequence. The parameters of the first 2 CBS operations are: kernel size 3, stride 1, padding value 1, and output channel number 4c t ; The parameters of the subsequent convolution operation are: kernel size 1, stride 1, padding value 1, and output channel number 4c t . This branch is used to calculate the classification loss of the target
[0083] Preferably, the traffic target detection model that fuses multi-scale features is trained through the following steps:
[0084] (4-1) Download the mixed dataset composed of the open-source BDD100K dataset and the KITTI dataset, and divide this mixed dataset into a training set and a test set according to the ratio of 8:2.
[0085] (4-2) Perform data preprocessing on the training set obtained in step (4-1) to obtain the preprocessed training set.
[0086] (4-3) Perform image enhancement processing on the preprocessed training set obtained in step (4-2) to obtain the enhanced training set.
[0087] (4-4) For each sample (with a dimension of 640×640×3) in the enhanced training set obtained in step (4-3), input this sample into the first layer of the traffic target detection model for processing to output the feature map corresponding to this sample with a dimension of 320×320×64.
[0088] (4-5) For each sample in the enhanced training set obtained in step (4-3), input the feature map corresponding to this sample with a dimension of 320×320×64 obtained in step (4-4) into the second layer of the traffic target detection model for processing to output the feature map corresponding to this sample with a dimension of 160×160×128.
[0089] (4-6) For each sample in the enhanced training set obtained in step (4-3), input the feature map corresponding to this sample with a dimension of 160×160×128 obtained in step (4-5) into the third layer of the traffic target detection model for multi-scale feature extraction to output the feature map corresponding to this sample with a dimension of 160×160×128.
[0090] (4-7) For each sample in the enhanced training set obtained in step (4-3), input the feature map corresponding to this sample with a dimension of 160×160×128 obtained in step (4-6) into the fourth layer of the traffic target detection model for downsampling to output the feature map corresponding to this sample with a dimension of 80×80×256.
[0091] (4-8) For each sample in the enhanced training set obtained in step (4-3), input the feature map corresponding to this sample with a dimension of 80×80×256 obtained in step (4-7) into the fifth layer of the traffic target detection model for multi-scale feature extraction to output the feature map corresponding to this sample with a dimension of 80×80×256.
[0092] (4-9) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 80×80×256 obtained in step (4-8) is input into the 6th layer of the traffic target detection model for downsampling to output a feature map corresponding to this sample with a dimension of 40×40×512.
[0093] (4-10) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 40×40×512 obtained in step (4-9) is input into the 7th layer of the traffic target detection model for multi-scale feature extraction to output a feature map corresponding to this sample with a dimension of 40×40×512.
[0094] (4-11) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 160×160×128 obtained in step (4-6), the feature map corresponding to this sample with a dimension of 80×80×256 obtained in step (4-8), and the feature map corresponding to this sample with a dimension of 40×40×512 obtained in step (4-10) are input into the 8th layer of the traffic target detection model for three-way semantic fusion to output a feature map corresponding to this sample with a dimension of 80×80×384.
[0095] (4-12) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 80×80×384 obtained in step (4-11) is input into the 9th layer of the traffic target detection model for processing to output a feature map corresponding to this sample with a dimension of 40×40×256.
[0096] (4-13) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 40×40×512 obtained in step (4-10) and each feature map with a dimension of 40×40×256 obtained in step (4-12) are input into the 10th layer of the traffic target detection model for processing to output a feature map corresponding to this sample with a dimension of 40×40×768.
[0097] (4-14) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 40×40×768 obtained in step (4-13) is input into the 11th layer of the traffic target detection model for processing to output a feature map corresponding to this sample with a dimension of 40×40×512.
[0098] (4-15) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 80×80×384 obtained in step (4-11) is input into the 12th layer of the traffic target detection model for upsampling to output a feature map corresponding to this sample with a dimension of 160×160×384.
[0099] (4-16) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 160×160×128 obtained in step (4-6) and each feature map with a dimension of 160×160×384 obtained in step (4-15) are input into the 13th layer of the traffic target detection model for processing to output a feature map corresponding to this sample with a dimension of 160×160×512.
[0100] (4-17) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 160×160×512 obtained in step (4-16) is input into the 14th layer of the traffic target detection model for processing to output a feature map corresponding to this sample with a dimension of 160×160×256.
[0101] (4-18) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 80×80×384 obtained in step (4-11), each feature map with a dimension of 40×40×512 obtained in step (4-14), and each feature map with a dimension of 160×160×256 obtained in step (4-17) are input into the 15th layer of the traffic target detection model for three-way semantic fusion to output a feature map corresponding to this sample with a dimension of 80×80×576.
[0102] (4-19) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a size of 80×80×576 obtained in step (4-18) is input into the 16th layer of the traffic target detection model for processing to output a feature map corresponding to this sample with a dimension of 40×40×256.
[0103] (4-20) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a size of 40×40×256 obtained in step (4-12), the feature map corresponding to this sample with a size of 40×40×512 obtained in step (4-14), and the feature map corresponding to this sample with a size of 40×40×256 obtained in step (4-19) into the 17th layer of the traffic target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 40×40×1024.
[0104] (4-21) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a size of 40×40×1024 obtained in step (4-20) into the 18th layer of the traffic target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 40×40×512.
[0105] (4-22) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a size of 80×80×576 obtained in step (4-18) into the 19th layer of the traffic target detection model for upsampling, so as to output a feature map corresponding to this sample with a dimension of 160×160×576.
[0106] (4-23) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a size of 160×160×384 obtained in step (4-15), the feature map corresponding to this sample with a size of 160×160×256 obtained in step (4-17), and the feature map corresponding to this sample with a size of 160×160×576 obtained in step (4-22) into the 20th layer of the traffic target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 160×160×1216.
[0107] (4-24) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a size of 160×160×1216 obtained in step (4-23) into the 21st layer of the traffic target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 160×160×256.
[0108] (4-25) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a size of 160×160×128 obtained in step (4-6) is input into the 22nd layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 160×160×256.
[0109] (4-26) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a size of 80×80×256 obtained in step (4-8) is input into the 23rd layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 80×80×576.
[0110] (4-27) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a size of 160×160×256 obtained in step (4-24) and the feature map corresponding to this sample with a size of 160×160×256 obtained in step (4-25) are input into the 24th layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 160×160×256.
[0111] (4-28) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a size of 80×80×576 obtained in step (4-18) and the feature map corresponding to this sample with a size of 80×80×576 obtained in step (4-26) are input into the 25th layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 80×80×576.
[0112] (4-29) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a size of 40×40×512 obtained in step (4-10) and the feature map corresponding to this sample with a size of 40×40×512 obtained in step (4-21) are input into the 26th layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 40×40×512.
[0113] (4-30) For each sample in the enhanced training set obtained in step (4-3), input the feature map corresponding to this sample with a size of 160×160×256 obtained in step (4-27), the feature map corresponding to this sample with a size of 80×80×576 obtained in step (4-28), and the feature map corresponding to this sample with a size of 40×40×512 obtained in step (4-29) into the 27th layer of the traffic target detection model for processing, so as to obtain the bounding box regression loss and classification loss corresponding to this sample.
[0114] (4-31) For each sample in the enhanced training set obtained in step (4-3), calculate the total loss L = αL CLS +βL DFL +γL WIoUv2 according to the bounding box regression loss and classification loss corresponding to this sample obtained in step (4-30), where α, β 和 γ represent loss weights.
[0115] (4-32) For each sample in the enhanced training set obtained in step (4-3), perform iterative training on the traffic target detection model that fuses multi-scale features according to the total loss corresponding to this sample obtained in step (4-31) and using gradient descent until the traffic target detection model that fuses multi-scale features reaches the preset number of iterations (which is 100 times in the present invention), and obtain the optimal parameters of the traffic target detection model that fuses multi-scale features at this time, so as to obtain a preliminarily trained traffic target detection model that fuses multi-scale features.
[0116] (4-33) Use the test set obtained in step (4-1) to test the traffic target detection model that fuses multi-scale features and is preliminarily trained in step (4-32) until the obtained detection accuracy reaches the optimum, so as to obtain a finally trained traffic target detection model that fuses multi-scale features.
[0117] Preferably, step (4-3) randomly selects one or any combination of the following 9 data augmentation methods for processing: image HSV augmentation, including hue, saturation, and brightness augmentation, and their augmentation factors are set to 0.015, 0.7, and 0.4 respectively; image translation, and its translation factor is 0.1; image scaling, and its scaling factor is 0.5; image left-right flipping, and its flipping factor is 0.5; multi-image stitching, and its stitching factor is 1.0; image erasing, and its erasing factor is 0.4; image cropping, and its cropping factor is 1.0.
[0118] Preferably, the classification loss L CLS is equal to the classification value predicted for this sample and the true classification value y clsThe binary cross-entropy loss between them is calculated as follows:
[0119]
[0120] According to another aspect of the present invention, there is provided a traffic target detection system that fuses multi-scale features, which is characterized by including:
[0121] A first module for acquiring a traffic image to be detected.
[0122] A second module for performing data preprocessing on the traffic image to be detected acquired by the first module to obtain a preprocessed image.
[0123] A third module for inputting the preprocessed image acquired by the second module into a pre-trained traffic target detection model that fuses multi-scale features to obtain a final detection result.
[0124] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0125] (1) The present invention designs a traffic target detection model through steps (4-1) to (4-33). By using a multi-scale feature extractor and a three-way semantic fusion module, combined with dilated convolution and a spatial enhancement perception mechanism, a multi-path structure is constructed to enhance feature diversity and spatial attention focusing ability. This model optimizes feature extraction and spatial perception, solves the problems of decreased detection accuracy, false detection, and missed detection caused by diverse target sizes, light changes, and occlusion, and significantly improves the detection performance, providing an efficient and accurate solution for traffic target detection;
[0126] (2) The present invention enhances the feature saliency of small targets through multi-image stitching and cropping in step (4-3); adopts three-level multi-scale feature extraction in steps (4-6) to (4-10) to strengthen the details of small targets; and designs a two-stage three-way semantic fusion in steps (4-11) to (4-18) to improve the completeness and distinctiveness of small target features. This solution effectively solves the problems of information loss, noise interference, and detection box perturbation caused by small proportion and low resolution of traffic small targets, and significantly improves the detection accuracy under occlusion and complex backgrounds.
[0127] (3) The present invention constructs a three-way dilated convolution module through steps (4-6), (4-8), and (4-10), uses multi-scale receptive fields to achieve fine extraction of target features, and combines the attention mechanism of cross-channel pooling and the Sigmoid function to adaptively allocate weights for key regions, enhancing feature discriminability and localization accuracy. In steps (4-11) and (4-18), depthwise separable convolution technology is adopted to design a multi-scale convolution kernel group, reducing the computational complexity while ensuring the feature extraction ability. This design solves the problem of insufficient real-time performance of traditional two-stage detection methods, meets the requirements of high real-time performance and high accuracy in traffic scenarios, and provides reliable decision-making support for autonomous driving and traffic monitoring systems;
[0128] (4) In steps (4-1) and (4-3) of the present invention, by combining an open-source dataset with data augmentation technology and using a large amount of diverse annotated data, the model can learn more comprehensive traffic target features, while reducing the production cost of new data samples and expanding the scale and diversity of the dataset. This method effectively solves the adaptability and generality problems of the model under different lighting conditions, weather, and cross-regional traffic sign design differences, significantly improving the generalization ability and detection accuracy of the model in complex traffic scenarios, and providing reliable support for the practical application of intelligent transportation systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0129] Figure 1 is a schematic diagram of the traffic target detection method using multi-scale feature fusion of the present invention;
[0130] Figure 2 is a schematic diagram of the training process of the traffic target detection method using multi-scale feature fusion of the present invention;
[0131] Figure 3 is a schematic diagram of the structure of the multi-scale feature extractor used in the present invention;
[0132] Figure 4 is a schematic diagram of the structure of the three-way dilated convolution module used in the present invention;
[0133] Figure 5 is a schematic diagram of the structure of the spatial enhancement perception module used in the present invention;
[0134] Figure 6 is a schematic diagram of the structure of the three-way semantic fusion module used in the present invention;
[0135] Figure 7 is a schematic diagram of the structure of the Detect module used in the present invention;
[0136] Figure 8 is a schematic diagram of the comparison of the detection results of the present invention and the existing method for complex traffic scene images with targets of diverse shapes and sizes;
[0137] Figure 9 It is a schematic diagram comparing the detection results of the present invention and existing methods for images of complex traffic scenes including light changes and object occlusions;
[0138] Figure 10 It is a schematic diagram comparing the detection results of the present invention and existing methods for images of complex traffic scenes including small targets. Detailed implementation manners
[0139] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0140] As Figure 1 shown, the present invention provides a traffic target detection method for fusing multi-scale features, including the following steps:
[0141] (1) Obtain a traffic image to be detected.
[0142] (2) Perform data preprocessing on the traffic image to be detected obtained in step (1) to obtain a preprocessed image.
[0143] Specifically, in this step, first, adjust the size of the traffic image to be detected to 640×640×3 (where 640×640 represents the resolution of the image and 3 represents the number of RGB channels); then, normalize the pixel values of the image with the adjusted size from the range [0, 255] to the range [0, 1], so as to obtain a preprocessed image.
[0144] (3) Input the preprocessed image obtained in step (2) into a pre-trained traffic target detection model for fusing multi-scale features to obtain a final detection result.
[0145] Specifically, the detection result obtained in this step exists in the form of a detection box, and each detection box marks the predicted traffic target position and traffic target category.
[0146] As Figure 2 shown, the traffic target detection model for fusing multi-scale features of the present invention includes 27 layers, and its model structure is as follows:
[0147] The first layer takes an image with a dimension of 640×640×3 as input. Perform a Convolution + Batch Normalization + SiLU (CBS) operation on this image. The parameters are: the number of output channels is 64, the convolution kernel size is 3, the stride is 2, the padding value is 1, and the output is a feature map with a dimension of 320×320×64.
[0148] The second layer takes the feature map output by the first layer as input. Continue to perform the CBS operation on this feature map. The parameters are: the number of output channels is 128, the convolution kernel size is 3, the stride is 2, the padding value is 1, and the output is a feature map with a dimension of 160×160×128.
[0149] The third layer takes the feature map output by the second layer as input and inputs this feature map into the first Multi-scale Feature Extractor (MFE) module for feature extraction, and the output is a feature map with a dimension of 160×160×128.
[0150] Specifically, the structure of the MFE module is as Figure 3 shown. It contains two sub-modules: the Triple-path Dilated Convolution (TDC) sub-module and the Spatial Augmented Perception (SAP) sub-module. The execution process of the MFE module is as follows: First, input the feature map output by the second layer into the TDC sub-module for processing; then, element-wise add the feature map processed by the TDC sub-module to the feature map output by the second layer; then, input the added feature map into the SAP sub-module for processing; finally, element-wise add the feature map processed by the SAP sub-module to the added feature map to obtain the feature map output by the MFE module, whose dimension is the same as that of the feature map output by the second layer.
[0151] Among them, the structure of the TDC sub-module is as Figure 4 shown. The TDC sub-module processes the feature map output by the second layer, and this process includes the following steps:
[0152] (a1) Perform a convolution operation on the feature map F output by the second layer (whose dimension is h×w×c) to obtain a feature map with a dimension of h×w×c. The parameters of the convolution operation are: the number of output channels is c, the convolution kernel size is 3, the stride is 1, and the padding value is 1.
[0153] (a2) Divide the feature map output in step (a1) into two output feature maps F21 and F22 equally by channel, and their dimensions are both h×w×0.5c.
[0154] (a3) Perform a dilated convolution operation on the feature map F22 obtained in step (a2) to obtain a feature map with dimensions h×w×0.5c, where the parameters of the convolution operation are: the number of output channels is 0.5c, the convolution kernel size is 3, the stride is 1, the padding value is 2, and the dilation rate is 2.
[0155] (a4) Perform a dilated convolution operation on the feature map obtained in step (a3) to obtain a feature map with dimensions h×w×0.5c, where the parameters of the dilated convolution operation are: the number of output channels is 0.5c, the convolution kernel size is 3, the stride is 1, the padding value is 3, and the dilation rate is 3.
[0156] (a5) Perform maximum and average operations on the cross-channel corresponding pixel points of the feature map F22 obtained in step (a2) respectively to obtain two feature maps with dimensions h×w×1, and concatenate these two feature maps along the channel direction to obtain a feature map with dimensions h×w×2.
[0157] (a6) Perform maximum and average operations on the cross-channel corresponding pixel points of the feature map obtained in step (a3) respectively to obtain two feature maps with dimensions h×w×1, and concatenate these two feature maps along the channel direction to obtain a feature map with dimensions h×w×2.
[0158] (a7) Perform maximum and average operations on the cross-channel corresponding pixel points of the feature map obtained in step (a4) respectively to obtain two feature maps with dimensions h×w×1, and concatenate these two feature maps along the channel direction to obtain a feature map with dimensions h×w×2.
[0159] (a8) Concatenate the feature maps obtained in steps (a5), (a6) and (a7) along the channel direction to obtain a feature map with dimensions h×w×6.
[0160] (a9) Perform a convolution operation on the feature map obtained in step (a8) to obtain a feature map with dimensions h×w×3, where the parameters of the convolution operation are: the number of output channels is 3, the convolution kernel size is 7, the stride is 1, and the padding value is 3.
[0161] (a10) Perform a Sigmoid activation operation on the feature map obtained in step (a9) to obtain a feature map with dimensions h×w×3.
[0162] (a11) Divide the feature map obtained in step (a10) equally into three feature maps F111, F112 and F113 along the channel, and their dimensions are all h×w×1.
[0163] (a12) Element - wise multiply the feature map F22 obtained in step (a2) with the feature map F111 obtained in step (a11) to obtain a feature map with dimensions h×w×0.5c.
[0164] (a13) Element - wise multiply the feature map obtained in step (a3) with the feature map F112 obtained in step (a11) to obtain a feature map with dimensions h×w×0.5c.
[0165] (a14) Element - wise multiply the feature map obtained in step (a4) with the feature map F113 obtained in step (a11) to obtain a feature map with dimensions h×w×0.5c.
[0166] (a15) Element - wise add the feature maps obtained in steps (a12), (a13), and (a14) to obtain a feature map with dimensions h×w×0.5c.
[0167] (a16) Concatenate the feature map F21 obtained in step (a2) and the feature map obtained in step (a15) along the channel dimension, and perform the SiLU activation operation on the concatenated result to obtain a feature map with dimensions h×w×c.
[0168] (a17) Perform a convolution operation on the feature map F output by the second layer to obtain a feature map with dimensions h×w×c, where the convolution operation parameters are: number of output channels c, convolution kernel size 1, stride value 1, and padding value 0.
[0169] (a18) Element - wise multiply the feature maps obtained in steps (a16) and (a17) to obtain a feature map with dimensions h×w×c.
[0170] (a19) Perform a convolution operation on the feature map obtained in step (a18) to obtain a feature map with dimensions h×w×c, where the convolution operation parameters are: number of output channels c, convolution kernel 1, stride 1, and padding value 0.
[0171] (a20) Element - wise add the feature map F output by the second layer and the feature map obtained in step (a19) to obtain a feature map with dimensions h×w×c, which is used as the output of the TDC sub - module.
[0172] The structure of the SAP sub - module is as Figure 5 shown, and it is used to enhance spatial attention. Its execution process is as follows: perform a convolution operation on the feature map (the convolution operation parameters are: number of output channels convolution kernel size 1, stride value 1, padding value 0) to obtain a feature map with dimensions The feature map after the convolution operation; subsequently, perform the SiLU activation operation on the feature map after the convolution operation to obtain the feature map after the activation operation; then, perform the convolution operation again on the feature map after the activation operation (the convolution operation parameters are: the number of output channels c, the convolution kernel 1, the stride 1, and the padding value 0) to obtain the feature map with the dimension of h×w×c as the final output of the SAP sub-module.
[0173] Layer 4, the input is the feature map output by Layer 3. First, perform the max pooling operation on this feature map (the pooling kernel is 2 and the stride is 2) to obtain the feature map with the dimension of 80×80×128; then, perform the CBS operation on the feature map with the dimension of 80×80×128 (parameters: the number of output channels 256, the convolution kernel size 1, the stride 1, and the padding value 0) to output the feature map with the dimension of 80×80×256.
[0174] Layer 5, the input is the feature map output by Layer 4. Input this feature map into the second MFE module for feature extraction to output the feature map with the dimension of 80×80×256.
[0175] Layer 6, the input is the feature map output by Layer 5. First, perform the max pooling operation on this feature map (the pooling kernel 2 and the stride 2) to obtain the feature map with the dimension of 40×40×256; then, perform the CBS operation on the feature map with the dimension of 40×40×256 (parameters: the number of output channels 512, the convolution kernel size 1, the stride 1, and the padding value 0) to output the feature map with the dimension of 40×40×512.
[0176] Layer 7, the input is the feature map output by Layer 6. Input this feature map into the third MFE module for feature extraction to output the feature map with the dimension of 40×40×512.
[0177] Specifically, the MFE modules in Layers 5 and 7 are exactly the same as that in Layer 3, which will not be elaborated here.
[0178] Layer 8, the input is the feature maps output by Layers 3, 5, and 7. Input these 3 feature maps into the triple-path semantic fusion module (abbreviated as TSF) for feature fusion to output the feature map with the dimension of 80×80×384.
[0179] Specifically, the structure of the TSF module is as Figure 6 shown. The process of its feature fusion for the feature maps output by Layers 3, 5, and 7 specifically includes the following steps:
[0180] (b1) For the feature map F output by Layer 3 l (with the dimension of h l ×wl ×dim l ) Perform the ADown operation to obtain a feature map with dimensions h m ×w m ×0.5dim m of the feature map.
[0181] Specifically, ADown is a lightweight downsampling operation, and its structure and functions are detailed in the paper "YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information" published on arXiv in 2024.
[0182] (b2) For the feature map F output by the 5th layer m (with dimensions h m ×w m ×dim m ) Perform the CBS operation to obtain a feature map with dimensions h m ×w m ×0.5dim m where the parameters of the CBS operation are: the number of output channels is 0.5dim m , the convolution kernel size is 1, the stride is 1, and the padding value is 0.
[0183] (b3) For the feature map F output by the 7th layer h (with dimensions h h ×w h ×dim h ) Perform the upsampling operation (parameters: the scaling factor is 2, and the mode is nearest), and then perform the CBS operation on the upsampled feature map to obtain a feature map with dimensions h m ×w m ×0.5dim m where the parameters of the CBS operation are: the number of output channels is 0.5dim m , the convolution kernel size is 1, the stride is 1, and the padding value is 0.
[0184] (b4) Concatenate the feature maps obtained in steps (b1), (b2), and (b3) along the channel dimension to obtain a feature map with dimensions h m ×w m ×1.5dim m of the feature map.
[0185] (b5) Perform the depthwise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map with dimensions h m ×w m ×1.5dim mThe feature map, where the parameters of the depthwise separable convolution operation are: the number of output channels is 1.5dim m , the convolution kernel size is 5, the stride is 1, the padding value is 2, and the number of groups is 1.5dim m .
[0186] (b6) Perform a depthwise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map with dimensions h m ×w m ×1.5dim m The feature map, where the parameters of the depthwise separable convolution operation are: the number of output channels is 1.5dim m , the convolution kernel size is 7, the stride is 1, the padding value is 3, and the number of groups is 1.5dim m .
[0187] (b7) Perform a depthwise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map with dimensions h m ×w m ×1.5dim m The feature map, where the parameters of the depthwise separable convolution are: the number of output channels is 1.5dim m , the convolution kernel size is 9, the stride is 1, the padding value is 4, and the number of groups is 1.5dim m .
[0188] (b8) Perform a depthwise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map with dimensions h m ×w m ×1.5dim m The feature map, where the parameters of the depthwise separable convolution operation are: the number of output channels is 1.5dim m , the convolution kernel size is 11, the stride is 1, the padding value is 5, and the number of groups is 1.5dim m .
[0189] (b9) Add the feature maps obtained in steps (b4), (b5), (b6), (b7), and (b8) element-wise to obtain a feature map with dimensions h m ×w m ×1.5dim m The feature map.
[0190] (b10) Perform a CBS operation on the feature map obtained in step (b9) to obtain a feature map with dimensions h m ×w m ×1.5dim m The feature map, where the parameters of the CBS operation are: the number of output channels is 1.5dim m , the convolution kernel size is 1, the stride is 1, and the padding value is 0.
[0191] (b11) Element-wise add the feature maps obtained in steps (b4) and (b10) to obtain a feature map with dimensions h m ×w m ×1.5dim m as the final output of the TFS sub-module.
[0192] Layer 9: The input is the feature map output from Layer 8. Perform a CBS operation on this feature map, and output a feature map with dimensions 40×40×256. The parameters of the CBS operation are: the number of output channels is 256, the convolutional kernel size is 3, the stride is 2, and the padding value is 1.
[0193] Layer 10: The inputs are the feature maps output from Layer 7 and Layer 9. Concatenate these two feature maps along the channel dimension, and output a feature map with dimensions 40×40×768.
[0194] Layer 11: The input is the feature map output from Layer 10. Perform a C2f operation on it, and output a feature map with dimensions 40×40×512. The parameters of the C2f operation are: the number of output channels is 512.
[0195] Specifically, the C2f operation is a module introduced by Ultralytics in the United States in YOLOv8. Its structure and functions are detailed in the open-source code library and official documentation of YOLOv8. Its parameter configuration is,
[0196] The first and last CBS operations: the number of output channels is c c , the convolutional kernel size is 1, the stride is 1, and the padding value is 0; it contains 3 Bottleneck modules, and each Bottleneck module contains 2 CBS operations. The first CBS operation: the number of output channels is 0.25c c , the convolutional kernel size is 3, the stride is 1, and the padding value is 1. The second CBS operation: the number of output channels is 0.5c c , the convolutional kernel size is 3, the stride is 1, and the padding value is 1.
[0197] Layer 12: The input is the feature map output from Layer 8. Perform an upsampling operation on it (parameters: scaling factor 2, mode is nearest), and output a feature map with dimensions 160×160×384.
[0198] Layer 13: The inputs are the feature maps output from Layer 3 and Layer 12. Concatenate these two feature maps along the channel dimension, and output a feature map with dimensions 160×160×512.
[0199] Layer 14: The input is the feature map output from Layer 13. Perform a C2f operation on it (parameters: the number of output channels is 256), and output a feature map with dimensions 160×160×256.
[0200] Specifically, the C2f operation of this layer is the same as that of the 11th layer, which will not be elaborated here.
[0201] For the 15th layer, the input is the feature maps output by the 8th, 11th, and 14th layers. The second TSF operation is performed on these 3 feature maps, and a feature map with a dimension of 80×80×576 is output.
[0202] Specifically, the TSF operation of this layer is the same as that of the 8th layer, which will not be elaborated here.
[0203] For the 16th layer, the input is the feature map output by the 15th layer. The CBS operation is performed on it, and a feature map with a dimension of 40×40×256 is output. The parameters of the CBS operation are: the number of output channels is 256, the convolution kernel size is 3, the stride is 2, and the padding value is 1.
[0204] For the 17th layer, the input is the feature maps output by the 9th, 11th, and 16th layers. These 3 feature maps are concatenated along the channel direction, and a feature map with a dimension of 40×40×1024 is output.
[0205] For the 18th layer, the input is the feature map output by the 17th layer. The C2f operation (parameters: the number of output channels is 512) is performed on it, and a feature map with a dimension of 40×40×512 is output.
[0206] Specifically, the C2f operation of this layer is the same as that of the 11th layer, which will not be elaborated here.
[0207] For the 19th layer, the input is the feature map output by the 15th layer. The upsampling operation (parameters: the scaling factor is 2, and the mode is nearest) is performed on it, and a feature map with a dimension of 160×160×576 is output.
[0208] For the 20th layer, the input is the feature maps output by the 12th, 14th, and 19th layers. These 3 feature maps are concatenated along the channel direction, and a feature map with a dimension of 160×160×1216 is output.
[0209] For the 21st layer, the input is the feature map output by the 20th layer. The C2f operation is performed on it, and a feature map with a dimension of 160×160×256 is output.
[0210] Specifically, the C2f operation of this layer is the same as that of the 11th layer, which will not be elaborated here.
[0211] For the 22nd layer, the input is the feature map output by the 3rd layer. The CBS operation is performed on it, and a feature map with a dimension of 160×160×256 is output. The parameters of the CBS operation are: the number of output channels is 256, the convolution kernel size is 1, the stride is 1, and the padding value is 0.
[0212] The 23rd layer takes the feature map output from the 5th layer as input. A CBS operation is performed on it, and a feature map with a dimension of 80×80×576 is output. The parameters of the CBS operation are: the number of output channels is 576, the convolution kernel size is 1, the stride is 1, and the padding value is 0.
[0213] The 24th layer takes the feature maps output from the 21st and 22nd layers as input, and performs a Fusion operation on these two feature maps, outputting a feature map with a dimension of 160×160×256.
[0214] Specifically, the details of the Fusion operation are described in the paper "DEA-Net: Single Image Dehazing Based on Detail-Enhanced Convolution and Content-Guided Attention" published by the team of the School of Aeronautics and Astronautics, Zhejiang University in IEEE Transactions on Image Processing in 2024.
[0215] The 25th layer takes the feature maps output from the 15th and 23rd layers as input, and performs a Fusion operation on these two feature maps, outputting a feature map with a dimension of 80×80×576.
[0216] The 26th layer takes the feature maps output from the 7th and 18th layers as input, and performs a Fusion operation on these two feature maps, outputting a feature map with a dimension of 40×40×512.
[0217] Specifically, the Fusion operations in the 25th and 26th layers are exactly the same as those in the 24th layer, and will not be elaborated here.
[0218] The 27th layer takes the feature maps output from the 24th, 25th, and 26th layers as input, and performs a Detect operation on these three feature maps to obtain the final detection result.
[0219] Specifically, the structure of the Detect operation is as Figure 7 shown, and the input feature map is divided into two branches for forward propagation:
[0220] Branch 1: The feature map passes through 2 CBS operations and 1 convolution operation in sequence. The parameters of the first 2 CBS operations are: the convolution kernel size is 3, the stride is 1, the padding value is 1, and the number of output channels is c t ; the parameters of the subsequent convolution operation are: the convolution kernel size is 1, the stride is 1, the padding value is 1, and the number of output channels is c t . This branch is used to calculate the bounding box regression loss of the target.
[0221] Branch 2: The feature map sequentially passes through 2 CBS operations and 1 convolution operation. The parameters of the first 2 CBS operations are: convolution kernel size 3, stride 1, padding value 1, and output channel number 4c t ; The parameters of the subsequent convolution operation are: convolution kernel size 1, stride 1, padding value 1, and output channel number 4c t . This branch is used to calculate the classification loss of the target.
[0222] The traffic target detection model that fuses multi-scale features of the present invention is obtained through the following steps of training:
[0223] (4-1) Download the mixed dataset composed of the open-source BDD100K dataset and the KITTI dataset, and divide the mixed dataset into a training set and a test set according to a ratio of 8:2.
[0224] The advantages of this step are as follows. First, the open-source dataset provides large-scale and diverse labeled data, covering traffic scenarios in different regions, lighting conditions, and weather, significantly improving the generalization ability of the model. Second, the dataset has been strictly labeled and verified, with high data quality, providing accurate supervision signals for model training. Finally, by fusing the BDD100K and KITTI datasets, the data diversity is further enhanced, making up for the limitations of a single dataset and enabling the model to perform more robustly in complex traffic scenarios.
[0225] (4-2) Perform data preprocessing on the training set obtained in step (4-1) to obtain the preprocessed training set.
[0226] Specifically, the data preprocessing process in this step is exactly the same as step (2) above and will not be elaborated here.
[0227] (4-3) Perform image enhancement processing on the preprocessed training set obtained in step (4-2) to obtain the enhanced training set.
[0228] Specifically, this step can randomly select one or any combination of the following 9 data enhancement methods for processing: image HSV enhancement, including hue, saturation, and brightness enhancement, with the enhancement factors set to 0.015, 0.7, and 0.4 respectively; image translation, with the translation factor of 0.1; image scaling, with the scaling factor of 0.5; image left-right flipping, with the flipping factor of 0.5; multi-image stitching, with the stitching factor of 1.0; image erasing, with the erasing factor of 0.4; image cropping, with the cropping factor of 1.0.
[0229] The advantage of this step is that this step expands the diversity of the training set through various data augmentation methods (such as HSV augmentation, translation, scaling, flipping, splicing, erasing, and cropping), simulating the lighting changes, perspective changes, and occlusion situations in the actual traffic scene. This not only improves the robustness and generalization ability of the model to complex scenes, but also alleviates the problem of insufficient data annotation, significantly improving the reliability and accuracy of the model in practical applications.
[0230] (4-4) For each sample (with a dimension of 640×640×3) in the augmented training set obtained in step (4-3), input this sample into the first layer of the traffic target detection model for processing to output a feature map corresponding to this sample with a dimension of 320×320×64.
[0231] (4-5) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a dimension of 320×320×64 obtained in step (4-4) into the second layer of the traffic target detection model for processing to output a feature map corresponding to this sample with a dimension of 160×160×128.
[0232] (4-6) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a dimension of 160×160×128 obtained in step (4-5) into the third layer of the traffic target detection model for multi-scale feature extraction to output a feature map corresponding to this sample with a dimension of 160×160×128.
[0233] The advantage of this step is that the multi-scale feature extractor MFE integrates the TDC and SAP sub-modules, achieving efficient and accurate feature extraction. The TDC sub-module, through a multi-path and multi-branch design, combines multi-dilation rate convolution and cross-channel pooling to enhance the diversity and expression ability of multi-scale features, and introduces a class attention mechanism to focus on key features; the SAP sub-module strengthens spatial attention through convolution and the SiLU activation function, improving the perception ability of complex spatial relationships. The two work together, not only ensuring the richness and non-linear expression ability of the features, but also optimizing the computational complexity, achieving efficient spatial feature extraction and multi-scale information fusion, significantly improving the adaptability and detection accuracy of the model to diverse traffic scenes, and providing strong technical support for traffic target detection.
[0234] (4-7) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a dimension of 160×160×128 obtained in step (4-6) into the fourth layer of the traffic target detection model for downsampling to output a feature map corresponding to this sample with a dimension of 80×80×256.
[0235] (4-8) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 80×80×256 obtained in step (4-7) is input into the 5th layer of the traffic target detection model for multi-scale feature extraction, so as to output the feature map corresponding to this sample with a dimension of 80×80×256.
[0236] (4-9) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 80×80×256 obtained in step (4-8) is input into the 6th layer of the traffic target detection model for downsampling, so as to output the feature map corresponding to this sample with a dimension of 40×40×512.
[0237] (4-10) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 40×40×512 obtained in step (4-9) is input into the 7th layer of the traffic target detection model for multi-scale feature extraction, so as to output the feature map corresponding to this sample with a dimension of 40×40×512.
[0238] The advantages of steps (4-6) to (4-10) are as follows: First, the feature expression ability is gradually improved through the three-level cascade structure of three MFE modules. The output of the previous module is used as the input of the next module, so that the features are continuously enriched and refined during transmission, and the multi-scale characteristics of the target are accurately captured. Second, the three-level cascade of the MFE module simulates the multi-scale feature extraction process from global to local, extracts global features first, and then focuses on local details, improving the detection ability for targets of different scales. Finally, this structure enhances the adaptability of the model to the diversity of target sizes, shapes and positions in traffic scenes, enabling it to achieve efficient and accurate target detection in different traffic scenes.
[0239] (4-11) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 160×160×128 obtained in step (4-6), the feature map corresponding to this sample with a dimension of 80×80×256 obtained in step (4-8), and the feature map corresponding to this sample with a dimension of 40×40×512 obtained in step (4-10) are input into the 8th layer of the traffic target detection model for three-way semantic fusion, so as to output the feature map corresponding to this sample with a dimension of 80×80×384.
[0240] The advantages of this step are as follows. First, based on the multi-scale feature fusion strategy, the TSF module integrates low-level, middle-level, and high-level feature maps to achieve multi-level feature fusion, effectively capturing local details and global semantic information, and significantly enhancing the multi-scale object detection ability in complex traffic scenarios. Second, lightweight operations such as the ADown efficient downsampling structure and depthwise separable convolution are adopted to reduce the computational complexity and the number of parameters, optimizing the lightweight design of the model. Finally, different-scale convolutional kernels are integrated through a multi-branch parallel structure to enhance the adaptability to target scale changes, and combined with the element-wise weighted fusion mechanism of feature maps to improve the robustness and discriminability of feature representation.
[0241] (4-12) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 80×80×384 obtained in step (4-11) is input into the 9th layer of the traffic object detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 40×40×256.
[0242] (4-13) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 40×40×512 obtained in step (4-10) and each feature map with a dimension of 40×40×256 obtained in step (4-12) are input into the 10th layer of the traffic object detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 40×40×768.
[0243] (4-14) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 40×40×768 obtained in step (4-13) is input into the 11th layer of the traffic object detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 40×40×512.
[0244] (4-15) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a dimension of 80×80×384 obtained in step (4-11) is input into the 12th layer of the traffic object detection model for upsampling, so as to output the feature map corresponding to this sample with a dimension of 160×160×384.
[0245] (4-16) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with dimensions 160×160×128 obtained in step (4-6) and each feature map with dimensions 160×160×384 obtained in step (4-15) are input into the 13th layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with dimensions 160×160×512.
[0246] (4-17) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with dimensions 160×160×512 obtained in step (4-16) is input into the 14th layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with dimensions 160×160×256.
[0247] (4-18) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with dimensions 80×80×384 obtained in step (4-11), each feature map with dimensions 40×40×512 obtained in step (4-14), and each feature map with dimensions 160×160×256 obtained in step (4-17) are input into the 15th layer of the traffic target detection model for three-way semantic fusion, so as to output the feature map corresponding to this sample with dimensions 80×80×576.
[0248] The advantages of steps (4-11) to (4-18) are that the first TSF module effectively captures the local details and global semantic information of the target by initially fusing low-level, middle-level, and high-level features; subsequently, the second TSF module further optimizes the feature fusion process on this basis and refines the feature representation, so as to more accurately extract the multi-scale target information in complex traffic scenes. This phased and hierarchical feature fusion strategy significantly improves the detection accuracy and robustness of the model in complex scenarios. It can not only more accurately identify traffic targets, but also effectively reduce the false detection and missed detection rates, showing excellent performance in practical applications.
[0249] (4-19) For each sample in the augmented training set obtained in step (4-3), the feature map corresponding to this sample with size 80×80×576 obtained in step (4-18) is input into the 16th layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with dimensions 40×40×256.
[0250] (4-20) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a size of 40×40×256 obtained in step (4-12), the feature map corresponding to this sample with a size of 40×40×512 obtained in step (4-14), and the feature map corresponding to this sample with a size of 40×40×256 obtained in step (4-19) into the 17th layer of the traffic target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 40×40×1024.
[0251] (4-21) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a size of 40×40×1024 obtained in step (4-20) into the 18th layer of the traffic target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 40×40×512.
[0252] (4-22) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a size of 80×80×576 obtained in step (4-18) into the 19th layer of the traffic target detection model for upsampling, so as to output a feature map corresponding to this sample with a dimension of 160×160×576.
[0253] (4-23) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a size of 160×160×384 obtained in step (4-15), the feature map corresponding to this sample with a size of 160×160×256 obtained in step (4-17), and the feature map corresponding to this sample with a size of 160×160×576 obtained in step (4-22) into the 20th layer of the traffic target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 160×160×1216.
[0254] (4-24) For each sample in the augmented training set obtained in step (4-3), input the feature map corresponding to this sample with a size of 160×160×1216 obtained in step (4-23) into the 21st layer of the traffic target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 160×160×256.
[0255] (4-25) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a size of 160×160×128 obtained in step (4-6) is input into the 22nd layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 160×160×256.
[0256] (4-26) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a size of 80×80×256 obtained in step (4-8) is input into the 23rd layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 80×80×576.
[0257] (4-27) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a size of 160×160×256 obtained in step (4-24) and the feature map corresponding to this sample with a size of 160×160×256 obtained in step (4-25) are input into the 24th layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 160×160×256.
[0258] (4-28) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a size of 80×80×576 obtained in step (4-18) and the feature map corresponding to this sample with a size of 80×80×576 obtained in step (4-26) are input into the 25th layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 80×80×576.
[0259] (4-29) For each sample in the enhanced training set obtained in step (4-3), the feature map corresponding to this sample with a size of 40×40×512 obtained in step (4-10) and the feature map corresponding to this sample with a size of 40×40×512 obtained in step (4-21) are input into the 26th layer of the traffic target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 40×40×512.
[0260] The advantages of steps (4-27) to (4-29) are that for the large, medium, and small target detection paths, the content-guided attention CGA and the fusion mechanism Fusion are respectively introduced. CGA dynamically focuses on key information, suppresses irrelevant features, and significantly reduces the phenomena of false detection and missed detection; the Fusion mechanism optimizes the feature expression by fusing multi-level features, further enhancing the model's discriminative ability for targets. This design not only improves the feature expression ability but also provides higher-quality feature inputs for the detection heads, achieving more accurate target recognition and localization, thereby significantly improving the overall detection performance of the model.
[0261] (4-30) For each sample in the enhanced training set obtained in step (4-3), the feature map of size 160×160×256 corresponding to this sample obtained in step (4-27), the feature map of size 80×80×576 corresponding to this sample obtained in step (4-28), and the feature map of size 40×40×512 corresponding to this sample obtained in step (4-29) are input into the 27th layer of the traffic target detection model for processing to obtain the bounding box regression loss and classification loss corresponding to this sample.
[0262] Specifically, the bounding box regression loss adopts WIoUv2 (Wise-IoU v2) and DEL (Distribution Focal Loss). The calculation formula and principle of WIoUv2 are detailed in the paper "Wise-IoU: Bounding Box Regression Loss with Dynamic Focusing Mechanism" published on arXiv in 2023, and the calculation formula and principle of DEL are detailed in the paper "Generalized Focal Loss: Towards Efficient Representation Learning for Dense Object Detection" published in IEEE Transactions on Pattern Analysis and Machine Intelligence in 2023.
[0263] The classification loss L CLS is equal to the binary cross-entropy loss between the predicted classification value of this sample cls and the true classification value y, and its calculation formula is:
[0264]
[0265] The advantage of this step is that it processes the two core tasks of traffic object detection, namely bounding box regression and classification, through a dual-branch structure. Among them, Branch 1 accurately captures spatial information with a smaller number of channels and optimized convolution parameters, significantly improving the accuracy of bounding box prediction; Branch 2 enhances the feature expression ability through a larger number of channels, further improving the classification accuracy. In addition, for traffic objects of different scales, large, medium, and small, the module processes them using feature maps of appropriate sizes respectively, achieving efficient multi-scale detection. This design not only greatly improves the adaptability of the model to complex traffic scenarios but also takes into account the high accuracy and real-time performance of detection, making it perform excellently in traffic object detection tasks.
[0266] (4-31) For each sample in the enhanced training set obtained in step (4-3), according to the bounding box regression loss and classification loss corresponding to this sample obtained in step (4-30), calculate the total loss L = αL CLS +βL DFL +γL WIoUv2 , where α, β 和 γ represent loss weights, and the value ranges of all three are positive real numbers. Preferably, α, β 和 γ are equal to 0.5, 1.5, and 7.5 respectively.
[0267] (4-32) For each sample in the enhanced training set obtained in step (4-3), according to the total loss corresponding to this sample obtained in step (4-31), and use gradient descent to iteratively train the traffic object detection model that fuses multi-scale features until the traffic object detection model that fuses multi-scale features reaches the preset number of iterations (in the present invention, it is 100 times), and obtain the optimal parameters of the traffic object detection model that fuses multi-scale features at this time, thereby obtaining a preliminarily trained traffic object detection model that fuses multi-scale features.
[0268] (4-33) Use the test set obtained in step (4-1) to test the traffic object detection model that fuses multi-scale features and is preliminarily trained in step (4-32) until the obtained detection accuracy reaches the optimal, thereby obtaining a finally trained traffic object detection model that fuses multi-scale features.
[0269] Example 1
[0270] To comprehensively evaluate the performance of the present invention, multi-dimensional evaluation indicators are adopted. In terms of object detection performance, average precision (AP) and mean average precision (mAP) are introduced. AP quantifies the detection performance of the model on a single category by calculating the area under the precision-recall curve; mAP evaluates the comprehensive detection performance of the model in a multi-category scenario by taking the average of the APs of all categories, reflecting its generalization ability and overall performance.
[0271] In terms of evaluating the practicality of the model, three key metrics are introduced: Frames Per Second (FPS), Giga Floating Point Operations Per Second (GFLOPs), and the number of model parameters (Parameters). FPS measures the real-time processing ability of the model, reflecting the number of frames processed per unit time; GFLOPs measures the computational complexity of the model, reflecting the amount of floating-point operations required for a single forward pass; Parameters measures the scale and storage requirements of the model, reflecting the total number of trainable parameters. These metrics jointly evaluate the real-time performance, computational resource requirements, and storage efficiency of the model.
[0272] Test Case 1
[0273] To verify the performance advantages of the method of the present invention in traffic object detection tasks, comparative tests were conducted on the KITTI dataset with 11 mainstream object detection methods. The test metrics include FPS, GFLOPs, the number of model parameters, AP, and mAP, and a comprehensive evaluation was carried out from aspects such as real-time performance, computational complexity, model scale, and detection accuracy. The experimental results are shown in Table 1 for details.
[0274] Table 1 Performance Comparison between the Present Invention and Existing Methods
[0275]
[0276]
[0277] According to the experimental results in Table 1, the method of the present invention has significant advantages in object detection performance: the mAP reaches 96.7%, and the AP values for the three categories of vehicles, bicycles, and pedestrians are 87.3%, 85.5%, and 89.8% respectively, all superior to the 11 comparative methods, indicating its excellent detection accuracy and robustness in complex traffic scenarios.
[0278] In terms of computational efficiency, the FPS of this method is 109.29, ranking third, demonstrating good real-time processing ability and meeting the requirements of real-time scenarios such as autonomous driving. Although the GFLOPs ranks tenth and the computational complexity is relatively high, considering its excellent detection performance, the complexity is still within an acceptable range and there is room for optimization. In addition, the number of model parameters is 13.2M, ranking second, achieving a good balance between model scale and storage efficiency.
[0279] In summary, the method of the present invention achieves an effective trade-off among key metrics such as detection accuracy, real-time performance, computational complexity, and model scale, has high practical value and promotion potential, and provides reliable technical support for practical applications in the field of traffic object detection.
[0280] Test Case 2
[0281] Through visual analysis on the KITTI dataset, the detection performance of the method of the present invention in complex traffic scenarios is further verified. As Figure 8 shown, the method of the present invention effectively avoids the problem of missed detection of two vehicles with different sizes by YOLOv8m, demonstrating better object detection capabilities, especially outstanding performance in small object detection.
[0282] As Figure 9 shown, in complex scenarios with light changes and occlusions, the method of the present invention successfully detects the semi-occluded vehicle missed by YOLOv8m, indicating its stronger robustness and detection accuracy under light changes and occlusion conditions.
[0283] As Figure 10 shown, in scenarios containing small bicycle objects, the method of the present invention successfully detects the small bicycle objects in the distance missed by YOLOv8m, further proving its significant advantages in small object detection tasks.
[0284] The above performance improvement is attributed to the innovative design of the present invention: through multi-scale feature extraction and three-way semantic fusion mechanism, the adaptability of the model to small objects, occluded objects, multi-scale objects and light changes in complex scenarios is enhanced, the missed detection rate is effectively reduced, and the detection reliability is improved, providing a better solution for object detection in complex traffic scenarios.
Claims
1. A traffic target detection method integrating multi-scale features, characterized in that: The following steps are involved: (1) Obtain the traffic image to be detected. (2) Performing data preprocessing on the traffic image to be detected obtained in step (1) to obtain a preprocessed image. (3) The preprocessed image obtained in step (2) is input into a pre-trained traffic target detection model that integrates multi-scale features to obtain the final detection result.
2. The traffic target detection method integrating multi-scale features according to claim 1 is characterized in that: Step (2) specifically includes, first, adjusting the size of the traffic image to be detected to 640×640×3; then, normalizing the pixel values of the resized image from the range of [0, 255] to the range of [0, 1], thereby obtaining a preprocessed image; The detection result obtained in step (3) is in the form of a detection frame, each of which indicates the predicted traffic target location and traffic target category.
3. The traffic target detection method integrating multi-scale features according to claim 1 or 2, characterized in that: The traffic target detection model integrating multi-scale features contains 27 layers, and its model structure is as follows: In the first layer, the input is an image with a dimension of 640× 640× 3. The convolution normalization activation CBS operation is performed on the image, and the output feature map has a dimension of 320×320×64. The second layer takes as input the feature map output by the first layer. The CBS operation is continued on the feature map, and the output feature map has a dimension of 160×160×128. The third layer inputs the feature map output by the second layer, which is input into the first multi-scale feature extractor MFE module for feature extraction, and the output dimension is a feature map of 160× 160× 128. The fourth layer takes as input the feature map output by the third layer. First, a maximum pooling operation is performed on the feature map to obtain a feature map with a dimension of 80× 80× 128. Then, a CBS operation is performed on the feature map with a dimension of 80× 80× 128 to output a feature map with a dimension of 80× 80× 256. The fifth layer takes as input the feature map output by the fourth layer. The feature map is input to the second MFE module for feature extraction to output a feature map with a dimension of 80×80×256. The sixth layer takes as input the feature map output by the fifth layer. First, a maximum pooling operation is performed on the feature map to obtain a feature map with a dimension of 40×40×256. Then, a CBS operation is performed on the feature map with a dimension of 40×40×256 to output a feature map with a dimension of 40×40×512. The input of the 7th layer is the feature map output by the 6th layer. The feature map is input to the 3rd MFE module for feature extraction to output a feature map with a dimension of 40×40×512. The 8th layer takes as input the feature maps output by the 3rd, 5th, and 7th layers. These three feature maps are input into the three-way semantic fusion module TSF for feature fusion, and the output feature map is a feature map with a dimension of 80×80×384. The ninth layer takes as input the feature map output by the eighth layer, performs the CBS operation on the feature map, and outputs a feature map with a dimension of 40×40×256. The 10th layer takes as input the feature maps output by the 7th and 9th layers. These two feature maps are concatenated along the channel direction, and the output dimension is a feature map of 40×40×768. The 11th layer takes as input the feature map output by the 10th layer. A C2f operation is performed on it, and the output feature map is a 40×40×512 feature map. The 12th layer takes as input the feature map output by the 8th layer. An upsampling operation is performed on it, and the output feature map has a dimension of 160×160×384. The input of the 13th layer is the feature map output by the 3rd and 12th layers. These two feature maps are concatenated along the channel direction, and the output dimension is a feature map of 160×160×512. The 14th layer takes as input the feature map output by the 13th layer. A C2f operation is performed on it, and the output feature map is 160×160×256 in size. At the 15th layer, the input is the feature map output by the 8th, 11th, and 14th layers. The second TSF operation is performed on these three feature maps, and the output dimension is a feature map of 80×80×576. The 16th layer takes as input the feature map output by the 15th layer. The CBS operation is performed on it, and the output feature map has a dimension of 40×40×256. The input of the 17th layer is the feature map output by the 9th, 11th, and 16th layers. These three feature maps are concatenated along the channel direction, and the output dimension is a feature map of 40×40×1024. The 18th layer takes as input the feature map output by the 17th layer. A C2f operation is performed on it, and the output feature map is a 40×40×512 feature map. The 19th layer takes as input the feature map output by the 15th layer. An upsampling operation is performed on it, and the output feature map has a dimension of 160×160×576. The input of the 20th layer is the feature map output by the 12th, 14th, and 19th layers. These three feature maps are concatenated along the channel direction, and the output dimension is a feature map of 160×160×1216. The 21st layer takes as input the feature map output by the 20th layer. A C2f operation is performed on it, and the output feature map has a dimension of 160×160×256. The 22nd layer takes as input the feature map output by the 3rd layer. A CBS operation is performed on it, and the output feature map has a dimension of 160×160×256. The 23rd layer takes as input the feature map output by the 5th layer. A CBS operation is performed on it, and the output feature map is 80×80×576 in size. The 24th layer takes as input the feature maps output by the 21st and 22nd layers. These two feature maps are fused and the output feature map is a 160×160×256 feature map. The 25th layer takes as input the feature maps output by the 15th and 23rd layers. These two feature maps are fused and the output feature map is 80×80×576 in dimension. The 26th layer takes as input the feature maps output by the 7th and 18th layers. These two feature maps are fused and the output feature map is a 40×40×512 feature map. The 27th layer takes as input the feature maps output by the 24th, 25th, and 26th layers, and performs Detect operations on these three feature maps to obtain the final detection results.
4. The method for traffic target detection by integrating multi-scale features according to any one of claims 1 to 3, characterized in that: The MFE module consists of two submodules: a three-way dilated convolution TDC submodule and a spatial enhancement perception SAP submodule; The execution process of the MFE module is as follows: first, the feature map output by the second layer is input into the TDC submodule for processing; then, the feature map processed by the TDC submodule is added element by element with the feature map output by the second layer; Then, the added feature map is input into the SAP submodule for processing; finally, the feature map processed by the SAP submodule and the added feature map are added element by element to obtain the feature map output by the MFE module, and its dimension is consistent with the feature map output by the second layer; The TDC submodule processes the feature map output by the second layer. This process includes the following steps: (a1) Perform a convolution operation on the feature map F (whose dimension is h×w×c) output by the second layer to obtain a feature map of dimension h×w×c, where the parameters of the convolution operation are: number of output channels c, convolution kernel size 3, stride 1, and padding value 1. (a2) The feature map output by step (a1) is equally divided into two output feature maps F21 and F22 according to the channel, and their dimensions are both h×w×0.5c. (a3) Perform a dilated convolution operation on the feature map F22 obtained in step (a2) to obtain a feature map with a dimension of h×w×0.5c, where the parameters of the convolution operation are: number of output channels 0.5c, convolution kernel size 3, step size 1, padding value 2, and dilation rate 2. (a4) Perform a dilated convolution operation on the feature map obtained in step (a3) to obtain a feature map with a dimension of h×w×0.5c, where the parameters of the dilated convolution operation are: output channel number 0.5c, convolution kernel size 3, step size 1, padding value 3, and dilation rate 3. (a5) Perform maximum and average operations on the corresponding pixel points across channels of the feature map F22 obtained in step (a2) to obtain two feature maps with dimensions of h×w×1, and concatenate the two feature maps along the channel direction to obtain a feature map with dimensions of h×w×2. (a6) Perform maximum and average operations on the corresponding pixels across channels of the feature map obtained in step (a3) to obtain two feature maps with dimensions of h×w×1, and concatenate the two feature maps along the channel direction to obtain a feature map with dimensions of h×w×2. (a7) Perform maximum and average operations on the corresponding pixels across channels of the feature map obtained in step (a4) to obtain two feature maps with a dimension of h×w×1, and concatenate the two feature maps along the channel direction to obtain a feature map with a dimension of h×w×2. (a8) Concatenate the feature maps obtained in steps (a5), (a6) and (a7) along the channel direction to obtain a feature map with a dimension of h×w×6. (a9) Perform a convolution operation on the feature map obtained in step (a8) to obtain a feature map of dimension h×w×3, where the convolution operation parameters are: output channel number 3, convolution kernel size 7, step size 1, and padding value 3. (a10) Perform a Sigmoid activation operation on the feature map obtained in step (a9) to obtain a feature map with a dimension of h×w×3. (a11) The feature map obtained in step (a10) is equally divided into three feature maps F111, F112 and F113 according to the channel, and their dimensions are all h×w×1. (a12) Multiply the feature map F22 obtained in step (a2) by the feature map F111 obtained in step (a11) element by element to obtain a feature map with a dimension of h×w×0.5c. (a13) Multiply the feature map obtained in step (a3) by the feature map F112 obtained in step (a1) element by element to obtain a feature map with a dimension of h×w×0.5c. (a14) Multiply the feature map obtained in step (a4) by the feature map F113 obtained in step (a1) element by element to obtain a feature map with a dimension of h×w×0.5c. (a15) Add the feature maps obtained in steps (a12), (a13), and (a14) element by element to obtain a feature map with a dimension of h×w×0.5c. (a16) The feature map F21 obtained in step (a2) and the feature map obtained in step (a15) are concatenated along the channel direction, and a SiLU activation operation is performed on the concatenated result to obtain a feature map with a dimension of h×w×c. (a17) Perform a convolution operation on the feature map F output by the second layer to obtain a feature map of dimension h×w×c, where the convolution operation parameters are: number of output channels c, convolution kernel size 1, step value 1, and padding value 0. (a18) Multiply the feature maps obtained in step (a16) and step (a17) element by element to obtain a feature map with a dimension of h×w×c. (a19) Perform a convolution operation on the feature map obtained in step (a18) to obtain a feature map of dimension h×w×c, where the convolution operation parameters are: number of output channels c, convolution kernel 1, step size 1, and padding value 0. (a20) adding the feature map F output by the second layer to the feature map obtained in step (a19) element by element to obtain a feature map with a dimension of h×w×c as the output of the TDC submodule; The execution process of the SAP submodule is as follows: Perform a convolution operation on the feature map (the convolution operation parameters are: the number of output channels The convolution kernel size is 1, the step value is 1, and the padding value is 0) to obtain the dimension The feature map after the convolution operation; Subsequently, a SiLU activation operation is performed on the feature map after the convolution operation to obtain the feature map after the activation operation; then, a convolution operation is performed again on the feature map after the activation operation (the convolution operation parameters are: number of output channels c, convolution kernel 1, step size 1, and padding value 0) to obtain a feature map with a dimension of h×w×c as the final output of the SAP submodule.
5. The traffic target detection method integrating multi-scale features according to claim 4 is characterized in that: The TSF module performs feature fusion on the feature maps output by the 3rd, 5th, and 7th layers. This process specifically includes the following steps: (b1) Perform an ADown operation on the feature map F1 (whose dimension is h1×w1×dim1) output by the third layer to obtain a feature map with a dimension of h m × m ×0.5dim m feature map. (b2) Feature map F of the 5th layer output m (The dimension is h m × m ×dim m ) performs CBS operation to obtain the dimension h m × m ×0.5dim m feature map. (b3) Feature map F output from layer 7 h (The dimension is h h × h ×dim h ) performs an upsampling operation and performs a CBS operation on the upsampled feature map to obtain a feature map with a dimension of h m × m ×0.5dim m feature map. (b4) Concatenate the feature maps obtained in steps (b1), (b2), and (b3) along the channel direction to obtain a feature map with a dimension of h m × m ×1.5dim m feature map. (b5) Perform a depth-wise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map of dimension h m × m ×1.5dim m feature map. (b6) Perform a depth-wise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map of dimension h m × m ×1.5dim m feature map. (b7) Perform a depth-wise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map of dimension h m × m ×1.5dim m feature map. (b8) Perform a depth-wise separable convolution operation on the feature map obtained in step (b4) to obtain a feature map of dimension h m × m ×1.5dim m feature map. (b9) Add the feature maps obtained in steps (b4), (b5), (b6), (b7) and (b8) element by element to obtain a feature map with a dimension of h. m × m ×1.5dim m feature map. (b10) Perform CBS operation on the feature map obtained in step (b9) to obtain a feature map with dimension h m ×W m ×1.5dim m feature map. (b1 1) Add the feature maps obtained in step (b4) and step (b10) element by element to obtain a feature map with dimension h m × m ×1.5dim m The feature map is used as the final output of the TFS sub-module.
6. The method for traffic target detection by integrating multi-scale features according to claim 5, characterized in that: The Detect operation divides the input feature map into two branches for forward propagation: Branch 1: The feature map passes through 2 CBS operations and 1 convolution operation in sequence. The parameters of the first 2 CBS operations are: convolution kernel size 3, step size 1, padding value 1, and the number of output channels is c t ; The subsequent convolution operation parameters are: convolution kernel size 1, step size 1, padding value 1, and the number of output channels is c t . This branch is used to calculate the bounding box regression loss of the target. Branch 2: The feature map passes through 2 CBS operations and 1 convolution operation in sequence. The parameters of the first 2 CBS operations are: convolution kernel size 3, step size 1, padding value 1, output channel number 4c t ; The subsequent convolution operation parameters are: convolution kernel size 1, step size 1, padding value 1, output channel number 4c t . This branch is used to calculate the classification loss of the target.
7. The method for traffic target detection by integrating multi-scale features according to claim 6, characterized in that: The traffic target detection model integrating multi-scale features is trained through the following steps: (4-1) Download the mixed dataset consisting of the open source BDD100K dataset and the KITTI dataset, and divide the mixed dataset into a training set and a test set in a ratio of 8:
2. (4-2) Perform data preprocessing on the training set obtained in step (4-1) to obtain a preprocessed training set. (4-3) Performing image enhancement processing on the preprocessed training set obtained in step (4-2) to obtain an enhanced training set. (4-4) For each sample (whose dimension is 640×640×3) in the training set after enhancement processing obtained in step (4-3), the sample is input into the first layer of the traffic target detection model for processing to output a feature map corresponding to the sample with a dimension of 320×320×64. (4-5) For each sample in the training set after the enhancement processing obtained in step (4-3), the feature map with a dimension of 320×320×64 corresponding to the sample obtained in step (4-4) is input into the second layer of the traffic target detection model for processing to output a feature map with a dimension of 160×160×128 corresponding to the sample. (4-6) For each sample in the training set after the enhancement processing obtained in step (4-3), the feature map with a dimension of 160×160×128 corresponding to the sample obtained in step (4-5) is input into the third layer of the traffic target detection model for multi-scale feature extraction to output a feature map with a dimension of 160×160×128 corresponding to the sample. (4-7) For each sample in the training set after the enhancement processing obtained in step (4-3), the feature map with a dimension of 160×160×128 corresponding to the sample obtained in step (4-6) is input into the 4th layer of the traffic target detection model for downsampling to output a feature map with a dimension of 80×80×256 corresponding to the sample. (4-8) For each sample in the training set after the enhancement processing obtained in step (4-3), the feature map with a dimension of 80×80×256 corresponding to the sample obtained in step (4-7) is input into the 5th layer of the traffic target detection model for multi-scale feature extraction to output a feature map with a dimension of 80×80×256 corresponding to the sample. (4-9) For each sample in the training set after the enhancement processing obtained in step (4-3), the feature map with a dimension of 80×80×256 corresponding to the sample obtained in step (4-8) is input into the 6th layer of the traffic target detection model for downsampling to output a feature map with a dimension of 40×40×512 corresponding to the sample. (4-10) For each sample in the training set after the enhancement processing obtained in step (4-3), the feature map with a dimension of 40×40×512 corresponding to the sample obtained in step (4-9) is input into the 7th layer of the traffic target detection model for multi-scale feature extraction to output a feature map with a dimension of 40×40×512 corresponding to the sample. (4-11) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map with a dimension of 160×160×128 corresponding to the sample obtained in step (4-6), the feature map with a dimension of 80×80×256 corresponding to the sample obtained in (4-8), and the feature map with a dimension of 40×40×512 corresponding to the sample obtained in (4-10) are input into the 8th layer of the traffic target detection model for three-way semantic fusion to output a feature map with a dimension of 80×80×384 corresponding to the sample. (4-12) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map with a dimension of 80×80×384 corresponding to the sample obtained in step (4-11) is input into the 9th layer of the traffic target detection model for processing to output a feature map with a dimension of 40×40×256 corresponding to the sample. (4-13) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map with a dimension of 40×40×512 corresponding to the sample obtained in step (4-10) and each feature map with a dimension of 40×40×256 obtained in step (4-12) are input into the 10th layer of the traffic target detection model for processing to output a feature map with a dimension of 40×40×768 corresponding to the sample. (4-14) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map with a dimension of 40×40×768 corresponding to the sample obtained in step (4-13) is input into the 11th layer of the traffic target detection model for processing to output a feature map with a dimension of 40×40×512 corresponding to the sample. (4-15) For each sample in the training set after the enhancement processing obtained in step (4-3), the feature map with a dimension of 80× 80× 384 corresponding to the sample obtained in step (4-11) is input into the 12th layer of the traffic target detection model for upsampling to output a feature map with a dimension of 160×160×384 corresponding to the sample. (4-16) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map with a dimension of 160×160×128 corresponding to the sample obtained in step (4-6) and each feature map with a dimension of 160×160×384 obtained in step (4-15) are input into the 13th layer of the traffic target detection model for processing to output a feature map with a dimension of 160×160×512 corresponding to the sample. (4-17) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map with a dimension of 160×160×512 corresponding to the sample obtained in step (4-16) is input into the 14th layer of the traffic target detection model for processing to output a feature map with a dimension of 160×160×256 corresponding to the sample. (4-18) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map with a dimension of 80×80×384 corresponding to the sample obtained in step (4-11), each feature map with a dimension of 40×40×512 obtained in step (4-14), and each feature map with a dimension of 160×160×256 obtained in step (4-17) are input into the 15th layer of the traffic target detection model for three-way semantic fusion to output a feature map with a dimension of 80×80×576 corresponding to the sample. (4-19) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 80×80×576 corresponding to the sample obtained in step (4-18) is input into the 16th layer of the traffic target detection model for processing to output a feature map of dimension 40×40×256 corresponding to the sample. (4-20) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 40×40×256 corresponding to the sample obtained in step (4-12), the feature map of size 40×40×512 corresponding to the sample obtained in step (4-14), and the feature map of size 40×40×256 corresponding to the sample obtained in step (4-19) are input into the 17th layer of the traffic target detection model for processing to output a feature map of dimension 40×40×1024 corresponding to the sample. (4-21) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 40×40×1024 corresponding to the sample obtained in step (4-20) is input into the 18th layer of the traffic target detection model for processing to output a feature map of dimension 40×40×512 corresponding to the sample. (4-22) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 80×80×576 corresponding to the sample obtained in step (4-18) is input into the 19th layer of the traffic target detection model for upsampling to output a feature map of dimension 160×160×576 corresponding to the sample. (4-23) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 160×160×384 corresponding to the sample obtained in step (4-15), the feature map of size 160×160×256 corresponding to the sample obtained in step (4-17), and the feature map of size 160×160×576 corresponding to the sample obtained in step (4-22) are input into the 20th layer of the traffic target detection model for processing to output a feature map of dimension 160×160×1216 corresponding to the sample. (4-24) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 160×160×1216 corresponding to the sample obtained in step (4-23) is input into the 21st layer of the traffic target detection model for processing to output a feature map of dimension 160×160×256 corresponding to the sample. (4-25) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 160×160×128 corresponding to the sample obtained in step (4-6) is input into the 22nd layer of the traffic target detection model for processing to output a feature map of dimension 160×160×256 corresponding to the sample. (4-26) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 80×80×256 corresponding to the sample obtained in step (4-8) is input into the 23rd layer of the traffic target detection model for processing to output a feature map of dimension 80×80×576 corresponding to the sample. (4-27) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 160×160×256 corresponding to the sample obtained in step (4-24) and the feature map of size 160×160×256 corresponding to the sample obtained in step (4-25) are input into the 24th layer of the traffic target detection model for processing to output a feature map of size 160×160×256 corresponding to the sample. (4-28) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 80×80×576 corresponding to the sample obtained in step (4-18) and the feature map of size 80×80×576 corresponding to the sample obtained in step (4-26) are input into the 25th layer of the traffic target detection model for processing to output a feature map of dimension 80×80×576 corresponding to the sample. (4-29) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 40×40×512 corresponding to the sample obtained in step (4-10) and the feature map of size 40×40×512 corresponding to the sample obtained in step (4-21) are input into the 26th layer of the traffic target detection model for processing to output a feature map of dimension 40×40×512 corresponding to the sample. (4-30) For each sample in the training set after the enhanced processing obtained in step (4-3), the feature map of size 160×160×256 corresponding to the sample obtained in step (4-27), the feature map of size 80×80×576 corresponding to the sample obtained in step (4-28), and the feature map of size 40×40×512 corresponding to the sample obtained in step (4-29) are input into the 27th layer of the traffic target detection model for processing to obtain the bounding box regression loss and classification loss corresponding to the sample. (4-31) For each sample in the training set after the enhancement processing obtained in step (4-3), the total loss L = αL is calculated based on the bounding box regression loss and classification loss corresponding to the sample obtained in step (4-30) CLS +βL DFL +γL WIoUv2 , where α`β and γ represent loss weights. (4-32) For each sample in the enhanced training set obtained in step (4-3), the traffic target detection model integrating multi-scale features is iteratively trained according to the total loss corresponding to the sample obtained in step (4-31) and using gradient descent until the traffic target detection model integrating multi-scale features reaches a preset number of iterations, and the optimal parameters of the traffic target detection model integrating multi-scale features at this time are obtained, thereby obtaining a preliminarily trained traffic target detection model integrating multi-scale features. (4-33) The test set obtained in step (4-1) is used to test the traffic target detection model integrating multi-scale features that was initially trained in step (4-32) until the detection accuracy is optimal, thereby obtaining the final trained traffic target detection model integrating multi-scale features.
8. The method for traffic target detection by integrating multi-scale features according to claim 7, characterized in that: Step (4-3) randomly selects one of the following nine data enhancement methods or any combination of them for processing: image HSV enhancement, including hue, saturation and brightness enhancement, and the enhancement factors are set to 0.015, 0.7 and 0.4 respectively; Image translation, the translation factor is 0.1; For image scaling, the scaling factor is 0.5; for image flipping, the flipping factor is 0.5; for multi-image stitching, the stitching factor is 1.0; for image erasing, the erasing factor is 0.4; for image cropping, the cropping factor is 1.
0.
9. The method for traffic target detection by integrating multi-scale features according to claim 8, characterized in that: Classification loss L CLS Equal to the classification value predicted for this sample and the true classification value y cls The binary cross entropy loss between is calculated as:
10. A traffic target detection system integrating multi-scale features, characterized in that: include: The first module is used to obtain the traffic image to be detected. The second module is used to perform data preprocessing on the traffic image to be detected acquired by the first module to obtain a preprocessed image. The third module is used to input the preprocessed image obtained by the second module into a pre-trained traffic target detection model integrating multi-scale features to obtain the final detection result.
Citation Information
Cited By
A traffic scene visual analysis system
CN122737876A