Traffic target detection algorithm based on multi-scale feature extraction and fusion

Through the traffic object detection algorithm that extracts and integrates multi-scale feature, the error detection and missed detection problems of small target detection in traffic scenarios are solved, the detection accuracy and recall of the model are improved, and the robustness of the model is enhanced.

CN120259989APending Publication Date: 2025-07-04CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510314272.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing traffic scene object detection models are prone to error detection and missed detection problems when detecting small targets, especially in complex environments, it is difficult to accurately identify small objects.

Method used

The traffic object detection algorithm based on multi-scale feature extraction and fusion is adopted, including feature extraction backbone network, cascade channel attention pyramid pooling module, path aggregation network, cross-layer feature collection module and multi-branch attention module. Through the fusion and interaction of multi-scale feature information, the model's perception ability of different scales and hierarchical features is improved.

Benefits of technology

The recall and detection accuracy of the model are improved, the utilization rate of global feature information is enhanced, and the robustness and accuracy of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259989A_ABST
    Figure CN120259989A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection, in particular to a traffic target detection algorithm based on multi-scale feature extraction and fusion, and the algorithm comprises the steps: obtaining an original traffic scene image and a corresponding label, and obtaining a traffic scene image data set; the traffic scene target detection model is trained by using a traffic scene image data set; and inputting traffic scene images in the traffic scene image data set into the traffic scene target detection model, and processing the traffic scene images by the traffic scene target detection model to obtain a detection result graph. According to the method, multi-scale semantic information is captured through the cascade channel attention pyramid pooling module, original image feature information is enhanced at the same time, the model recall rate is improved, and the detection accuracy is improved. A backbone network cross-layer feature collection module, a path aggregation network cross-layer feature collection module and a cross-layer feature fusion module are provided, the feature expression ability of the model is enhanced, a multi-branch attention module is provided, attention weights under different scales are used for improving the attention degree of the model on a detection target, and the overall detection precision of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a traffic object detection algorithm based on multi-scale feature extraction and fusion. Background Art

[0002] With the continuous progress of deep learning and computer vision technologies, the application of object detection in the traffic field has been further deepened, promoting the intelligent and automated development of the traffic system and making important contributions to achieving a safer, more efficient and environmentally friendly traffic environment. Object detection is one of the core technologies of traffic safety management. By real-time detecting and identifying objects in traffic scenes (such as pedestrians, other vehicles, traffic signs, traffic lights, etc.), potential dangers can be discovered in time and responses can be made, thus reducing the occurrence of traffic accidents. In autonomous driving technology, object detection is a core component of vehicle environment perception. Autonomous vehicles obtain real-time surrounding environment data through sensors such as cameras, lidar (LiDAR), millimeter-wave radars, etc., and perform object detection through deep learning models to identify dynamic and static objects in the environment. In intelligent traffic management systems, object detection can provide decision-making support for traffic scheduling through the analysis of real-time traffic data. For example, cameras and other sensors can be used to monitor traffic flow, detect congestion and traffic accidents. This data can help the traffic management center adjust traffic lights in real time, control vehicle flow, optimize lane allocation, etc., thereby improving road usage efficiency and reducing traffic congestion.

[0003] The background in traffic scenes usually contains various interference elements, such as buildings, road signs, billboards, trees, etc., which are likely to interfere with the accuracy of object detection models, especially in complex environments such as urban areas or intersections. The objects in traffic scenes (such as pedestrians, vehicles, cyclists, etc.) are often partially or completely blocked by other objects. The objects in traffic scenes may have large scale differences. For example, vehicles or pedestrians in the distance may be very small, while those nearby are larger. Due to the size differences of different objects, object detection algorithms need to have strong multi-scale detection capabilities to simultaneously detect objects of different sizes at different distances. In traffic scenes, especially in complex urban environments, small objects (such as pedestrians, cyclists, low obstacles, etc.) are often difficult to be accurately identified. Especially in crowded scenes, the contrast between the target object and the background is low, making target recognition more difficult. These difficulties make the object detection task in traffic scenes a challenging task. The existing object detection models in traffic scenes are mainly divided into single-stage object detectors, two-stage object detectors, and Transformer-based object detectors, which have problems of misdetection and missed detection when detecting small targets and still have room for improvement. Summary of the Invention

[0004] The object of the present invention is to provide a traffic target detection algorithm based on multi-scale feature extraction and fusion, aiming to solve the problems of misdetection and missed detection of small targets in the existing traffic scene target detection models.

[0005] To achieve the above object, the present invention provides a traffic target detection algorithm based on multi-scale feature extraction and fusion, including the following steps:

[0006] Obtain the original traffic scene image and the corresponding label to obtain a traffic scene image dataset;

[0007] The traffic scene target detection model is trained using the traffic scene image dataset;

[0008] Input the traffic scene images in the traffic scene image dataset into the traffic scene target detection model, and after being processed by the traffic scene target detection model, a detection result map is obtained.

[0009] Among them, the traffic scene target detection model includes a feature extraction backbone network, a cascaded channel attention pyramid pooling module, a path aggregation network, a backbone network cross-layer feature collection module, a path aggregation network cross-layer feature collection module, a cross-layer feature fusion module, and a multi-branch attention module.

[0010] Among them, the specific manner in which the traffic scene target detection model is trained using the traffic scene image dataset:

[0011] Input the original traffic scene image into the feature extraction backbone network for feature extraction to obtain four features with different resolution sizes, and input the features with different resolution sizes into the backbone network cross-layer feature collection module to obtain backbone network cross-layer features;

[0012] Input the feature with the smallest resolution into the cascaded channel attention pyramid pooling module to obtain cascaded features;

[0013] Input the cascaded features into the path aggregation network, and in the top-down feature extraction process of the path aggregation network, fuse the backbone network cross-layer features with the local features obtained through convolution and upsampling during the feature extraction process;

[0014] Input the remaining three features with different resolution sizes in the top-down feature extraction process of the path aggregation network into the path aggregation network cross-layer feature collection module to obtain path aggregation network cross-layer features;

[0015] In the bottom-up feature extraction process of the path aggregation network, fuse the path aggregation network cross-layer features with the local features obtained through downsampling and convolution operations during the feature extraction process;

[0016] The three obtained output feature maps are passed through the multi-branch attention module and then fed into the detection head to complete the training of the traffic scene target detection model.

[0017] Among them, in the cascaded channel attention pyramid pooling module, the input feature map is first passed through channel attention, and then during the pooling process, the feature map passed through channel attention is used as supplementary features to establish a hierarchical association path between pooling layers to retain detailed information in the feature map.

[0018] Among them, in the path aggregation network, the backbone network cross-layer feature collection module and the path aggregation network cross-layer feature collection module are added. The backbone network cross-layer feature collection module and the path aggregation network cross-layer feature collection module respectively collect cross-layer feature information, reduce the resolution by the method of the maximum pooling layer, improve the resolution by the linear interpolation method to align feature maps of different sizes, and extract cross-layer feature information by constructing a convolutional structure.

[0019] A traffic target detection algorithm based on multi-scale feature extraction and fusion of the present invention obtains an original traffic scene image and a corresponding label to obtain a traffic scene image dataset; a traffic scene target detection model is trained using the traffic scene image dataset; the traffic scene image in the traffic scene image dataset is input into the traffic scene target detection model, and after being processed by the traffic scene target detection model, a detection result map is obtained. This method uses ResNet-50 as the backbone network of the model to extract multi-scale feature information with strong robustness, including features F1-F4 of four resolution sizes. Through the CCAPP (Cascading Channel Attention Pyramid Pooling) cascading channel attention pyramid pooling module, it makes up for the deficiency of losing the original feature information when the spatial pyramid pooling result extracts multi-scale features; through the BCFC (Backbone Cross Feature Collection) backbone network cross-layer feature collection module, it collects the global features of the backbone network; through the PANCFC (PAN Cross Feature Collection) path aggregation network cross-layer feature collection module, it collects the global features of the path aggregation network in the PANet network. Through the CFF (Cross Feature Fusion) cross-layer feature fusion module, it realizes the interaction between cross-layer feature information and local feature information, aiming at the problems of losing cross-layer information when the lateral connection of the path aggregation network fuses the backbone network information and losing feature information during the feature extraction process. Aiming at the problems of misdetection and missed detection of small targets in the existing traffic scene target detection models, the model proposed by the present invention improves the model's perception ability of different scale and hierarchical features, effectively improves the recall rate of the model, improves the utilization rate of global feature information, and enhances the robustness and accuracy of the target detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 It is a schematic structural diagram of the target detection model in the traffic scene of the present invention.

[0022] Figure 2 It is a schematic structural diagram of the progressive pyramid pooling module of the present invention.

[0023] Figure 3 It is a schematic structural diagram of the global feature information extraction module of the present invention.

[0024] Figure 4 It is a schematic structural diagram of the global feature information fusion module of the present invention.

[0025] Figure 5 It is a schematic structural diagram of the multi-branch attention module of the present invention.

[0026] Figure 6 It is a flowchart of a traffic target detection algorithm based on multi-scale feature extraction and fusion provided by the present invention.

[0027] Figure 7 It is a flowchart of the specific manner in which the traffic scene target detection model is trained using the traffic scene image dataset. Specific implementation manner

[0028] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0029] Please refer to Figures 1 to 7 , in the first aspect, the present invention provides a traffic target detection algorithm based on multi-scale feature extraction and fusion, including the following steps:

[0030] S1 Obtain the original traffic scene image and the corresponding label to obtain a traffic scene image dataset;

[0031] In the embodiments of the present invention, the training dataset used is BDD100K. When used for object detection tasks, it mainly provides object annotations in various driving scenarios to help researchers and developers train and evaluate visual perception models in autonomous driving systems. In the BDD100K dataset for object detection tasks, the accurate positions of the following types of objects are mainly annotated: vehicles, pedestrians, cyclists, traffic signs, and obstacles. The BDD100K dataset contains more than 100,000 image frames. The training set, validation set, and test set of BDD100K are divided into 70,000 images, 10,000 images, and 20,000 images respectively, providing rich data to help train and evaluate object detection tasks. The training set is used to train the model and contains a large number of images and annotations, covering different driving environments and scenarios, such as urban streets, highways, rural roads, etc. The validation set is used to adjust the hyperparameters of the model (such as learning rate, network structure, etc.) and conduct mid-term evaluations. The scale of the validation set is small, but it is sufficient to reflect the generalization ability of the model. The test set: is used to finally evaluate the performance of the model. The test set is usually used to report the final accuracy, recall rate, mAP and other metrics of the model. For object detection, each image frame in the training and test sets may contain multiple object instances, ensuring that the model can handle complex multi-object scenarios.

[0032] The traffic scene object detection model S2 is trained using the traffic scene image dataset;

[0033] In the embodiments of the present invention, the traffic scene object detection model includes a feature extraction backbone network, a CCAPP (Cascading Channel Attention Pyramid Pooling) cascading channel attention pyramid pooling module, a PANet path aggregation network, a BCFC (Backbone Cross Feature Collection) backbone network cross-layer feature collection module, a PANCFC (PAN Cross Feature Collection) path aggregation network cross-layer feature collection module, a CFF (Cross Feature Fusion) cross-layer feature fusion module, and a multi-branch attention module.

[0034] Specific method:

[0035] S21 Input the original traffic scene image into the feature extraction backbone network for feature extraction to obtain four features with different resolution sizes, and input the features with different resolution sizes into the backbone network cross-layer feature collection module to obtain backbone network cross-layer features;

[0036] In the embodiment of the present invention, the original traffic scene image is input into the feature extraction backbone network for feature extraction to obtain features F1 - F4 with different resolution sizes; and the features F1 - F4 with four resolution sizes in the feature extraction backbone network are input into the BCFC (Backbone Cross Feature Collection) backbone network cross - layer feature collection module to obtain the backbone network cross - layer feature B G Specifically, ResNet - 50 is used as the core network architecture. In the object detection task, the feature extraction backbone network is the core part responsible for extracting features from the input image. It is stacked by a convolutional neural network (CNN) and is used to extract low - level and high - level feature information of the image. These features will serve as the basis for the subsequent detection network to perform object localization and classification. After an image is input into the backbone network of the model, a series of multi - scale features with different resolution sizes can be obtained. These features decrease in resolution size and increase in feature dimension step by step. For a feature map with a larger resolution, each feature point (pixel) contains more detailed spatial information and can describe low - level information such as details, edges, and textures in the image. A feature map with a smaller resolution has a smaller spatial size, and each feature point represents a larger area in the image and contains more high - level semantic information, such as the overall shape, category, and background of the object. Therefore, the four features with different resolution sizes are all passed into the BCFC (Backbone Cross Feature Collection) backbone network cross - layer feature collection module. Please refer to Figure 3 , in this embodiment, after the image is extracted by the backbone features, four features with resolution sizes from high to low are obtained, which are respectively marked as F1, F2, F3, and F4. F1, F2, F3, and F4 are all passed into the BCFC (Backbone Cross Feature Collection) backbone network cross - layer feature collection module. F1 and F2 reduce the resolution by the method of max - pooling layer, and F4 increases the resolution by the linear interpolation method. After directly adding the four aligned feature maps and passing through a convolutional structure, the output B of the BCFC (Backbone CrossFeature Collection) backbone network cross - layer feature collection module is obtained G .

[0037] S22 Input the feature with the smallest resolution into the cascaded channel attention pyramid pooling module to obtain a cascaded feature;

[0038] In the embodiment of the present invention, the feature F4 with the smallest resolution is input into the CCAPP module to obtain the feature S. Specifically, as Figure 2 shown, the output of the feature extraction backbone network first passes through global average pooling to obtain X′, and then passes through a fully - connected layer and an activation function to obtain the channel attention weight Weightc , which is weighted with the original input to obtain the output X″ of the channel attention branch. At the same time, the original input passes through a convolutional layer to obtain the input X of the CCAPP module, and then passes through three pooling layers with a pooling window size of 5×5 in sequence. The input of each pooling layer is the result F obtained by adding the pooling result of the previous layer and the original input image i-1 +X″ (the input of the first pooling layer is only X). The output F of each layer of pooling obtained thereby i corresponds to three different receptive field sizes of 5×5, 9×9, and 13×13. The outputs F1 - F3 of the three different receptive field sizes after pooling are concatenated with the original input X through the Concat function, and then pass through convolution, normalization, and activation functions in sequence to obtain the output S of the CCAPP module.

[0039] S23 inputs the concatenated features into the path aggregation network, and fuses the cross-layer features of the backbone network with the local features obtained by convolution and upsampling during the top-down feature extraction process of the path aggregation network;

[0040] In the embodiment of the present invention, the feature S is input into the PANet network; during the top-down feature extraction process of the PANet network, B G is fused with the local features obtained by convolution and upsampling during the feature extraction process. Specifically, please refer to Figure 4 , and the global feature B of the feature extraction backbone network G is used as an input G of the CFF (Cross Feature Fusion) cross-layer feature fusion module feature , and the intermediate layer features in the PANet network are used as another input L feature . After G feature passes through convolution, one branch is upsampled by the Bilinear linear interpolation method as the result G of the global feature information feature ′, and after G feature passes through convolution, it is input into another branch for convolution and upsampling operations by the Bilinear linear interpolation method as the weight weight of the global feature information G and L feature are dot-multiplied to obtain the result L feature ′ of the local feature information weighted by the global feature. Finally, the global feature information G feature ′ and the local feature information L feature ′ are added to obtain the final output.

[0041] S24 inputs the remaining three features of different resolution sizes during the top-down feature extraction process of the path aggregation network into the cross-layer feature collection module of the path aggregation network to obtain the cross-layer features of the path aggregation network;

[0042] In the embodiment of the present invention, the features P1 - P3 with three resolution sizes in the top - down feature extraction process of the PANet network are input into the cross - layer feature collection module of the PANCFC (PAN Cross Feature Collection) path aggregation network to obtain the global feature P of the path aggregation network G , specifically, please refer to Figure 3 , in the cross - layer feature collection module of the PANCFC (PAN Cross Feature Collection) path aggregation network, the three feature maps P1 - P3 in the top - down direction of the PANet network are aligned. P3 is upsampled by linear interpolation method, P1 is downsampled by max - pooling, and the three aligned feature maps are directly added and then passed through a convolutional structure to obtain the global feature P of the path aggregation network G .

[0043] S25 In the bottom - up feature extraction process of the path aggregation network, the cross - layer feature of the path aggregation network is fused with the local feature obtained through downsampling and convolution operations in the feature extraction process;

[0044] In the embodiment of the present invention, in the bottom - up feature extraction process of the PANet network, P G is fused with the local feature obtained through downsampling and convolution operations in the feature extraction process. Specifically, please refer to Figure 4 , in the CFF (Cross Feature Fusion) cross - layer feature fusion module, the global feature P of the path aggregation network G is used as an input G feature of the module, and the intermediate layer feature in the PANet network is used as another input L feature . After G feature is convolved, one branch is upsampled by the Bilinear linear interpolation method as the result G feature ' of the global feature information. After G feature is convolved, it is passed into another branch for convolution and upsampling by the Bilinear linear interpolation method as the weight weight G of the cross - layer feature information. It is dot - multiplied with L feature to obtain the result L feature ' of the local feature information weighted by the cross - layer feature. Finally, the cross - layer feature information G feature ' and the local feature information L feature ' are added to obtain the final output.

[0045] S26 Pass the obtained three output feature maps through the multi-branch attention module and then input them into the detection head to complete the training of the traffic scene target detection model.

[0046] In the embodiment of the present invention, the obtained three output feature maps P1, P2′, and P3′ are passed through the multi-branch attention module and then input into the detection head to complete the training of the traffic scene target detection model. Specifically, please refer to Figure 5 , The multi-branch attention module takes P1, P2′, and P3′ as inputs respectively. The input X enters three different branches. One branch passes through a 3×3 convolution to obtain X 3×3 , Another branch passes through a 1×1 convolution to obtain X 1×1 , By processing (average pooling) the spatial dimensions (width and height) of X 1×1 , a weight map X of the width and height of the feature map is obtained h and X w 、Take X h 、X w After passing through the Sigmoid activation function, Concat is performed to obtain the attention weight weight 3×3 , Then the original input X 3×3 is weighted. Take X 1×1 、X 3×3 After average pooling respectively, pass through the Softmax function to obtain X 1×1 ′、X 3×3 ′、Multiply X 1×1 with X 3×3 ′、Multiply X 3×3 with X 1×1 ′ and add the results, and finally obtain the attention weight weight through the Sigmoid function att . The last branch takes P1, P2′, and P3′ as inputs. After global average pooling and a fully connected layer, pass through the Sigmoid function to obtain the channel weight, and then weight it with P1, P2′, and P3′ to obtain the channel branch output. The channel branch output is weighted with the attention weight weight att to obtain the final result.

[0047] This model is implemented using the Pytorch framework and trained on a Tesla V100 GPU with 32GB of video memory. Our model requires 100 epochs to complete training, and the batch size for each training is set to 16. We use the Adam optimizer and set the initial learning rate to 1e-4. At the same time, the poly learning rate is adopted as the strategy to adjust the learning rate during training. In the training stage and the testing stage, the size of all input images is adjusted to 640×640. The data augmentation strategy adopted is random rotation, vertical flipping, and horizontal flipping.

[0048] S3 inputs the traffic scene images in the traffic scene image dataset into the traffic scene target detection model, and after being processed by the traffic scene target detection model, a detection result map is obtained.

[0049] In a second aspect, the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the traffic target detection algorithm based on multi-scale feature extraction and fusion as described above.

[0050] In a third aspect, the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores the computer program, and when the computer program is executed by the processor, it causes the processor to execute the traffic target detection algorithm based on multi-scale feature extraction and fusion as described above.

[0051] What is disclosed above is only a preferred embodiment of the traffic target detection algorithm based on multi-scale feature extraction and fusion of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A traffic target detection algorithm based on multi-scale feature extraction and fusion, characterized in that, It includes the following steps: Obtain the original traffic scene images and corresponding labels to obtain a traffic scene image dataset; The traffic scene target detection model is trained using the traffic scene image dataset; Input the traffic scene images in the traffic scene image dataset into the traffic scene target detection model, and after being processed by the traffic scene target detection model, obtain a detection result map.

2. The traffic target detection algorithm based on multi-scale feature extraction and fusion according to claim 1, wherein; The traffic scene target detection model includes a feature extraction backbone network, a cascaded channel attention pyramid pooling module, a path aggregation network, a backbone network cross-layer feature collection module, a path aggregation network cross-layer feature collection module, a cross-layer feature fusion module, and a multi-branch attention module.

3. The traffic target detection algorithm based on multi-scale feature extraction and fusion according to claim 2, wherein; The specific way that the traffic scene target detection model is trained using the traffic scene image dataset: Input the original traffic scene image into the feature extraction backbone network for feature extraction to obtain four features with different resolution sizes, and input the features with different resolution sizes into the backbone network cross-layer feature collection module to obtain backbone network cross-layer features; Input the feature with the smallest resolution into the cascaded channel attention pyramid pooling module to obtain cascaded features; Input the cascaded features into the path aggregation network, and during the top-down feature extraction process of the path aggregation network, fuse the backbone network cross-layer features with the local features obtained through convolution and upsampling during the feature extraction process; Input the remaining three features with different resolution sizes during the top-down feature extraction process of the path aggregation network into the path aggregation network cross-layer feature collection module to obtain path aggregation network cross-layer features; During the bottom-up feature extraction process of the path aggregation network, fuse the path aggregation network cross-layer features with the local features obtained through downsampling and convolution operations during the feature extraction process; Pass the three obtained output feature maps through the multi-branch attention module and then input them into the detection head to complete the training of the traffic scene target detection model.

4. The traffic target detection algorithm based on multi-scale feature extraction and fusion according to claim 1, wherein; In the cascaded channel attention pyramid pooling module, first pass the input feature map through channel attention, and then use the feature map passed through channel attention as a supplementary feature during the pooling process to establish a hierarchical association path between pooling layers to retain the detailed information in the feature map.

5. The traffic target detection algorithm based on multi-scale feature extraction and fusion according to claim 1, wherein; In the path aggregation network, the backbone network cross-layer feature collection module and the path aggregation network cross-layer feature collection module are added. The backbone network cross-layer feature collection module and the path aggregation network cross-layer feature collection module respectively collect cross-layer feature information, reduce the resolution by the method of the max pooling layer, enhance the resolution by the bilinear interpolation method to align feature maps of different sizes, and extract cross-layer feature information by constructing a convolutional structure.

Citation Information

Cited By

  • Traffic scene small target detection method, device, equipment and medium

    CN121305050A