Traffic scene small target detection method based on adaptive spatial aggregation pyramid

Through adaptive spatial aggregation pyramid and multi-scale aggregation attention, combined with the small target loss function NE_IoU, the problems of weak semantic information and low positioning accuracy in small target detection are solved, and efficient small target detection effect is achieved.

CN120599223APending Publication Date: 2025-09-05CHONGQING JIAOTONG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510738766.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing deep learning small target detection algorithms have problems in complex traffic scenarios, such as weak semantic information of small targets, insufficient fusion, and low positioning accuracy, which leads to missed detections and false detections, especially poor detection effects under long-distance and low-resolution conditions.

Method used

The adaptive spatial aggregation pyramid (SAFPN) is combined with multi-scale aggregation attention (MAA) and small target loss function NE_IoU. The dynamic convolution TVConv is used to enhance feature fusion, construct a decoupling head structure, and adopt the Soft-NMS strategy to optimize feature representation ability and positioning accuracy.

Benefits of technology

It significantly improves the feature representation capability and positioning accuracy of small target detection, reduces missed detections and false detections, meets real-time detection requirements, and has better detection accuracy than other methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599223A_ABST
    Figure CN120599223A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic scene small target detection method based on an adaptive spatial aggregation pyramid, and belongs to the field of computer vision and small target detection, and the method comprises the steps: designing a multi-scale aggregation attention mechanism to enhance the texture, shape and context information of a small target for the problem of weak semantic information of the small target; aiming at the problem of insufficient multi-scale fusion of small targets, designing an adaptive space aggregation pyramid depth fusion multi-size feature map; a small target loss function NEIoU is constructed to solve the problem that small target detection is sensitive to position deviation, and the convergence speed is increased; and meanwhile, a Soft-NMS strategy is adopted to alleviate the problem that small targets are mistakenly suppressed due to IoU calculation deviation or dense arrangement. According to the traffic scene small target detection method based on the adaptive spatial aggregation pyramid, the precision of small target detection is remarkably improved by improving semantic richness, multi-scale fusion and positioning precision; meanwhile, deployment is easy, the operation speed is high, and the requirement for real-time detection is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and small target detection, and in particular to a small target detection method in traffic scenes based on an adaptive spatial aggregation pyramid. Background Art

[0002] In recent years, with the rapid development of computer vision and its interdisciplinary applications, small object detection has become a hot research topic in the field of object detection due to its critical role in complex traffic scenarios. Current small object detection algorithms are primarily categorized into two main groups: traditional handcrafted feature methods and deep learning-based detection methods. Traditional methods are limited by the expressive power of manually designed features and exhibit poor robustness in complex backgrounds and scenarios with multi-scale variations. In contrast, deep learning-based methods, through end-to-end feature learning, significantly improve detection accuracy and real-time performance, and are gradually becoming mainstream.

[0003] Currently, mainstream deep learning algorithms for small target detection include the YOLO series, FCOS, VFNet, and DeformableDETR. These algorithms achieve high detection accuracy while maintaining low computational complexity by optimizing strategies such as multi-scale feature fusion and dynamic receptive field adjustment. They have been widely used in scenarios such as remote sensing image analysis, traffic sign detection, and video surveillance. However, in practical applications, small target detection (such as pedestrians at a distance or traffic signs at a relatively long distance) is often required at long distances and low resolutions, which places extremely high demands on the sensitivity and positioning accuracy of the detection algorithm. Although existing deep learning methods have made significant progress, they still face challenges in terms of weak semantic information of small targets, insufficient fusion, and low positioning accuracy. Summary of the Invention

[0004] The purpose of the present invention is to provide a small target detection method in traffic scenes based on an adaptive spatial aggregation pyramid. By designing multi-scale aggregate attention (MAA), adaptive spatial aggregation pyramid (SAFPN) and small target loss function NE_IoU, the characterization ability of small target features can be significantly improved; while maintaining real-time performance, it can effectively alleviate the problems of missed detection and false detection of small targets caused by weak semantic information, insufficient fusion and low positioning accuracy.

[0005] To achieve the above object, the present invention provides a method for detecting small targets in traffic scenes based on an adaptive spatial aggregation pyramid, comprising the following steps:

[0006] S1. Prepare the typical small object detection datasets TT100K and CityPersons, and convert the data annotation file format to the MS COCO data format;

[0007] S2. Use DarkNet-53 as the backbone network for small target detection, and output the backbone network C3, C4, and C5 feature layers to the adaptive spatial aggregation pyramid for feature fusion;

[0008] S3, design an adaptive spatial aggregation pyramid;

[0009] S3.1. Based on the asymptotic feature pyramid, dynamic convolution TVConv is added, that is, the traditional ConvsBNSILU structure is replaced by TVConvBNSILU structure;

[0010] S3.2. Construct an adaptive multi-scale aggregation attention module MAA and embed it into the adaptive spatial aggregation pyramid;

[0011] S4, construct a decoupled head structure detection head to separate the classification and detection heads;

[0012] S4.1, the classification loss uses binary cross entropy BCE, and the positioning loss function uses the designed small target loss function NE_IOU;

[0013] S4.2. Soft-NMS is used to alleviate the problem of small objects being incorrectly suppressed due to IoU calculation deviation or dense arrangement.

[0014] S5. Input the training set into the overall detection network for training and save the optimal model; input the test set into the optimal model saved during training for testing to verify the detection effect of the improved model.

[0015] Preferably, in S3.1, dynamic convolution is added to the asymptotic fusion, and the dynamic convolution TVConv enhances the adaptability of local features through layout-aware translation-variant convolution.

[0016] Preferably, in S3.2, the multi-scale aggregation attention module MAA first performs wavelet convolution WTConv and deconvolution TPConv on the input n feature maps F i After sampling to the same size, the fusion is performed; then the overall information of the fused feature map is enhanced through global average pooling; then the multi-scale aggregation attention module MAA considers the cross-channel interaction information of each channel and its k neighbors through one-dimensional convolution. Through experiments, k=3 is set to map the channel weight after one-dimensional convolution to the range of [0,1] through SILU activation function; finally, the n feature maps are processed by their respective weight coefficients α i Fuse it with the mapped feature channels to obtain the final output feature map F'. The following formula is the calculation formula for multi-scale aggregate attention MAA;

[0017]

[0018] Among them, F i represents the feature map of the i-th branch; n represents the number of feature maps through the multi-scale aggregation attention module MAA; F' represents the fused feature map; M represents the channel attention weight after channel attention; α j Indicates the weight corresponding to j branches; b j is a learnable three-element one-dimensional tensor with all initial values ​​1.

[0019] Preferably, in S4.1, the small target loss function NE_IoU is composed of the small target detection loss function NWD, the α_EIOU loss function proposed in the BANet paper, and the adaptive coefficient γ. The calculation formula of α_EIOU is as follows:

[0020]

[0021] Among them, w, h, w gt and h gt Represents the width and height of the predicted box and the real box respectively, c is the diagonal length of the minimum bounding box covering the two boxes, ρ 2α (b,b gt ) is the distance between the center points of the two boxes, ρ 2α (w,w gt ) is the distance between the two box widths, ρ 2α (h,h gt ) is the distance between the two box heights, α = 3.1;

[0022] The NWD loss function is introduced based on α_EIOU. NWD uses the normalized Wasserstein distance to reduce the sensitivity to position deviations of small objects and effectively improve the detection performance of small targets, but the convergence is slow. The calculation formula of the NDS loss function is as follows:

[0023]

[0024] Among them, the real bounding box A=(x A ,y A ,w A ,h A ), predicted bounding box B = (x B ,y B ,w B ,h B ), x, y, w, and h are the center coordinates and width and height of the box respectively; NA and NB are Gaussian distributions modeled by the true bounding box A and the predicted bounding box B; is a distance metric; c is a constant closely related to the dataset;

[0025] By setting the adaptive parameters, the combination ratio of NWD and α_EIOU is dynamically adjusted as the final bounding box regression loss function. The following formula is the formula expression of the adaptive coefficient:

[0026]

[0027] Among them, λ is the balance coefficient, taking λ = 0.35, w and h represent the height and width of the real sample frame respectively, w l With h l Represent the height and width of the large target sample, w l With h l Both are 96;

[0028] Therefore, the formula expression of the NE_IOU regression loss function is:

[0029]

[0030] Preferably, the traditional NMS in S4.2 directly deletes all candidate boxes whose IoU with the highest score box exceeds the threshold, while Soft-NMS dynamically reduces the score of the overlapping box according to the IoU value instead of directly setting it to zero, thereby retaining potential valid detection results.

[0031] Preferably, the experimental evaluation indicators used for quantitative evaluation in S5 are mean average precision (mAP), floating-point operations (FLOPs), and parameter quantity (Params); wherein, mAP@0.5:0.95 represents the average precision of IoU from 0.5 to 0.95, floating-point operations (FLOPs) are used to measure the complexity of the model, and parameter quantity (Params) represents the number of parameters of the model, which measures the demand for graphics card performance.

[0032] Therefore, the present invention adopts the above-mentioned small target detection method in traffic scenes based on adaptive spatial aggregation pyramid, which has the following advantages:

[0033] (1) The constructed adaptive spatial aggregation pyramid strengthens the semantic information of small targets, such as fuzzy features such as texture and shape, through progressive fusion (ASFPN), multi-scale aggregate attention (MAA) and variational convolution (TVConv), and optimizes multi-scale feature fusion. It adaptively combines shallow high-resolution features (preserving details) with deep semantic features to further improve the effectiveness of feature fusion and enhance the model detection performance.

[0034] (2) The constructed small target loss function NE_IoU adjusts the ratio of NWD loss to α_EIoU loss function through the adaptive parameter γ, solving the problem that small target detection is sensitive to position deviation and improving the convergence speed. At the same time, the Soft-NMS strategy is adopted to alleviate the problem that small targets are mistakenly suppressed due to IoU calculation deviation or dense arrangement, further improving the positioning accuracy of small target detection.

[0035] (3) Experimental results show that the method based on adaptive spatial aggregation pyramid performs well on typical small target detection datasets such as TT100K and CityPersons. Its detection accuracy is significantly better than other methods, and its parameter count is small, meeting the requirements of real-time detection.

[0036] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is an overall flow chart of an embodiment of a method for detecting small targets in traffic scenes based on an adaptive spatial aggregation pyramid according to the present invention;

[0038] Figure 2 2. It is a diagram of a small target detection network structure according to an embodiment of a method for detecting small targets in traffic scenes based on an adaptive spatial aggregation pyramid according to the present invention;

[0039] Figure 3 2. This is a network structure diagram of an AFPNet module according to an embodiment of a method for detecting small targets in traffic scenes based on an adaptive spatial aggregation pyramid according to the present invention;

[0040] Figure 4 This is a structural diagram of the MAA multi-scale aggregation attention mechanism of an embodiment of the small target detection method in traffic scenes based on an adaptive spatial aggregation pyramid of the present invention. DETAILED DESCRIPTION

[0041] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0042] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.

[0043] Example 1

[0044] like Figure 1 As shown, the present invention provides a small target detection method in traffic scenes based on adaptive spatial aggregation pyramid, comprising the following steps:

[0045] S1, such as Figure 2 As shown in the figure, prepare the typical small target detection datasets TT100K and CityPersons, and convert the data annotation file format to the MS COCO data format.

[0046] The TT100K dataset contains 9176 traffic scene images, of which 6105 are used for training and 3071 are used for testing. Although the TT100K dataset contains 221 categories, the number of instances of most categories is small. In the experiment, 176 categories with less than 100 instances are excluded, and 45 categories with more than 100 instances are used (w55, pm20, pl20, pg, pl70, pm55, p27, il100, w13, ph4, p19, pm30, ph5, wo, p6, w32, pn, pne, i5, p11, pl40, po, pl50, pl80, io, pl60, p26, i4, pl100, il80, The CityPersons dataset is built based on the Cityscapes dataset. Through fine manual annotation, it covers the task of human detection in urban environments. The dataset contains 2975 pictures for training, 500 pictures for verification and 1575 pictures for testing. There are an average of 7 pedestrians in each picture, and annotations of visible areas and full bodies are provided.

[0047] S2. Use DarkNet-53 as the backbone network for small target detection, and use the three-layer multi-scale feature maps of Dark3, Dark4, and Dark5 as the input of the adaptive spatial aggregation pyramid (SAFPN). The sizes of the C3, C4, and C5 feature maps are 80*80*256, 40*40*512, and 20*20*1024, respectively.

[0048] S3. Design an adaptive spatial aggregation pyramid. Currently, small object detection networks widely use multi-scale fusion technology for feature fusion and interaction. Multi-scale fusion is a key technology for small object detection. By integrating shallow details with deep semantic features, it significantly improves edge preservation and semantic differentiation of small objects. However, it faces core challenges such as the trade-off between feature resolution and semantics, and the effectiveness of cross-scale alignment bias.

[0049] To overcome this limitation, an adaptive spatial aggregation pyramid structure (ASFPN) is proposed. It strengthens the semantic information of small targets through progressive fusion, dynamic convolution (TVConv) and multi-scale aggregate attention (MAA), optimizes multi-scale feature fusion, and adaptively combines shallow high-resolution features (preserving details) with deep semantic features to further improve the effectiveness of feature fusion and enhance model detection performance.

[0050] The idea of ​​the gradual fusion structure comes from the gradual fusion feature pyramid AFPNet, which starts by fusing two adjacent low-level features and gradually incorporates high-level features into the fusion process. In this way, the large semantic gap between non-adjacent feature layers can be avoided. The model structure is as follows Figure 3 However, AFPNet has three major issues: First, assigning spatial weights to features at different levels through the Adaptive Spatial Fusion Module (ASF) can lead to over-attention to certain feature layers, resulting in loss of potential information; second, continuous downsampling can cause loss of information about small objects; and third, failure to consider the positional features of individual feature maps can lead to errors in the alignment of multiple feature maps. To address these issues, the present invention provides improvements and solutions in S3.1 and S3.2.

[0051] S3.1. Based on the asymptotic feature pyramid, dynamic convolution TVConv is added, replacing the traditional ConvsBNSILU structure with the TVConvBNSILU structure; dynamic convolution is added to the asymptotic fusion. Dynamic convolution TVConv enhances local feature adaptability through layout-aware translation-variable convolution. The translation-variable nature of dynamic convolution TVConv enables the model to focus on local features such as the boundaries and textures of the target. In traffic scene images, for the detection of traffic signs and pedestrians, TVCony can apply different weights to convolution operations for different positions such as traffic sign shapes, textures that are different from the environment, and pedestrian limbs, thereby more accurately extracting unique small target position features and inputting them into multi-scale aggregated attention.

[0052] The C3, C4 and C5 feature maps output by CSPDarkNet53 (whose sizes are 80*80*256, 40*40*512 and 20*20*1024 respectively) are reduced to a quarter of the input depth (whose sizes are 80*80*64, 40*40*128 and 20*20*256 respectively) through 1×1 convolution in the CBS module and then connected to the TVCBS module. The compact affinity maps in TVCons can be learned to capture the relationship between input pixel pairs, and this information is used to generate variable weights to enhance the unique small object position features.

[0053] S3.2. Input the enhanced feature map into the adaptive multi-scale aggregation attention module MAA to construct the adaptive multi-scale aggregation attention module MAA. The structure of the adaptive multi-scale aggregation attention module MAA is as follows: Figure 4 As shown. Upsampling and downsampling are optimized through wavelet convolution (WTConv) and deconvolution (TPConv), and the weights of important channels of fusion feature maps are strengthened through global average pooling and one-dimensional convolution. jMulti-scale fusion is performed with the mapped feature channels. The fused feature maps have rich position information and semantic information and reduce cross-scale alignment deviation.

[0054] The multi-scale aggregate attention module MAA first uses wavelet convolution WTConv and deconvolution TPConv to input n feature maps F i After sampling to the same size, the fusion is performed; then the overall information of the fused feature map is enhanced through global average pooling; then the multi-scale aggregation attention module MAA considers the cross-channel interaction information of each channel and its k neighbors through one-dimensional convolution. Through experiments, k=3 is set to map the channel weight after one-dimensional convolution to the range of [0,1] through SILU activation function; finally, the n feature maps are processed by their respective weight coefficients α i Fuse it with the mapped feature channels to obtain the final output feature map F'. The following formula is the calculation formula for multi-scale aggregate attention MAA;

[0055]

[0056] Among them, F i represents the feature map of the i-th branch; n represents the number of feature maps through the multi-scale aggregation attention module MAA; F' represents the fused feature map; M represents the channel attention weight after channel attention; α j Indicates the weight corresponding to j branches; b j is a learnable three-element one-dimensional tensor with all initial values ​​1.

[0057] After feature fusion through MAA, four residual units are used to continue learning features. These residual units are similar to ResNet, and each residual unit includes two 3×3 convolutions. The above process is repeated twice as the low-level, high-level, and top-level features are gradually integrated. Finally, the TVCBS module is used again to further enhance the unique small object position features. The enhanced features are input into the CBS module and restored to their original sizes (80*80*256, 40*40*512, and 20*20*1024) before being input into the detection head. These improvements help to strengthen the semantics of small objects and the effectiveness of multi-scale fusion, improving detection performance.

[0058] S4, construct a decoupled head structure detection head to separate the classification and detection heads;

[0059] S4.1. The NWD loss function has excellent performance for small target detection, but its convergence is slow. Therefore, we propose an adaptive small target loss function that adaptively combines the NWD loss function with the α_EIOU loss function to improve the small target detection rate while accelerating the convergence speed. This function is named the NE_IOU loss function. The classification loss uses binary cross entropy (BCE), and the localization loss function uses the designed small target loss function NE_IOU.

[0060] The small target loss function NE_IoU is composed of the small target detection loss function NWD, the α_EIOU loss function proposed in the BANet paper, and the adaptive coefficient γ. The calculation formula of α_EIOU is as follows:

[0061]

[0062] Among them, w, h, w gt and h gt Represents the width and height of the predicted box and the real box respectively, c is the diagonal length of the minimum bounding box covering the two boxes, ρ 2α (b,b gt ) is the distance between the center points of the two boxes, ρ 2α (w,w gt ) is the distance between the two box widths, ρ 2α (h,h gt ) is the distance between the two box heights, α = 3.1;

[0063] The NWD loss function is introduced based on α_EIOU. NWD uses the normalized Wasserstein distance to reduce the sensitivity to position deviations of small objects and effectively improve the detection performance of small targets, but the convergence is slow. The calculation formula of the NDS loss function is as follows:

[0064]

[0065] Among them, the real bounding box A=(x A ,y A ,w A ,h A ), predicted bounding box B = (x B ,y B ,w B ,h B ), x, y, w, and h are the center coordinates and width and height of the box respectively; NA and NB are Gaussian distributions modeled by the true bounding box A and the predicted bounding box B; is a distance metric; c is a constant closely related to the dataset;

[0066] By setting the adaptive parameters, the combination ratio of NWD and α_EIOU is dynamically adjusted as the final bounding box regression loss function. The following formula is the formula expression of the adaptive coefficient:

[0067]

[0068] Among them, λ is the balance coefficient, taking λ = 0.35, w and h represent the height and width of the real sample frame respectively, w l With h l Represent the height and width of the large target sample, w l With h l Both are 96;

[0069] Therefore, the formula expression of the NE_IOU regression loss function is:

[0070]

[0071] The α_EIOU loss function solves the problem of slowing convergence due to the overly complex aspect ratio measurement of CIoU. By taking into account the overlapping area, center distance, and aspect ratio, the convergence speed is accelerated. However, the introduction of a large number of position parameters increases the sensitivity to position deviations of small objects. This is not conducive to small target detection. To overcome the above problems, an adaptive small target loss function is proposed that adaptively combines the NWD loss function with the α_EIOU loss function to improve the small target detection rate while accelerating the convergence speed. It is named the NE_IOU loss function. Since the adaptive coefficient requires the height and width of the target bounding box, the target bounding box needs to be scaled before being filtered by the sample assigner to project it onto three detection heads of different sizes (80*80, 40*40, and 20*20) for loss calculation. Directly obtaining the size data is the scaled data. Therefore, before scaling the true sample frame, a copy of the sample is made, and then the copied true frame sample of the original size and the scaled true frame sample are input into the sample distributor for screening. In the NE_IoU loss function, the original true sample size after screening is used to obtain the corresponding γ, thereby obtaining the NE_IoU loss.

[0072] S4.2. Soft-NMS is used to alleviate the problem of small targets being mistakenly suppressed due to IoU calculation bias or dense arrangement. Traditional NMS directly deletes all candidate boxes whose IoU with the highest-scoring box exceeds a threshold, while Soft-NMS dynamically reduces the score of overlapping boxes based on the IoU value instead of directly setting it to zero, thereby retaining potential valid detection results.

[0073] S5. Input the training set into the overall detection network for training and save the optimal model; input the test set into the saved optimal model for testing to verify the detection effect of the improved model. The model training uses the Ubuntu 20.04 operating system, the CPU is AMD Ryzen 97945HX, the GPU is a single NVIDIA GeForce RTX 4060 (8GB GDRR6 / 140W), the operating environment is Python 3.8, PyTorch 1.11.0, CUDA11.3, and the target detection toolbox mmdetection is used;

[0074] During data preprocessing, data augmentation uses operations such as random horizontal flipping, mosaic, random affine transformation, and mixup to increase the diversity of training data, simulating the various variations and noise found in real scenes, and helping the model better adapt to different environments and scenarios. The images are then resized to 640*640 pixels and fed into the Darknet-53 backbone network. The SGD optimizer is used with momentum set to 0.9, weight decay to 0.0005, and an initial learning rate to 0.001. The batch size is set to 4, and the number of training epochs is set to 300. The test set is fed into the optimal model saved from training for testing, and the proposed method is compared to current mainstream small object detection networks to verify its effectiveness.

[0075] To demonstrate the effectiveness of the proposed method for small object detection based on adaptive spatial aggregation pyramids, we compared the proposed method with several typical small object detection algorithms on the TT100K and CityPersons datasets. The detection results were evaluated using mean average precision (mAP) 0.5:0.95 and small object detection accuracy (APs). The proposed method achieved good results on both the TT100K and CityPersons datasets, as shown in Tables 1 and 2.

[0076] Table 1 Comparison of detection effects of different algorithms on the TT100K dataset

[0077] Model Input size mAP0.5:0.9 <![CDATA[AP s ]]> FLOPs / Params / M Faster-RCNN 1024×10244 58.9 27.7 52.5 41.6 RetinaNet 2048×2048 49.2 - 35.1 27.8 Cascade-RCNN 1024×1024 63.5 36.2 84.5 69.3 FCOS 2048×2048 63.8 - 49.5 39.2 PSG-YOLOv5 1024×1024 53.2 35.6 131.3 106.2 AIE-YOLO 512×512 61.7 - - - YOLOv8s 640×640 61.3 37.3 14.8 11.6 Ghost-YOLOv8 640×640 44.7 - 3.5 2.8 TF-Fusion 640×640 39.5 43.6 - - Ours 640×640 64.6 45.5 17.1 13.5

[0078] Table 2 Comparison of detection effects of different algorithms on the CityPersons dataset

[0079] Model Input size mAP0.5:0.9 <![CDATA[AP s ]]> FLOPs / Params / M Faster-RCNN 1024×10244 27.4 15.1 52.5 41.6 RetinaNet 2048×2048 35.4 20.6 35.1 27.8 Sparse-RCNN 1024×1024 34.3 13.9 74.5 59.3 TOLOv7tiny 640×640 30.2 19.8 7.2 6.5 YOLOv8s 640×640 39.6 28.6 14.8 11.6 Ours 640×640 43.7 32.2 17.1 13.5

[0080] The experimental evaluation indicators used for quantitative evaluation are average precision (mAP), floating-point operations (FLOPs), and parameter count (Params); mAP@0.5:0.95 represents the average precision of IoU from 0.5 to 0.95, and AP SIt represents the detection accuracy of small objects. The floating-point operation number FLOPs is used to measure the complexity of the model. The parameter amount Params represents the number of parameters of the model and measures the demand for graphics card performance.

[0081] As can be seen from Tables 1 and 2, the small target detection method proposed in this paper achieves the highest mean average accuracy (mAP0.5) of 0.95, and also achieves the best small target detection precision (APs) compared to other algorithms. Furthermore, the model proposed in this paper uses fewer parameters and fewer floating-point operations than most mainstream small target detection models. This demonstrates that the small target detection method proposed in this paper, based on adaptive spatial aggregation pyramids, is beneficial in both improving detection accuracy and meeting real-time detection requirements.

[0082] Therefore, the present invention adopts the above-mentioned small target detection method in traffic scenes based on adaptive spatial aggregation pyramid. By designing multi-scale aggregate attention (MAA), adaptive spatial aggregation pyramid (SAFPN) and small target loss function NE_IoU, the representation ability of small target features is significantly improved; while maintaining real-time performance, it can effectively alleviate the problems of missed detection and false detection of small targets caused by weak semantic information, insufficient fusion and low positioning accuracy.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A small target detection method in traffic scenes based on an adaptive spatial aggregation pyramid, characterized by: The following steps are involved: S1. Prepare the typical small object detection datasets TT100K and CityPersons, and convert the data annotation file format to the MS COCO data format; S2. Use DarkNet-53 as the backbone network for small target detection, and output the backbone network C3, C4, and C5 feature layers to the adaptive spatial aggregation pyramid for feature fusion; S3, design an adaptive spatial aggregation pyramid; S3.

1. Based on the asymptotic feature pyramid, dynamic convolution TVConv is added, that is, the traditional ConvsBNSILU structure is replaced by TVConvBNSILU structure; S3.

2. Construct an adaptive multi-scale aggregation attention module MAA and embed it into the adaptive spatial aggregation pyramid; S4, construct a decoupled head structure detection head to separate the classification and detection heads; S4.1, the classification loss uses binary cross entropy BCE, and the positioning loss function uses the designed small target loss function NE_IOU; S4.

2. Soft-NMS is used to alleviate the problem of small objects being incorrectly suppressed due to IoU calculation deviation or dense arrangement. S5. Input the training set into the overall detection network for training and save the optimal model; input the test set into the optimal model saved during training for testing to verify the detection effect of the improved model.

2. The method for detecting small targets in traffic scenes based on adaptive spatial aggregation pyramid according to claim 1, characterized in that: In S3.1, dynamic convolution is added to the asymptotic fusion, and the dynamic convolution TVConv enhances the adaptability of local features through layout-aware translation-variant convolution.

3. The method for detecting small targets in traffic scenes based on adaptive spatial aggregation pyramid according to claim 1, characterized in that: In S3.2, the multi-scale aggregation attention module MAA first performs wavelet convolution WTConv and deconvolution TPConv on the input n feature maps F i After sampling to the same size, the fusion is performed; then the overall information of the fused feature map is enhanced through global average pooling; then the multi-scale aggregation attention module MAA considers the cross-channel interaction information of each channel and its k neighbors through one-dimensional convolution. Through experiments, k=3 is set to map the channel weight after one-dimensional convolution to the range of [0,1] through SILU activation function; finally, the n feature maps are processed by their respective weight coefficients α i Fuse it with the mapped feature channels to obtain the final output feature map F'. The following formula is the calculation formula for multi-scale aggregate attention MAA; Among them, F i represents the feature map of the i-th branch; n represents the number of feature maps through the multi-scale aggregation attention module MAA; F' represents the fused feature map; M represents the channel attention weight after channel attention; α j Indicates the weight corresponding to j branches; b j is a learnable three-element one-dimensional tensor with all initial values ​​1.

4. The method for detecting small targets in traffic scenes based on adaptive spatial aggregation pyramid according to claim 1, characterized in that: In S4.1, the small target loss function NE_IoU is composed of the small target detection loss function NWD, the α_EIOU loss function proposed in the BANet paper, and the adaptive coefficient γ. The calculation formula of α_EIOU is as follows: Among them, w, h, w gt and h gt Represents the width and height of the predicted box and the real box respectively, c is the diagonal length of the minimum bounding box covering the two boxes, ρ 2α (b,b gt ) is the distance between the center points of the two boxes, ρ 2α (w,w gt ) is the distance between the two box widths, ρ 2α (h,h gt ) is the distance between the two box heights, α = 3.1; The NWD loss function is introduced based on α_EIOU. NWD uses the normalized Wasserstein distance to reduce the sensitivity to position deviations of small objects and effectively improve the detection performance of small targets, but the convergence is slow. The calculation formula of the NDS loss function is as follows: Among them, the real bounding box A=(x A ,y A ,w A ,h A ), predicted bounding box B = (x B ,y B ,w B ,h B ), x, y, w, and h are the center coordinates and width and height of the box respectively; NA and NB are Gaussian distributions modeled by the true bounding box A and the predicted bounding box B; is a distance metric; c is a constant closely related to the dataset; By setting the adaptive parameters, the combination ratio of NWD and α_EIOU is dynamically adjusted as the final bounding box regression loss function. The following formula is the formula expression of the adaptive coefficient: Among them, λ is the balance coefficient, taking λ = 0.35, w and h represent the height and width of the real sample frame respectively, w l With h l Represent the height and width of the large target sample, w l With h l Both are 96; Therefore, the formula expression of the NE_IOU regression loss function is:

5. The method for detecting small targets in traffic scenes based on adaptive spatial aggregation pyramid according to claim 1, characterized in that: In S4.2, the traditional NMS directly deletes all candidate boxes whose IoU with the highest-scoring box exceeds a threshold, while Soft-NMS dynamically reduces the score of overlapping boxes based on the IoU value instead of directly setting it to zero, thereby retaining potential valid detection results.

6. The method for detecting small targets in traffic scenes based on adaptive spatial aggregation pyramid according to claim 1, characterized in that: The experimental evaluation indicators used for quantitative evaluation in the S5 are mean average precision (mAP), floating-point operations (FLOPs), and parameter count (Params). Among them, mAP@0.5:0.95 represents the average precision with an IoU range of 0.5 to 0.95, FLOPs is used to measure the complexity of the model, and Params represents the number of model parameters, which measures the demand for graphics card performance.

Citation Information

Cited By

  • Rice brown spot detection and counting method, system and equipment

    CN121482047A

  • A method, system and equipment for detecting and counting brown spot disease in rice.

    CN121482047B

  • Unmanned aerial vehicle overlooking small target detection method based on density perception and spatial hierarchy

    CN121640326A