Lightweight traffic vehicle detection method under view angle of unmanned aerial vehicle

By improving the YOLOv11 algorithm, using high-resolution detection layer, shared detection head, multi-band feature decomposition and Soft-NMS algorithm, multiple challenges of vehicle detection from the perspective of the drone are solved, and high-precision and lightweight vehicle detection effect is achieved.

CN120388311APending Publication Date: 2025-07-29浪潮智慧城市科技有限公司
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510545890.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Urban traffic vehicle detection from the perspective of drones faces multiple challenges such as small target size, large scale changes, complex background interference and occlusion, and it is difficult for the existing technology to achieve efficient and lightweight vehicle detection.

Method used

Based on the YOLOv11 algorithm, the detection performance is improved through multi-resolution detection architecture design, multi-scale feature extraction optimization, scalable receptive field embedding and soft threshold post-processing strategies, including the addition of a high-resolution detection layer, shared lightweight detection head, multi-band feature decomposition, scalable receptive field module and Soft-NMS algorithm.

Benefits of technology

It realizes high-precision and lightweight vehicle detection from the perspective of drones, improves the accuracy and efficiency of small target detection, and is suitable for equipment with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388311A_ABST
    Figure CN120388311A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight traffic vehicle detection method under the view angle of an unmanned aerial vehicle, and relates to the technical field of target detection, and the method is based on a YOLOv11 algorithm, and achieves the improvement of the vehicle detection performance through the following aspects: on the basis of maintaining a three-scale detection architecture, adding a high-resolution detection layer for a small target, more detail features are reserved; a shared lightweight detection head is adopted, so that the network parameter quantity and the calculation quantity are reduced; wavelet transform is used in the backbone and neck network to replace traditional convolution, and low-frequency and high-frequency components are used to extract multi-scale features; an RDS module is embedded in the C3K2 module to realize high-level feature sensing range expansion and deep and shallow feature fusion; in a target frame screening stage, a Soft-NMS algorithm is adopted to replace NMS, high-confidence prediction in an overlapped frame is reserved through a softening suppression strategy, and missing detection of small targets due to hard threshold filtering is avoided. The method is applied to vehicle detection under the view angle of the unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and specifically to a lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle (UAV). Background Art

[0002] In recent years, with the rapid growth of the number of motor vehicles, the problem of traffic congestion has become more serious. The traditional human-based traffic management mode has significant deficiencies in real-time monitoring, response speed, and data processing capabilities, and it is difficult to meet the needs of modern urban development. The efficiency and safety of urban traffic systems are crucial for sustainable development, but the rapid increase in population and vehicles makes it difficult for traditional management means to cope with the increasingly complex traffic conditions. Intelligent transportation systems (ITS) achieve information interaction between vehicles and traffic facilities through vehicle-road cooperation technology, improving traffic management efficiency. In intelligent transportation systems, vehicle detection and tracking technology from the perspective of UAVs has become an important research direction due to its flexibility and high efficiency. Different from traditional fixed cameras with limited coverage and high deployment costs, UAVs, with their high mobility and wide-area coverage capabilities, have become an efficient technology option for wide-area traffic monitoring. By accurately locating and quickly identifying the positions of vehicles in UAV images and combining real-time information transmission, intelligent transportation systems can not only effectively relieve traffic congestion but also reduce the accident rate through data analysis. In addition, UAVs have shown their important potential for target detection and are widely used in military, civilian, and commercial fields. Compared with traditional methods, UAV technology combining high-resolution sensors and deep learning algorithms has significantly improved the detection efficiency and achieved a higher level of automation. The rapid development of this technology has not only accelerated the construction of intelligent transportation systems but also expanded the broad application prospects of UAV perspective target detection in smart cities.

[0003] In urban traffic scenarios, vehicle detection from long-distance UAV photography faces multiple challenges such as small target sizes, large scale variations, complex background interference, and occlusion. Summary of the Invention

[0004] In view of the requirements and deficiencies of the current technological development, the present invention provides a lightweight traffic vehicle detection method from the perspective of a UAV to achieve more accurate and lightweight vehicle detection and better adapt to the urban vehicle detection task from the perspective of a UAV.

[0005] The technical solution adopted by the lightweight traffic vehicle detection method from the perspective of a UAV of the present invention to solve the above technical problems is as follows:

[0006] A lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle (UAV). This method is based on the YOLOv11 algorithm and comprehensively improves the target detection performance of low-altitude UAVs through four aspects: multi-resolution detection architecture design, multi-scale feature extraction optimization, scalable receptive field embedding, and soft threshold post-processing strategy. Specifically, it includes:

[0007] (A) On the basis of maintaining the three-scale detection architecture of the YOLOv11 algorithm, a high-resolution detection layer for small targets is newly added to retain more detailed features to enhance the positioning and feature extraction of small-sized vehicles. At the same time, a shared lightweight detection head (SLD) is adopted, and the number of network parameters and computational complexity are reduced through a parameter sharing mechanism;

[0008] (B) In the backbone network CSPDarkNet and the neck network PAFPN of the YOLOv11 algorithm, the pre-specified convolutional layers are replaced with wavelet pooling modules, and multi-scale feature extraction of vehicle targets from the perspective of UAVs is achieved through multi-band feature decomposition and downsampling operations;

[0009] (C) In the C3K2 module of the YOLOv11 algorithm, a scalable receptive field (RDS) module is embedded. Through multi-branch dilated convolution, residual structure, and depthwise separable convolution, the perception range of high-level features is extended and the deep and shallow layer features are fused to improve the detection accuracy of small targets and alleviate the vanishing gradient;

[0010] (D) In the target box screening stage of the YOLOv11 algorithm, the Soft-NMS algorithm is used to replace NMS. By means of a soft suppression strategy, high-confidence predictions in overlapping boxes are retained, avoiding the missed detection of small targets due to hard threshold filtering and improving the detection accuracy.

[0011] Optionally, when performing (A), the newly added high-resolution detection layer for small targets specifically performs the following operations:

[0012] The size of the deep feature map output by the backbone network CSPDarkNet is enlarged to 160×160×64 through three upsampling operations, so that small targets occupy more pixels on the feature map, retaining the detailed features of tire textures and headlight contours, and improving the positioning accuracy and feature expression ability;

[0013] At the same time, the number of channels of the backbone network CSPDarkNet is compressed, and the number of channels of the last three layers of the backbone network CSPDarkNet is reduced from 1024 to 512. By reducing the thickness of the feature map, the number of parameters is reduced, avoiding semantic dilution of small target information caused by high-dimensional convolution in the deep network.

[0014] Further optionally, when performing (A), the shared lightweight detection head (SLD) realizes the dual optimization of detection performance and efficiency through the following operations:

[0015] Adjust the number of channels of the input feature layer using 1×1 convolutions, while reducing the number of parameters and retaining key features;

[0016] Pool the processed feature layers into a shared convolutional module, which consists of 1×1 convolutions for extracting inter-channel interaction features, 3×3 convolutions for capturing local context information, and pooling operations for extracting multi-scale features, to achieve efficient extraction of multi-scale features;

[0017] During the training phase, the independent branches of the shared lightweight detection head SLD act on the input feature map respectively:

[0018] Classification branch: Output the target class probability through the formula Y1 = W1 * X + b1, where X represents the input feature map, and W1 and b1 are parameters optimized through backpropagation during the training process;

[0019] Regression branch: Output the bounding box coordinates through the formula Y2 = W2 * X + b2, where X represents the input feature map, and W2 and b2 are parameters optimized through backpropagation during the training process;

[0020] Scale-aware branch: Enhance the robustness to targets of different scales through the formula Y3 = Pooling(X) + b3, where Pooling represents the pooling operation, X represents the input feature map, and b3 is a parameter optimized through backpropagation during the training process.

[0021] Further optionally, operation (B) specifically includes:

[0022] (B1) The wavelet pooling module processes the input feature map in parallel through a convolutional layer containing 4 wavelet filter kernels LL, LH, HL, and HH, where: the size of each wavelet filter kernel is k×k, and the depth matches the number of channels of the input feature map; the convolution operation adopts the grouped convolution mode, and the number of groups is set to the number of input channels to ensure that each channel of the input feature map independently uses all filters;

[0023] (B2) The wavelet pooling module uses the predefined 4 wavelet filter kernels LL, LH, HL, and HH to decompose the input feature map into the low-frequency global feature map LL and the high-frequency detail feature maps LH, HL, and HH in one-time through a single convolution operation;

[0024] (B3) Set the stride in the convolution operation to 2 to reduce the size of the output feature map by 1 / 2, and achieve downsampling while decomposing the multi-band;

[0025] (B4) The pooling process of the wavelet pooling module is represented by the following formula:

[0026] F out = Conv2D(F in,{h LL ,h LH ,h HL ,h HH}, stride = 2),

[0027] where, F in is the input feature map; F out is the decomposed multi-scale feature map, including low-frequency global features and high-frequency edge features; {h LL ,h LH ,h HL ,h HH} are 4 wavelet filter kernels;

[0028] (B5) The low-frequency global feature map LL extracts the global structural features of the image through low-pass filtering, which is used for the high-level abstraction of subsequent semantic features. At the same time, the high-frequency detail feature maps LH, HL, and HH capture horizontal edges, vertical edges, and diagonal textures respectively through high-pass filtering, enhancing the sensitivity to the edge details of small targets;

[0029] (B6) The low-frequency global feature map LL is input into the deep layer of the backbone network through PAFPN to participate in the high-level abstraction of vehicle semantic features. The high-frequency detail feature maps LH, HL, and HH are fused with the shallow features of PAFPN to strengthen the expression of the edge details of small targets.

[0030] Preferably, the size of each wavelet filter kernel is k×k, and the value is 3.

[0031] Further optionally, operation (C) specifically includes:

[0032] (C1) Embed the scalable receptive field RDS module in the C3K2 module of YOLOv11, replacing the traditional post RFB (Receptive Field Block) structure to form the C3K2_RDS module;

[0033] (C2) The RDS module receives the input feature map passed from the C3K2 module. First, it performs a 3×3 per-channel convolution, independently performing convolution operations on each channel of the input feature map, and then performs a 1×1 pointwise convolution to integrate feature information through cross-channel linear combination;

[0034] (C3) The RDS module adopts a three-parallel-branch structure, and the three branches use 3×3 depthwise separable convolutions with different dilation rates to extract features from different perspectives;

[0035] (C4) After depthwise separable convolutions with different dilation rates through three parallel branches, each branch outputs corresponding feature maps, which contain feature information at different scales and perspectives; the BN layer normalizes the feature maps output by each branch, making the mean of the feature maps 0 and the variance 1; after the BN layer processing, the normalized feature maps are subtracted from the input feature maps to generate semantic residuals;

[0036] (C5) After feature extraction through three parallel branches and BN layer processing, multiple feature maps containing feature information at different scales and perspectives are obtained. The RDS module uses 1×1 pointwise convolutions for feature fusion; meanwhile, the 1×1 pointwise convolutions adjust the number of channels of the feature maps, fusing multiple feature maps into a comprehensive feature map to provide a richer and more effective feature representation for subsequent object detection tasks;

[0037] (C6) After multi-branch feature extraction, BN layer processing, and 1×1 pointwise convolution fusion, a fused feature map is obtained. The RDS module directly superimposes the input feature map onto the fused output feature map in a residual connection manner.

[0038] Preferably, in step (C3), the RDS module adopts a three-parallel-branch structure, and the three branches respectively use 3×3 depthwise separable convolutions with dilation rates of 1, 3, and 5:

[0039] The first branch uses a 3×3 depthwise separable convolution with a dilation rate of 1, which is responsible for capturing local detail information in the feature map;

[0040] The second branch uses a 3×3 depthwise separable convolution with a dilation rate of 3, enabling the convolution kernel to span a set interval during convolution and expanding the receptive field;

[0041] The third branch uses a 3×3 depthwise separable convolution with a dilation rate of 5 to further expand the receptive field and capture more macroscopic global features.

[0042] Further optionally, operation (D) specifically includes:

[0043] (D1) Output all predicted target boxes and their corresponding confidence scores, and sort them from high to low confidence to generate an ordered list;

[0044] (D2) Select the box with the highest current confidence from the sorted ordered list as the main box, and calculate the intersection over union (IoU) between it and all the remaining boxes in the ordered list;

[0045] (D3) Gradually decay the confidence of the overlapping boxes according to the size of the intersection over union, and retain the boxes with high confidence;

[0046] (D4) Add the currently selected main bounding box to the final detection result list and remove it from the original list. Then repeat steps (D2)-(D3) for the remaining bounding boxes until the ordered list is empty. Finally, retain all the bounding boxes with confidence levels higher than the preset threshold as the detection results for output.

[0047] A lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle according to the present invention has the following beneficial effects compared with the prior art:

[0048] Based on the YOLOv11 algorithm, the present invention comprehensively improves the performance of low-altitude unmanned aerial vehicle vehicle detection through four aspects: multi-resolution detection architecture design, multi-scale feature extraction optimization, scalable receptive field embedding, and soft-threshold post-processing strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] FIG. Figure 1 is a schematic diagram of the optimized content of the method described in Embodiment 1 of the present invention;

[0050] FIG. Figure 2 is a network structure diagram of the detection model designed based on the detection method in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] To make the technical solutions, the technical problems to be solved, and the technical effects of the present invention more clearly understood, the following describes the technical solutions of the present invention clearly and completely in conjunction with specific embodiments.

[0052] Embodiment 1:

[0053] Referring to FIG. Figure 1 , this embodiment proposes a lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle. Based on the YOLOv11 algorithm, this method comprehensively improves the performance of low-altitude unmanned aerial vehicle target detection through four aspects: multi-resolution detection architecture design, multi-scale feature extraction optimization, scalable receptive field embedding, and soft-threshold post-processing strategy. Specifically, it includes:

[0054] (A) On the basis of maintaining the three-scale detection architecture of the YOLOv11 algorithm, add a high-resolution detection layer for small targets to retain more detailed features to enhance the positioning and feature extraction of small-sized vehicles. At the same time, adopt the shared lightweight detection head SLD to reduce the number of network parameters and the amount of calculation through the parameter sharing mechanism.

[0055] Performing (A), the newly added high-resolution detection layer for small targets specifically performs the following operations:

[0056] The size of the deep feature map output by the backbone network CSPDarkNet is enlarged to 160×160×64 through three upsampling operations, enabling small targets to occupy more pixels on the feature map, preserving the detailed features of tire textures and headlight contours, and improving the positioning accuracy and feature expression ability;

[0057] Meanwhile, the number of channels of the backbone network CSPDarkNet is compressed, and the number of channels of the last three layers of the backbone network CSPDarkNet is reduced from 1024 to 512. By reducing the thickness of the feature map, the number of parameters is reduced, avoiding semantic dilution of small target information caused by high-dimensional convolution in the deep network.

[0058] Execute (A). The shared lightweight detection head SLD realizes the dual optimization of detection performance and efficiency through the following operations:

[0059] Adjust the number of channels of the input feature layer using 1×1 convolution, preserving key features while reducing the number of parameters;

[0060] Pool the processed feature layer into a shared convolution module, which consists of 1×1 convolution for extracting inter-channel interaction features, 3×3 convolution for capturing local context information, and pooling operations for extracting multi-scale features, achieving efficient extraction of multi-scale features;

[0061] In the training stage, the independent branches of the shared lightweight detection head SLD act on the input feature map respectively:

[0062] Classification branch: Output the target class probability through the formula Y1 = W1*X + b1, where X represents the input feature map, and W1 and b1 are parameters optimized through backpropagation during the training process;

[0063] Regression branch: Output the bounding box coordinates through the formula Y2 = W2*X + b2, where X represents the input feature map, and W2 and b2 are parameters optimized through backpropagation during the training process;

[0064] Scale perception branch: Enhance the robustness to targets of different scales through the formula Y3 = Pooling(X) + b3, where Pooling represents the pooling operation, X represents the input feature map, and b3 is a parameter optimized through backpropagation during the training process.

[0065] (B) In the backbone network CSPDarkNet and the neck network PAFPN of the YOLOv11 algorithm, replace the pre-specified convolutional layer with a wavelet pooling module, and realize the extraction of multi-scale features of vehicle targets from the perspective of drones through multi-band feature decomposition and downsampling operations; this process specifically includes:

[0066] (B1) The wavelet pooling module performs parallel processing on the input feature map through a convolutional layer containing four wavelet filter kernels LL, LH, HL, and HH, where: each wavelet filter kernel has a size of k×k, k takes the value of 3, and the depth matches the number of channels of the input feature map; the convolution operation adopts the grouped convolution mode, and the number of groups is set to the number of input channels to ensure that each channel of the input feature map independently uses all filters;

[0067] (B2) The wavelet pooling module uses four predefined wavelet filter kernels LL, LH, HL, and HH to decompose the input feature map into a low-frequency global feature map LL and high-frequency detail feature maps LH, HL, and HH in one-time through a single convolution operation;

[0068] (B3) In the convolution operation, the stride is set to 2, reducing the size of the output feature map by 1 / 2, and achieving downsampling while decomposing multiple frequency bands;

[0069] (B4) The pooling process of the wavelet pooling module is represented by the following formula:

[0070] F out =Conv2D(F in ,{h LL ,h LH ,h HL ,h HH},stride = 2),

[0071] In the formula, F in is the input feature map; F out is the decomposed multi-scale feature map, containing low-frequency global features and high-frequency edge features; {h LL ,h LH ,h HL ,h HH} are four wavelet filter kernels;

[0072] (B5) The low-frequency global feature map LL extracts the global structural features of the image through low-pass filtering for subsequent high-level abstraction of semantic features. At the same time, the high-frequency detail feature maps LH, HL, and HH capture horizontal edges, vertical edges, and diagonal textures through high-pass filtering respectively, enhancing the sensitivity to small target edge details;

[0073] (B6) The low-frequency global feature map LL is input to the deep layer of the backbone network through PAFPN to participate in the high-level abstraction of vehicle semantic features, and the high-frequency detail feature maps LH, HL, and HH are fused with the shallow features of PAFPN to strengthen the edge detail expression of small targets.

[0074] (C) Embed the expandable receptive field RDS module in the C3K2 module of the YOLOv11 algorithm. Through multi-branch dilated convolution, residual structure, and depthwise separable convolution, expand the perception range of high-level features and fuse shallow and deep features, improving the detection accuracy of small targets and alleviating the vanishing gradient. This process specifically includes:

[0075] (C1) Embed the expandable receptive field RDS module in the C3K2 module of YOLOv11, replacing the traditional post RFB (Receptive Field Block) structure to form the C3K2_RDS module;

[0076] (C2) The RDS module receives the input feature map passed from the C3K2 module. First, perform a 3×3 per-channel convolution, independently convolving each channel of the input feature map, and then perform a 1×1 pointwise convolution to integrate feature information through cross-channel linear combination;

[0077] (C3) The RDS module adopts a three-parallel-branch structure, and the three branches respectively use 3×3 depthwise separable convolutions with dilation rates of 1, 3, and 5:

[0078] The first branch uses a 3×3 depthwise separable convolution with a dilation rate of 1, responsible for capturing local detail information in the feature map;

[0079] The second branch uses a 3×3 depthwise separable convolution with a dilation rate of 3, enabling the convolution kernel to span a set interval during convolution and expanding the receptive field;

[0080] The third branch uses a 3×3 depthwise separable convolution with a dilation rate of 5 to further expand the receptive field and capture more macroscopic global features;

[0081] (C4) After the different dilation rate depthwise separable convolutions of the three parallel branches, each branch outputs the corresponding feature map. These feature maps contain feature information of different scales and perspectives. The BN layer normalizes the feature maps output by each branch, making the mean of the feature map 0 and the variance 1. After the BN layer processing, subtract the normalized feature map from the input feature map to generate a semantic residual;

[0082] (C5) After the feature extraction of the three parallel branches and the BN layer processing, multiple feature maps containing feature information of different scales and perspectives are obtained. The RDS module uses a 1×1 pointwise convolution for feature fusion. At the same time, the 1×1 pointwise convolution adjusts the number of channels of the feature map, fusing multiple feature maps into a comprehensive feature map to provide a richer and more effective feature representation for subsequent object detection tasks;

[0083] (C6) After multi-branch feature extraction, BN layer processing, and 1×1 pointwise convolution fusion, the fused feature map is obtained. The RDS module directly superimposes the input feature map onto the fused output feature map in a residual connection manner.

[0084] (D) In the target box screening stage of the YOLOv11 algorithm, the Soft-NMS algorithm is used to replace NMS. By the soft suppression strategy, high-confidence predictions in overlapping boxes are retained, avoiding the missed detection of small targets due to hard threshold filtering, and improving the detection accuracy.

[0085] This process specifically includes:

[0086] (D1) Output all predicted target boxes and their corresponding confidence scores, and sort them in descending order of confidence to generate an ordered list;

[0087] (D2) Select the box with the highest current confidence from the sorted ordered list as the main box, and calculate the intersection over union (IoU) between it and all the remaining boxes in the ordered list;

[0088] (D3) Gradually attenuate the confidence of overlapping boxes according to the size of the IoU, and retain high-confidence boxes;

[0089] (D4) Add the currently selected main box to the final detection result list and remove it from the original list. Then repeat steps (D2)-(D3) for the remaining boxes until the ordered list is empty. Finally, retain all target boxes with confidence higher than the preset threshold as the detection results for output.

[0090] Example Two:

[0091] A lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle proposed in Example One aims to improve the detection performance by improving the backbone network, neck network, detection head, and post-processing process of YOLOv11. Based on the detection method of this Example One, referring to Appendix Figure 2 , in this example, a lightweight traffic vehicle detection model YOLO-SWR is designed. Its architecture aims to achieve the descriptions in (A)-(D) through four core modules: multi-resolution detection architecture design, multi-scale feature extraction optimization, scalable receptive field embedding, and soft threshold post-processing strategy.

[0092] The experimental data of the lightweight traffic vehicle detection model YOLO-SWR involved in this embodiment is selected from the VisDrone2019 dataset. This dataset is collected from the perspective of drones in urban environments, covering a variety of environments, weather, and lighting scenarios, comprehensively covering the dynamic scene features in urban traffic environments. It contains various target categories such as pedestrians and cars, among which small targets account for up to 61%. Given that the focus of the research is on urban traffic environments, four types of traffic-related targets, namely cars, vans, buses, and trucks, are selected from this dataset. After screening, the training set, test set, and validation set contain 6,471, 1,610, and 548 images respectively. In addition, the VisDrone2019-MOT dataset is also used to further evaluate the multi-object tracking performance.

[0093] Details of the hardware, software, and experimental environment used in the experiment are shown in Table 1, and the training parameter settings are shown in Table 2. Other strategies and hyperparameters follow the benchmark model. To ensure the fairness and consistency of the model training effect, no pre-trained weights are used in all ablation experiments and comparative experiments.

[0094] Table 1 Experimental Environment

[0095]

[0096] Table 2 Training Parameters

[0097]

[0098] To verify the effectiveness of the lightweight traffic vehicle detection model YOLO-SWR involved in this embodiment, 7 groups of ablation experiments are designed and completed, and the experimental results are shown in Table 3. In Table 3, the first group is the experimental result of the basic YOLOv11 model; the "√" mark in the table represents the experiment conducted on the basis of the corresponding improvement point. Each group of experiments conducts ablation verification on each component of the basic model, aiming to quantify the contribution of individual improvement strategies to the overall performance.

[0099] Table 3 Ablation Experiments of YOLO-SWR Model

[0100]

[0101]

[0102] By gradually improving and analyzing the YOLOv11 model, the specific contribution of each module to the performance can be evaluated.

[0103] Adding an SLD detection head and a small object detection layer (Group ①) to the base model increased the recall rate from 42.4% to 44.1%, and mAP95 increased from 31.6% to 34.4%. At the same time, the number of parameters decreased significantly to 1.2M. This indicates that the SLD module and the small object detection layer can enhance the extraction of small object features without increasing resource consumption.

[0104] Subsequently, the C3k2_RDS module (Group ②) was introduced. By combining the residual structure and depthwise separable convolution, the model's ability to obtain high-level network features was improved, and mAP50 increased from 46.7% to 48.4%. Although the number of parameters increased slightly, its stability in processing complex object features was significantly enhanced. In addition, the WaveletPool module (Group ③), namely the wavelet pooling module, was adopted to replace the traditional convolution operation, strengthening the extraction of multi-scale features, increasing mAP95 to 31.9%, and reducing the computational overhead at the same time. GFLOPs decreased from 6.3G to 5.4G. This result shows that the WaveletPool module can effectively retain key information while reducing the computational amount.

[0105] After introducing the Soft-NMS module (Group ④), the bounding box screening strategy was optimized, further improving the target localization accuracy and increasing mAP95 to 40.2%.

[0106] The experimental results of combining the SLD module and the C3k2_RDS module (Group ⑤) showed that the synergistic effect of the two increased the recall rate to 47.7% and mAP50 reached 50.9%, fully demonstrating their complementarity in terms of accuracy and recall.

[0107] Adding the WaveletPool module (Group ⑥) on the basis of Group ⑤, although the recall rate decreased slightly, the overall accuracy still maintained high stability. At the same time, the number of parameters was further reduced to 1.1M. Finally, the YOLO-SWR model that combines all improvement points integrated the synergistic effects of each module in feature extraction, bounding box optimization, and computational efficiency, reaching 56.3% and 42.3% in mAP50 and mAP95 respectively, which were 9.6% and 10.7% higher than the base model YOLOv11. At the same time, the number of parameters was only 1.1M, and GFLOPs remained at 6.5G, verifying the efficiency of its lightweight design, being able to achieve a balance between detection accuracy and efficiency, and being suitable for complex application scenarios of real-time object detection by drones.

[0108] Based on the detection method of Embodiment 1, the lightweight traffic vehicle detection model YOLO-SWR designed in this embodiment solves the problems existing in the basic model YOLOv11, improves the detection performance of small targets from the perspective of drones, and is applicable to resource-constrained devices. By adding a small target detection layer and adopting an SLD detection head, YOLO-SWR reduces the model complexity while improving the detection accuracy of small targets. Secondly, in the backbone network and the neck network, wavelet pooling (WaveletPool) is used to replace the traditional convolution, enhancing the multi-scale feature extraction ability and simplifying the calculation process through convolution operations. The designed C3k2_RDS module strengthens the cross-layer connection and information transfer of the feature map while avoiding the problem of gradient disappearance. Finally, the Soft-NMS module is introduced to optimize the processing of detailed features and reduce the information loss caused by fixed thresholds.

Claims

1. A lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle, characterized in that, This method is based on the YOLOv11 algorithm and comprehensively improves the performance of low-altitude UAV target detection through four aspects: multi-resolution detection architecture design, multi-scale feature extraction optimization, scalable receptive field embedding, and soft-threshold post-processing strategy. Specifically, it includes: (A) On the basis of maintaining the three-scale detection architecture of the YOLOv11 algorithm, a high-resolution detection layer for small targets is added, and more detailed features are retained to enhance the localization and feature extraction of small-sized vehicles. At the same time, the shared lightweight detection head SLD is adopted, and the number of network parameters and computational volume are reduced through the parameter sharing mechanism; (B) In the backbone network CSPDarkNet and the neck network PAFPN of the YOLOv11 algorithm, the pre-specified convolutional layer is replaced by a wavelet pooling module, and multi-scale feature extraction of vehicle targets from the UAV perspective is realized through multi-band feature decomposition and downsampling operations; (C) The scalable receptive field RDS module is embedded in the C3K2 module of the YOLOv11 algorithm. Through multi-branch dilated convolution, residual structure, and depthwise separable convolution, the perception range of high-level features is extended and the features of shallow and deep layers are fused, improving the detection accuracy of small targets and alleviating the vanishing gradient; (D) In the target box screening stage of the YOLOv11 algorithm, the Soft-NMS algorithm is used to replace NMS. Through the soft suppression strategy, the high-confidence predictions in the overlapping boxes are retained, avoiding the missed detection of small targets due to hard-threshold filtering and improving the detection accuracy.

2. The lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle according to claim 1, characterized in that When performing (A), the newly added high-resolution detection layer for small targets specifically performs the following operations: The size of the deep feature map output by the backbone network CSPDarkNet is enlarged to 160×160×64 through three upsampling operations, so that small targets occupy more pixels on the feature map, retaining the detailed features of tire textures and headlight contours, and improving the localization accuracy and feature expression ability; At the same time, the number of channels of the backbone network CSPDarkNet is compressed, and the number of channels of the last three layers of the backbone network CSPDarkNet is reduced from 1024 to 512, reducing the number of parameters by reducing the thickness of the feature map and avoiding semantic dilution of small target information due to high-dimensional convolution in the deep network.

3. The lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle according to claim 2, wherein When performing (A), the shared lightweight detection head SLD realizes the dual optimization of detection performance and efficiency through the following operations: The number of channels of the input feature layer is adjusted by using 1×1 convolution, retaining key features while reducing the number of parameters; The processed feature layer is pooled into a shared convolution module, which consists of a 1×1 convolution for extracting inter-channel interaction features, a 3×3 convolution for capturing local context information, and a pooling operation for extracting multi-scale features, realizing the efficient extraction of multi-scale features; In the training stage, the independent branches of the shared lightweight detection head SLD act on the input feature map respectively: Classification branch: The target category probability is output through the formula Y1 = W1*X + b1, where X represents the input feature map, and W1 and b1 are parameters optimized through backpropagation during the training process; Regression branch: The bounding box coordinates are output through the formula Y2 = W2 * X + b2, where X represents the input feature map, and W2 and b2 are parameters optimized through backpropagation during the training process; Scale perception branch: The robustness to objects of different scales is enhanced through the formula Y3 = Pooling(X) + b3, where Pooling represents the pooling operation, X represents the input feature map, and b3 is a parameter optimized through backpropagation during the training process.

4. The lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle according to claim 3, wherein The operation (B) specifically includes: (B1) The wavelet pooling module parallelly processes the input feature map through a convolutional layer containing four wavelet filter kernels LL, LH, HL, and HH, where: the size of each wavelet filter kernel is k×k, and the depth matches the number of channels of the input feature map; the convolution operation adopts the grouped convolution mode, and the number of groups is set to the number of input channels to ensure that each channel of the input feature map independently uses all filters; (B2) The wavelet pooling module uses the predefined four wavelet filter kernels LL, LH, HL, and HH to decompose the input feature map into a low-frequency global feature map LL and high-frequency detail feature maps LH, HL, and HH in one go through a single convolution operation; (B3) The stride in the convolution operation is set to 2, reducing the size of the output feature map by 1 / 2, and realizing downsampling while decomposing multiple frequency bands; (B4) The pooling process of the wavelet pooling module is represented by the following formula: F out = Conv2D(F in , {h LL , h LH , h HL , h HH}, stride = 2), Where F in is the input feature map; F out is the decomposed multi-scale feature map, including low-frequency global features and high-frequency edge features; {h LL , h LH , h HL , h HH} are 4 wavelet filter kernels; (B5) The low-frequency global feature map LL extracts the global structural features of the image through low-pass filtering for subsequent high-level abstraction of semantic features. At the same time, the high-frequency detail feature maps LH, HL, and HH capture horizontal edges, vertical edges, and diagonal textures respectively through high-pass filtering, enhancing the sensitivity to the edge details of small objects; (B6) The low-frequency global feature map LL is input into the deep layer of the backbone network through PAFPN to participate in the high-level abstraction of vehicle semantic features, and the high-frequency detail feature maps LH, HL, and HH are fused with the shallow features of PAFPN to strengthen the expression of edge details of small objects.

5. A lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle according to claim 4, characterized in that, The size of each wavelet filter kernel is k×k, and the value is 3.

6. The lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle according to claim 4, wherein, The operation (C) specifically includes: (C1) Embed the expandable receptive field RDS module in the C3K2 module of YOLOv11 to replace the traditional post RFB (Receptive Field Block) structure, forming the C3K2_RDS module; (C2) The RDS module receives the input feature map passed from the C3K2 module. First, it performs a 3×3 per-channel convolution to independently perform convolution operations on each channel of the input feature map, and then performs a 1×1 pointwise convolution to integrate feature information through cross-channel linear combination; (C3) The RDS module adopts a three-parallel-branch structure, and the three branches use 3×3 depthwise separable convolutions with different dilation rates to extract features from different perspectives; (C4) After different dilation rate depthwise separable convolutions through three parallel branches, each branch outputs corresponding feature maps, which contain feature information of different scales and perspectives; the BN layer normalizes the feature maps output by each branch, making the mean of the feature maps 0 and the variance 1; after the BN layer processing, the normalized feature maps are subtracted from the input feature maps to generate semantic residuals; (C5) After feature extraction through three parallel branches and BN layer processing, multiple feature maps containing feature information of different scales and perspectives are obtained. The RDS module uses 1×1 pointwise convolution for feature fusion; at the same time, the 1×1 pointwise convolution adjusts the number of channels of the feature maps, fusing multiple feature maps into a comprehensive feature map to provide a richer and more effective feature representation for subsequent object detection tasks; (C6) After multi-branch feature extraction, BN layer processing, and 1×1 pointwise convolution fusion, a fused feature map is obtained. The RDS module directly superimposes the input feature map onto the fused output feature map in a residual connection manner.

7. A lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle according to claim 6, characterized in that, Execute step (C3). The RDS module adopts a three-parallel-branch structure. The three branches respectively use 3×3 depthwise separable convolutions with dilation rates of 1, 3, and 5: The first branch uses a 3×3 depthwise separable convolution with a dilation rate of 1, responsible for capturing local detail information in the feature map; The second branch uses a 3×3 depthwise separable convolution with a dilation rate of 3, enabling the convolutional kernel to span a set interval during convolution to expand the receptive field; The third branch uses a 3×3 depthwise separable convolution with a dilation rate of 5 to further expand the receptive field and capture more macroscopic global features.

8. A lightweight traffic vehicle detection method from the perspective of an unmanned aerial vehicle according to claim 6, characterized in that, The operation (D) specifically includes: (D1) Output all predicted target boxes and their corresponding confidence scores, and sort them from high to low confidence to generate an ordered list; (D2) Select the box with the highest current confidence from the sorted ordered list as the main box, and calculate its intersection over union with all the remaining boxes in the ordered list; (D3) Gradually decay the confidence of the overlapping boxes according to the size of the intersection over union, and retain the boxes with high confidence; (D4) Add the currently selected main box to the final detection result list and remove it from the original list, and then repeat steps (D2)-(D3) for the remaining boxes until the ordered list is empty. Finally, retain all the target boxes with confidence higher than the preset threshold as the detection result output.

Citation Information

Cited By

  • Urban vehicle detection method based on ORB-YOLO model

    CN121214399A

  • A method for urban vehicle detection based on the ORB-YOLO model

    CN121214399B

  • SAR ship detection method based on multi-scale edge information enhancement and lightweight decoupling detection head

    CN121259534A

  • Complex traffic scene-oriented small object and shelter detection method and system, electronic equipment and storage medium

    CN121811337A