An unmanned aerial vehicle aerial small target recognition method for urban scenes
Patent Information
- Application Number
- CN202611230551.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明的目的在于解决现有技术存在对小目标感知能力不足、聚焦能力有限、缺乏自适应尺度调整能力、定位精度低、检测结构复杂、计算开销较大的问题,提供一种面向城市场景的无人机航拍小目标识别方法,能够保证轻量化和实时性的前提下,同时增强多尺度局部特征提取、关键区域聚焦以及小目标定位精度
[0066](1)为增强主干网络对遥感图像中不同尺度目标的感知能力,本发明设计了PMDFU模块,该模块通过引入并行多尺度膨胀卷积,在不显著增加参数量的前提下扩大了感受野,使模型能够同时捕获局部细节与更大范围的上下文信息。相比传统单一路径卷积结构,PMDFU模块更适合处理遥感场景中尺寸变化大、目标分布密集的情况,能够有效缓解小目标特征在深层网络中逐渐衰减的问题。
Smart Images

Figure CN122841998A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method for small target recognition in drone aerial photography for urban scenarios. Background Technology
[0002] With their high mobility, flexible deployment, and unique top-down perspective, unmanned aerial vehicles (UAVs) have been widely used in infrastructure inspection, agricultural monitoring, and disaster detection, and are gradually becoming an important platform for intelligent sensing and automated decision-making. As the demand for low-altitude economics and intelligent inspection continues to grow, how to quickly and accurately identify targets from high-altitude images acquired by UAVs, images with frequently changing scales and complex backgrounds, has become an important research direction in the fields of computer vision and remote sensing image analysis. Unlike target detection from a ground perspective, targets captured by UAVs often exhibit characteristics such as small scale, high density, blurred edges, frequent occlusion, and strong background interference. Especially in long-distance shooting or large-area coverage situations, targets occupy only a very small number of pixels in the image, making their discriminative information extremely limited, posing significant challenges to feature extraction, semantic modeling, and spatial localization of detection models.
[0003] Traditional target detection methods primarily rely on manual feature extraction techniques such as Viola-Jones, HOG, and DPM for target localization and recognition. However, these methods have limited adaptability to complex backgrounds, scale variations, and target deformation, making it difficult to balance speed and accuracy. With the development of convolutional neural networks, two-stage detectors such as Faster R-CNN and Cascade R-CNN first locate the target and then extract features for recognition, improving detection accuracy. However, their computational processes are complex and their inference speed is slow, making them unsuitable for real-time UAV detection. In contrast, single-stage detectors like SSD and the YOLO series unify target classification and bounding box regression into an end-to-end framework, simultaneously performing both localization and classification. This results in higher inference efficiency and deployment potential. Among them, YOLOv8n, employing a more efficient feature extraction and decoupled detection approach, achieves a high balance between speed and accuracy, thus becoming an important baseline model for current research on improving small target detection in UAVs.
[0004] Despite the positive progress made by existing methods in improving small target detection on UAVs, many problems remain unresolved. Firstly, with the development of deep learning, more and more researchers are focusing on semantic feature extraction from deep networks, resulting in insufficient network perception of multi-scale targets. Shallow detail features are easily lost during layer-by-layer downsampling and feature fusion. Secondly, in complex backgrounds and low-contrast environments, existing models have limited ability to focus on target regions. Many attention mechanisms only enhance features from a single dimension, failing to effectively suppress background noise and highlight sparse target responses, leading to false positives and false negatives. Thirdly, most conventional detection heads are designed for general-scale targets, lacking fine-grained feature enhancement mechanisms and adaptive scale adjustment capabilities for small targets. In densely distributed, heavily occluded, and extremely small target scenarios, localization errors and classification confusion are likely to occur. Finally, while some high-performance methods improve detection accuracy, their complex structures and high computational costs make them difficult to meet the deployment requirements of real-time detection on UAVs and remote sensing devices. Therefore, how to simultaneously enhance multi-scale local feature extraction, key region focusing, and small target localization accuracy while ensuring lightweight design and real-time performance remains a critical problem to be solved. Summary of the Invention
[0005] The purpose of this invention is to address the problems of insufficient perception of small targets, limited focusing ability, lack of adaptive scale adjustment capability, low positioning accuracy, complex detection structure, and large computational overhead in existing technologies. This invention provides a method for small target recognition in UAV aerial photography for urban scenarios, which can enhance multi-scale local feature extraction, key area focusing, and small target positioning accuracy while ensuring lightweight and real-time performance.
[0006] To achieve the above-mentioned objectives, the embodiments of the present invention provide the following technical solutions:
[0007] A method for small target recognition in drone aerial photography for urban scenarios includes the following steps: acquiring remote sensing images and using a PMDED-YOLO remote sensing target detection model based on an improved YOLOv8n model to identify targets in the remote sensing images;
[0008] The PMDED-YOLO remote sensing target detection model introduces a multi-kernel perception module and a multi-kernel attention enhancement module; replaces the C2f module in the YOLOv8n model with a parallel multi-scale dilated receptive field module; and replaces the detection head in the YOLOv8n model with a detail enhancement detection head, which has four scales.
[0009] The internal processing flow of the parallel multi-scale dilatation receptive field module is as follows:
[0010] Input feature map X inSimultaneously, the input feature map X is fed into a multi-scale convolutional branch (MC) and a double depth dilation convolutional branch (DC); in the multi-scale convolutional branch (MC), the data passes through convolutional layers and depth convolutional layers before being combined with the input feature map X. in By adding element by element, we obtain the feature F. MC In the double-depth dilated convolution branch DC, after passing through deep convolutional layers with different dilation rates, the feature F is obtained. DC ;Feature F is learned through weights MC and feature F DC After fusion, channel calibration is performed, and finally the result is compared with the input feature X. in By adding elements one by one, we obtain the output feature map F. out .
[0011] Furthermore, the PMDED-YOLO remote sensing target detection model includes a backbone network;
[0012] Multi-core sensing modules are represented by MKP;
[0013] Parallel multiscale dilatational receptive field modules are represented by PMDFU;
[0014] The backbone network includes MKP-1, MKP-2, PMDFU-1, MKP-3, PMDFU-2, MKP-4, PMDFU-3, MKP-5, PMDFU-4, and SPPF connected in sequence.
[0015] Furthermore, the PMDED-YOLO remote sensing target detection model also includes a neck network;
[0016] Multi-core attention enhancement modules are denoted by MSCA;
[0017] The hierarchical connection relationship within the neck network is as follows: Feature map F4 output by SPPF is input into the MSCA module to obtain feature map F5; Feature map F5 is upsampled and fused with feature map F3 output by PMDFU-3, and the fused feature map is input into PMDFU-5 to obtain feature map F6; Feature map F6 is upsampled and fused with feature map F2 output by PMDFU-2, and the fused feature map is input into PMDFU-6 to obtain feature map F7; Feature map F7 is upsampled and fused with feature map F1 output by PMDFU-1, and the fused feature map is input into PMDFU-7 to obtain feature map F8; Feature map F8 is fused with feature map F7 after passing through CBS-1, and the fused feature map is input into PMDFU-8 to obtain feature map F9; Feature map F9 is fused with feature map F6 after passing through CBS-2, and the fused feature map is input into PMDFU-9 to obtain feature map F10; Feature map F10 is fused with feature map F5 after passing through CBS-3, and the fused feature map is input into PMDFU-10 to obtain feature map F11.
[0018] Furthermore, the PMDED-YOLO remote sensing target detection model also includes a head network;
[0019] The detail enhancement detection head is designated as DED-Head;
[0020] The head network includes DED-Head-1, DED-Head-2, DED-Head-3, and DED-Head-4; the feature map F8 output by PMDFU-7 is input to DED-Head-1, the feature map F9 output by PMDFU-8 is input to DED-Head-2, the feature map F10 output by PMDFU-9 is input to DED-Head-3, and the feature map F11 output by PMDFU-10 is input to DED-Head-4.
[0021] Furthermore, MKP consists of a series of convolutional layers, a 3×3 depth convolutional layer, a convolutional layer, a 5×5 depth convolutional layer, a convolutional layer, and a 7×7 depth convolutional layer. The output feature of MKP is the feature obtained by adding its input feature and the output feature of the 7×7 depth convolutional layer element by element.
[0022] Furthermore, the internal processing flow of PMDFU is as follows:
[0023] Input feature map X in Simultaneously feed in the multi-scale convolution branch MC and the double depth dilation convolution branch DC;
[0024] Input feature map X in In the multi-scale convolution branch MC, the input features first pass through a convolutional layer, then are fed into parallel 1×1, 3×3, and 5×5 convolutional layers. The outputs of each scale are then summed and fused to obtain the multi-scale feature map E. After passing through the convolutional layer, the multi-scale feature map E is then combined with the input feature map X. in Element-wise addition yields the output feature F of the multi-scale convolution branch MC. MC ;
[0025] Input feature map X in In the dual-depth dilated convolution branch DC, the output features F are obtained by sequentially passing through a 3×3 depth convolutional layer with dilation d=1, a 3×3 depth convolutional layer with dilation d=2, batch normalization, and the SiLU activation function. DC ;
[0026] The fused feature F is obtained by summing the elements of the learnable weights. fuse :
[0027]
[0028] Among them, w Aw B These are the learnable weights for the multi-scale convolution branch MC and the double depth dilation convolution branch DC, respectively.
[0029] Fusion feature F fuse After channel calibration by CBS-4, it is then compared with the input feature X. in By performing element-wise addition, the feature map F output by the PMDFU module is finally obtained. out :
[0030]
[0031] Where SiLU is the SiLU activation function; BN is batch normalization; Conv 1×1 It is a 1×1 convolutional layer; This is an element-wise addition.
[0032] Furthermore, the internal processing flow of MSCA is as follows:
[0033] Three parallel depthwise separable convolutional branches respectively process the input feature G in Perform depthwise convolution:
[0034]
[0035] Where k∈{3,5,7}, Y k For the output features corresponding to the depthwise separable convolution branches, k=3 represents the first depthwise separable convolution branch, k=5 represents the second depthwise separable convolution branch, and k=7 represents the third depthwise separable convolution branch; DWConv k×k A k×k depth convolutional layer; PWConv is a pointwise convolutional layer; BN is a normalization layer; SiLU is the SiLU activation function;
[0036] The feature Y output by three parallel depthwise separable convolution branches k The feature Y is obtained by averaging and fusing:
[0037]
[0038] Feature Y passes through a channel attention unit to generate channel attention weights A. c :
[0039]
[0040] Where W1 and W2 are the weights of the fully connected layer, and GAP is the average pooling layer; Use the Sigmoid activation function;
[0041] Feature Y is processed by a spatial attention unit to generate spatial attention weights A. s :
[0042]
[0043] Where Max is the max pooling layer; Concat is the fusion operation;
[0044] Feature map S is obtained after CBS-5:
[0045]
[0046] in, This means that convolution transforms the channel dimension from 2C to C, where C is the input feature G. in The number of channels; For element-wise multiplication;
[0047] MSCA's output characteristics G out for:
[0048]
[0049] Where α is the residual scaling factor.
[0050] Furthermore, the DED-Head comprises CGS-1, IDEConv-1, LCA-1, IDEConv-2, and LCA-2 connected in sequence;
[0051] Input feature X i Feature map T is obtained through CGS-1. out1 :
[0052]
[0053] Where SiLU is the SiLU activation function; GN is group normalization; Conv 1×1 It is a 1×1 convolutional layer;
[0054] Feature map T out1 Feature map T is obtained through IDEConv-1. out2 :
[0055]
[0056]
[0057]
[0058] Among them, W fused W represents the convolution kernel weights. cdc W adc W vdc W hdc These are the weights for the center, diagonal, vertical, and horizontal difference convolutions, respectively; Wfused For mixed bias; W 3×3 b represents the weights of a standard 3×3 convolution; cdc b adc b vdc b hdc These are the biases for the center, diagonal, vertical, and horizontal difference convolutions, respectively; b 3×3 The bias for a 3×3 standard convolution; For element-wise multiplication;
[0059] Feature map T out2 Feature map T is obtained through LCA-1. out3 :
[0060]
[0061] in, The Sigmoid activation function is used; GAP is an average pooling layer.
[0062] DED-Head's output characteristic T out4 for:
[0063]
[0064] Among them, IDEConv1 is the IDEConv-1 operation; LCA1 is the LCA-1 operation; IDEConv2 is the IDEConv-2 operation; and LCA2 is the LCA-2 operation.
[0065] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0066] (1) To enhance the backbone network's ability to perceive targets of different scales in remote sensing images, this invention designs the PMDFU module. This module expands the receptive field without significantly increasing the number of parameters by introducing parallel multi-scale dilated convolution, enabling the model to capture local details and broader contextual information simultaneously. Compared to traditional single-path convolution structures, the PMDFU module is more suitable for handling remote sensing scenes with large size variations and dense target distribution, and can effectively alleviate the problem of small target features gradually decaying in deep networks.
[0067] (2) To address the issues of complex backgrounds and weak target saliency in remote sensing images, this invention further introduces the MSCA module. This module, while expanding the receptive field, jointly extracts sparse attention information from both spatial and channel dimensions, and highlights key target regions and suppresses background noise interference through an efficient fusion mechanism. Compared with traditional single attention mechanisms, the MSCA module is more suitable for feature selection and enhancement in small target detection tasks, and can help the model more accurately locate target regions in low-contrast, densely distributed, and occluded scenes.
[0068] (3) Finally, the present invention designs a DED-Head detection head, which uses a cross-scale shared detail enhancement mechanism to specifically enhance the fine-grained information representation of extremely small targets, while taking into account parameter efficiency and computational overhead, thereby effectively improving the positioning accuracy and overall detection performance of small targets. Attached Figure Description
[0069] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0070] Figure 1 This is a network structure diagram of the PMDED-YOLO remote sensing target detection model of the present invention.
[0071] Figure 2 This is a network structure diagram of the Multi-Core Perception Module (MKP) module of the present invention.
[0072] Figure 3 This is a network structure diagram of the parallel multi-scale dilated receptive field (PMDFU) module of the present invention.
[0073] Figure 4 This is a network structure diagram of the Multi-Core Attention Enhancement (MSCA) module of the present invention.
[0074] Figure 5 This is a network structure diagram of the detail enhancement detection head (DED-Head) of the present invention.
[0075] Figure 6 The results of heatmap detection of partially enlarged images on the UC dataset using the YOLOv8n model in Example 2 and this method are shown below. Figure 6 (a) in the image is a UC sample image with real annotations. Figure 6 Image (b) is an enlarged view of the labeled portion of the UC sample. Figure 6 (c) in the image shows the heatmap results of the YOLOv8n model on a partially enlarged image. Figure 6 (d) in the figure represents the heatmap result of the model of this scheme for a partially enlarged image.
[0076] Figure 7 The results are heatmaps of the YOLOv8n model and this solution on the VisDrone 2019 test set from Example 2. Figure 7 Image (a) in the image is from the VisDrone 2019 test set. Figure 7 (b) in the image shows the heatmap results of the YOLOv8n model. Figure 7 (c) in the figure represents the heat map result of the model in this scheme.
[0077] Figure 8 The results are heatmaps of the YOLOv8n model and this scheme on the HazyDet test set from Example 2. Figure 8 (a) in the image is an image from the HazyDet test set. Figure 8 (b) in the image shows the heatmap results of the YOLOv8n model. Figure 8 (c) in the diagram represents the heat map result of this scheme. Detailed Implementation
[0078] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0079] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0080] Example 1:
[0081] This invention is achieved through the following technical solution: a method for identifying small targets in urban scenes using drone aerial photography, comprising the following steps: acquiring remote sensing images and using a PMDED-YOLO remote sensing target detection model based on an improved YOLOv8n model to identify targets in the remote sensing images.
[0082] This scheme improves upon the traditional YOLOv8n model to obtain the PMDED-YOLO remote sensing target detection model. For example... Figure 1As shown, the PMDED-YOLO remote sensing target detection model includes a backbone network, a neck network, and a head network. The main improvements are: 1) Introducing a multi-kernel perception (MKP) module into the backbone network to enhance the model's perception capabilities at multiple scales; 2) Replacing the original C2f module in the model with a parallel multi-scale dilated receptive field (PMDFU) module to enhance the model's feature extraction capabilities; 3) Introducing a multi-kernel attention enhancement (MSCA) module into the output of the backbone network, allowing the model to focus more on important feature regions and improve detection accuracy in detail; 4) Expanding the detection pyramid to four scales (P2-P5) in the head network, and combining a cross-scale shared detail enhancement module with various differential convolutions to form a detail enhancement detection head (DED-Head) to filter key features of small targets; 5) Introducing a scale dynamic loss function L... SDIoU The original CIoU loss function is optimized, and shape decoupling and dynamic weighting are used to enhance the location and category recognition of small targets. These improvements are made around four core objectives: multi-scale perception enhancement, key region focusing, detail feature enhancement, and lightweight design. This allows the PMDED-YOLO remote sensing target detection model to improve the detection accuracy and precision of dense, small, and multi-category targets while maintaining real-time inference speed.
[0083] For more details, please continue reading. Figure 1 The backbone network includes MKP-1, MKP-2, PMDFU-1, MKP-3, PMDFU-2, MKP-4, PMDFU-3, MKP-5, PMDFU-4, and SPPF connected in sequence.
[0084] The hierarchical connection relationship within the neck network is as follows: Feature map F4 output by SPPF is input into the MSCA module to obtain feature map F5; Feature map F5 is upsampled and fused with feature map F3 output by PMDFU-3, and the fused feature map is input into PMDFU-5 to obtain feature map F6; Feature map F6 is upsampled and fused with feature map F2 output by PMDFU-2, and the fused feature map is input into PMDFU-6 to obtain feature map F7; Feature map F7 is upsampled and fused with feature map F1 output by PMDFU-1, and the fused feature map is input into PMDFU-7 to obtain feature map F8; Feature map F8 is processed by CBS-1 and fused with feature map F7, and the fused feature map is input into PMDFU-8 to obtain feature map F9; Feature map F9 is processed by CBS-2 and fused with feature map F6, and the fused feature map is input into PMDFU-9 to obtain feature map F10; Feature map F10 is processed by CBS-3 and fused with feature map F5, and the fused feature map is input into PMDFU-10 to obtain feature map F11.
[0085] The head network includes DED-Head-1, DED-Head-2, DED-Head-3, and DED-Head-4; the feature map F8 output by PMDFU-7 is input to DED-Head-1, the feature map F9 output by PMDFU-8 is input to DED-Head-2, the feature map F10 output by PMDFU-9 is input to DED-Head-3, and the feature map F11 output by PMDFU-10 is input to DED-Head-4.
[0086] It should be noted that the "-1", "-2", etc. at the end of each level in the PMDED-YOLO remote sensing target detection model are only used to distinguish different levels with the same structure. For example, MKP-1 and MKP-2 belong to the same multi-kernel sensing module; PMDFU-1 and PMDFU-2 belong to the same parallel multi-scale dilated receptive field module; and CBS-1 and CBS-2 belong to the same convolutional unit, which are all composed of sequentially connected convolutional layers (Conv), batch normalization (BatchNorm), and SiLU activation functions. CBS is an existing technology and will not be described in detail here.
[0087] like Figure 2 As shown, the Multi-Kernel Perception Module (MKP) module includes a 1×1 convolutional layer (Conv), a 3×3 depth convolutional layer (DWConv), a 1×1 convolutional layer, a 5×5 depth convolutional layer, a 1×1 convolutional layer, and a 7×7 depth convolutional layer connected in sequence. The feature output by the MKP module is the feature obtained by adding its input feature and the output feature of the 7×7 depth convolutional layer element by element.
[0088] In UAV-generated remote sensing images of small targets, objects are densely packed, and there are overlapping or occlusion issues, posing a significant challenge to accurate target detection. To enhance the model's perception capabilities in complex backgrounds and multi-scale targets, this solution replaces the original C2f module with a parallel multi-scale dilated receptive field (PMDFU) module, such as... Figure 3 As shown, the core improvement lies in replacing the traditional bottleneck unit inside the C2f module with a dual-branch parallel fusion unit consisting of a multi-scale convolutional branch MC and a dual depth dilation convolutional branch DC. This improvement introduces a dual-branch parallel fusion unit and a learnable weighted fusion mechanism while maintaining the compatibility of the original structure of the YOLOv8n model, enabling the model to better capture small targets and edge and texture features that are easily disturbed by the background.
[0089] like Figure 3As shown, the Parallel Multi-Scale Dilated Receptive Field (PMDFU) module includes a dual-branch parallel fusion unit and a CBS-4; the dual-branch parallel fusion unit includes a multi-scale convolutional branch MC and a dual depth dilated convolutional branch DC; the multi-scale convolutional branch MC includes a 1×1 convolutional layer, a 1×1 depth convolutional layer, a 3×3 depth convolutional layer, a 5×5 depth convolutional layer, and a 1×1 convolutional layer; the dual depth dilated convolutional branch DC includes a 3×3 depth convolutional layer with dilation rate d=1, a 3×3 depth convolutional layer with dilation rate d=2, batch normalization, and the SiLU activation function.
[0090] Processing flow: Input the feature map X of the PMDFU module in The data is simultaneously fed into a multi-scale convolutional branch (MC) and a double depth dilation convolutional branch (DC). In the multi-scale convolutional branch (MC), the data first passes through a 1×1 convolutional layer to compress the channels and reduce computation. Then, it is fed into parallel 1×1 depth convolutional layers, 3×3 depth convolutional layers, and 5×5 depth convolutional layers. The outputs of each scale are added and fused to obtain a multi-scale feature map (E). This enables the joint extraction of local details and multi-receptive field textures, effectively alleviating the problem of detail loss caused by insufficient image pixels for small targets. This can be expressed by the formula:
[0091]
[0092] Among them, X in The feature map of the input PMDFU module; Conv 1×1 It is a 1×1 convolutional layer; DWConv k×k It is a k×k depth convolutional layer. .
[0093] The multi-scale feature map E is then passed through a 1×1 convolutional layer to compress the channels back to their original dimensions, and then connected to the input feature X via residual connections. in Element-wise addition yields the output feature F of the multi-scale convolution branch MC. MC :
[0094]
[0095] in, This is an element-wise addition.
[0096] The receptive field is expanded in the dual-depth dilated convolutional branch DC without loss of resolution and with a smaller number of parameters, capturing broader contextual information, enhancing global perception of complex scenes, and improving the recognition of small targets. The output feature F of the dual-depth dilated convolutional branch DC is obtained. DC :
[0097]
[0098] in, It is a 3×3 depth convolutional layer with an expansion ratio d=1; It is a 3×3 convolutional layer with an expansion ratio d=2; BN is batch normalization; SiLU is the SiLU activation function.
[0099] Then, at the output of the dual-branch parallel fusion unit, the features are output by element-wise weighted summation using learnable weights to achieve adaptive fusion and obtain the fused features F. fuse :
[0100]
[0101] Among them, w A w B These are the learnable weights for the multi-scale convolutional branch MC and the double depth dilation convolutional branch DC, respectively.
[0102] Next, the fusion feature F fuse The channel calibration is then performed using CBS-4, and the processed features are then connected to the input feature X via residual concatenation. in By performing element-wise addition, we finally obtain the feature map F output by the PMDFU module, which has richer semantic information and more prominent detailed regions. out :
[0103]
[0104] As network layers deepen, the detailed information in the feature maps becomes increasingly diluted, and small targets that need to be identified are more easily confused with objects in the background or with similar textures, leading to problems such as localization bias and missed detections. To improve this situation, this solution proposes a Multi-Core Attention Enhancement (MSCA) module, which aims to maximize the preservation of information richness in the feature maps after the feature extraction process of the backbone network to achieve efficient feature fusion and extraction. Figure 4As shown, the MSCA module includes three parallel depthwise separable convolutional branches, channel attention units, spatial attention units, a 1×1 CBS-5 layer, and a Drpout layer. The first depthwise separable convolutional branch consists of a 3×3 depthwise convolutional layer (DWConv), a 1×1 pointwise convolutional layer (PWConv), batch normalization, and a SiLU activation function, connected in sequence. The second depthwise separable convolutional branch consists of a 5×5 depthwise convolutional layer (DWConv), a 1×1 pointwise convolutional layer (PWConv), batch normalization, and a SiLU activation function, connected in sequence. The third depthwise separable convolutional branch consists of a 7×7 depthwise convolutional layer (DWConv), a 1×1 pointwise convolutional layer (PWConv), batch normalization, and a SiLU activation function, connected in sequence. The channel attention unit includes an average pooling layer (AvgPool), a 1×1 CBS-6, and a 1×1 CBSM-1 connected in sequence; the spatial attention unit includes an average pooling layer (AvgPool) and a max pooling layer (MaxPool) connected in parallel, and then a 3×3 CBS-7 and a 3×3 CBSM-2 connected in series; wherein both CBSM-1 and CBSM-2 include a convolutional layer (Conv), a batch normalization layer (BatchNorm), and a sigmoid activation function connected in sequence.
[0105] Processing flow: Three parallel depthwise separable convolutional branches process the features G of the input MSCA module respectively. in By performing depthwise convolution while maintaining the same number of channels, and then fusing channel information through pointwise convolution, the MSCA module achieves lightweight cross-channel multi-scale feature extraction, avoiding the problem of excessive computational cost of ordinary convolution. This can be expressed by the following formula:
[0106]
[0107] Among them, Y k The output features of the corresponding depth-separable convolutional branches; DWConv k×k The deep convolutional layer is a k×k convolutional layer, where k∈{3,5,7}; PWConv is a pointwise convolutional layer; BN is a normalization layer; and SiLU is the SiLU activation function. The kernel value of the deep convolutional layer is approximately logarithmically spaced, which can achieve sparse sampling from local to larger ranges within the same layer. When k=3, the first deep separable convolutional branch with a kernel size of 3×3 can capture the fine local texture of small objects; when k=5, the second deep separable convolutional branch with a kernel size of 5×5 enhances the ability to capture information of medium to large small objects; and when k=7, the second deep separable convolutional branch with a kernel size of 7×7 further enhances the discrimination of the surrounding environment.
[0108] The feature Y output by three parallel depthwise separable convolution branches kThe feature Y is obtained by averaging and fusing:
[0109]
[0110] Feature Y passes through channel attention units, where CBS-6 fully extracts fine-grained features of small targets, and CBSM-1 uses a sigmoid activation function to generate weights for feature selection. The combination of these two allows small targets to focus on salient features and textures in the channels, emphasizing the extraction of input feature G. in The channel attention unit generates channel attention weights A by extracting meaningful feature information and reducing the impact of complex image backgrounds on classification method training. c :
[0111]
[0112] Where W1 and W2 are the weights of the fully connected layer, and GAP is the average pooling layer; This is the Sigmoid activation function.
[0113] Feature Y is processed by the spatial attention unit, where average pooling and max pooling are performed along the channel dimension. After concatenation, the fully connected layers are replaced by 3×3 CBS-7 and 3×3 CBSM-2 layers, allowing the model to focus on effective pixel regions in the image and improving its ability to distinguish targets from similar backgrounds. This complements the channel attention unit. The spatial attention unit generates spatial attention weights A. s :
[0114]
[0115] Wherein, GAP is the average pooling layer; Max is the max pooling layer; and Concat is the fusion operation.
[0116] Next, the channel attention weight A c Spatial attention weight A s Then it is fused with feature Y, and then with input feature G. in Feature concatenation is performed; then, the concatenated features are processed using 1×1 CBS-5 to obtain a feature map S with the same characteristics as the original channels.
[0117]
[0118] in, This means that convolution transforms the channel dimension from 2C to C, where C is the input feature G. in The number of channels; This is for element-wise multiplication.
[0119] Finally, the MSCA module introduces a residual scaling factor α that can be dynamically adjusted during training and Dropout layer regularization to enhance the self-regulation and generalization capabilities of feature quality. Ultimately, the MSCA module outputs feature G. out :
[0120]
[0121] The initial value of the residual scaling factor α is 0.1.
[0122] Traditional YOLOv8n models use conventional 3×3 or 1×1 convolutions for their detection heads. Their feature fusion methods within the receptive field are not sensitive enough to detailed information such as edges, corners, and texture variations of target features. Small targets, occupying only a few dozen pixels, are easily smoothed or even obscured in deep networks. To better integrate scene information to assist the model in object detection, this solution proposes a Detail Enhancement Detection Head (DED-Head), such as... Figure 5 As shown.
[0123] First, to improve the ability to detect extremely small targets, the Detail Enhancement Detection Head (DED-Head) draws feature layers P2, P3, P4, and P5 from the backbone network, which serve as inputs to DED-Head-1, DED-Head-2, DED-Head-3, and DED-Head-4, respectively. This improves the model's accuracy in detecting extremely small targets smaller than 8×8 pixels. Let the set of input feature maps for the four feature layers be... L=4 represents the number of feature layers. C i H i W i These are the number of channels, height, and width, respectively. For example, when i=1, X1 is the feature of input DED-Head-1, i.e., the feature map F8 mentioned earlier; when i=2, X2 is the feature of input DED-Head-2, i.e., the feature map F9 mentioned earlier; when i=3, X3 is the feature of input DED-Head-3, i.e., the feature map F10 mentioned earlier; when i=4, X4 is the feature of input DED-Head-4, i.e., the feature map F11 mentioned earlier.
[0124] Taking any DED-Head as an example, the DED-Head consists of CGS-1, IDEConv-1, LCA-1, IDEConv-2, and LCA-2 connected sequentially. Input feature X i The feature map T will be obtained by processing with CGS-1. out1CGS-1 consists of sequentially connected convolutional layers (Conv), group normalization (GroupNorm), and SiLU activation function, expressed by the formula:
[0125]
[0126] Wherein, GN represents group normalization.
[0127] Next, to address the issue of standard convolutions in conventional detection heads being insensitive to local gradient changes and exhibiting numerous redundant parameters in the detail feature extraction part, this solution incorporates a detail enhancement convolution (IDEConv) module within the DED-Head. The IDEConv-1 module fuses four differential convolutions (central, diagonal, vertical, and horizontal) with a 3×3 standard convolution in the parameter space, enabling the final single convolution kernel to possess both intensity-aware and gradient-aware capabilities. In addition to weight fusion, the bias terms of each branch are also fused following the same linear superposition principle. This design significantly enhances the response to local gradient changes, reducing the inference computation to the level of a single standard convolution without sacrificing any detail-aware accuracy. The weights W of the fused convolution kernel are... fused and mixed bias b fused The calculation formula is as follows:
[0128]
[0129]
[0130] Among them, W cdc W adc W vdc W hdc These are the weights for the center, diagonal, vertical, and horizontal difference convolutions, respectively; W 3×3 b represents the weights of a standard 3×3 convolution; cdc b adc b vdc b hdc These are the biases for the center, diagonal, vertical, and horizontal difference convolutions, respectively; b 3×3 The bias is for a 3×3 standard convolution.
[0131] The IDEConv-1 module performs only one standard CGS-2 convolution operation during forward propagation, with the same number of parameters and computational cost as a regular 3×3 convolution. After the convolution kernels are fused, the feature map undergoes distribution calibration using group normalization within CGS-2 to mitigate the statistical instability of mini-batch training. Subsequently, a non-linear mapping is introduced through the SiLU activation function to further enhance the model's discriminative power for low-resolution small object contours. Finally, the IDEConv-1 module outputs feature map T. out2 :
[0132]
[0133] The IDEConv-1 module enhances the ability to perceive details in the spatial dimension, but the contributions of different channels to object detection vary significantly. To filter key channels and suppress background noise, this scheme designs a lightweight channel attention (LCA) module, with feature map T... out2 The LCA-1 module reweights the channel features using average pooling layers, 1×1 convolutional layers, and the Sigmoid activation function. The LCA-1 module outputs a feature map T. out3 :
[0134]
[0135] After that, feature map T out3 The process will continue through the IDEConv-2 and LCA-2 modules, with the same workflow as the IDEConv-1 and LCA-1 modules, which will not be repeated here. The two combined, stacked structures can be collectively referred to as the cross-scale shared detail enhancement module. All scales share the same set of parameters for IDEConv and LCA. This not only reduces the total number of parameters from L·P to P (where L is the number of feature layers, L=4, and P is the number of parameters per feature layer), but also provides effective regularization on small target datasets with limited training samples, mitigating overfitting. The output feature map T after passing through the cross-scale shared detail enhancement module is... out4 :
[0136]
[0137] Wherein, IDEConv1 represents the operation of the IDEConv-1 module; LCA1 represents the operation of the LCA-1 module; IDEConv2 represents the operation of the IDEConv-2 module; and LCA2 represents the operation of the LCA-2 module. out4 The feature map output by the LCA-2 module.
[0138] In the final output stage, to enhance the regression adaptability of DED-Head to targets of different scales, especially to address the localization difficulties caused by the narrow numerical range of offsets for small targets, this scheme introduces a lightweight learnable scale factor module. This module equips each feature layer with a learnable scalar parameter s. i (Initialized to 1.0), used to adaptively scale the output features of the bounding box regression branch to compensate for the narrow numerical range of small target offsets. Scaled output. for:
[0139]
[0140] Among them, R i The original bounding box regression feature map for the corresponding feature layer; This is the scaled bounding box regression feature map.
[0141] Example 2:
[0142] This embodiment presents experimental results based on Embodiment 1 above. To verify the effectiveness of the proposed PMDED-YOLO remote sensing target detection model, extensive validation was performed on the following datasets, and detailed evaluation results are presented, including performance comparisons of the overall and detail categories, as well as visualization results.
[0143] (a) Dataset and training environment.
[0144] Currently, urban management patrols for illegal construction heavily rely on drone aerial photography. Traditional manual image screening suffers from low efficiency, high false negative rates, and high labor costs. Most publicly available remote sensing datasets focus on large buildings and road segmentation, lacking specific labeled data for the detection of small, scattered, and irregularly shaped illegal colored steel sheds. This dataset, UC, contains 1184 drone aerial images, divided into 828 training images and 356 test images, all labeled using unified YOLO annotation. The annotation categories are: blue canopy sheds, other colored canopy sheds, and green shack sheds. The dataset covers typical scenes including suburbs, riverside residences in water towns, and farmland, providing highly practical real-world data support for target detection of small targets in complex rural backgrounds, effectively filling the data gap in existing publicly available remote sensing datasets for detecting illegal constructions in rural areas.
[0145] To further validate the detection performance of this scheme, training experiments were also conducted on large-scale datasets with different characteristics, namely VisDrone2019 and HazyDet. VisDrone2019 is a large-scale UAV-view dataset specifically designed for computational tasks. This dataset contains 10,209 still images, covering various scenes captured by UAVs in different urban and rural environments. The diversity of image types, high resolution, and complete object annotations, especially the fact that 60% of the object instances in this dataset are smaller than 20 pixels and 25% are between 20 and 30 pixels in size, place higher demands on small object detection methods. This experiment used 6,471 training images, 1,610 test images, and 548 validation images from this dataset. The labeled categories are 10: pedestrians (PS), crowds (PP), bicycles (BC), cars (CA), vans (VA), trucks (TU), tricycles (TC), awning tricycles (AT), buses (BU), and motorcycles (MT).
[0146] The HazyDet dataset is a large-scale dataset for drone-based hazy scenes, containing 383,000 annotated instances. These instances are collected from natural hazy environments and artificially added hazy effects in normal scenes. The creation process combines depth estimation and atmospheric scattering models, ensuring the data's authenticity and diversity. The dataset includes 8,000 training images, 1,000 validation images, and 2,000 test images, covering three common target categories: cars, trucks, and buses. Compared to datasets like VisDrone2019, which are taken under normal weather conditions, the HazyDet dataset provides high-quality data support for drone-perspective target detection in extreme weather conditions.
[0147] In this embodiment, all training processes were performed from scratch without using any pre-trained weights. The specific experimental environment included the pyTorch 2.0.1 framework, using an NVIDIA GeForce RTX 3090 GPU and CUDA 11.8 for training and inference testing. Training parameter settings were as follows: input image size was 640×640, initial learning rate was 0.1, learning rate decay coefficient was 0.001, image batch size was 32, training epochs were 300, the model also used the SGD optimizer, weight decay coefficient was 0.0005, and momentum was set to 0.937.
[0148] (ii) Evaluation criteria and loss function.
[0149] To quantitatively evaluate the performance of the proposed PMDED-YOLO remote sensing target detection model, this embodiment uses precision, recall, mean precision with an IoU threshold of 0.5, number of model parameters, and computational cost as the main evaluation metrics. The number of model parameters and computational cost can be considered together to measure the model's computational overhead. The number of parameters directly determines the model's size and is a key indicator for assessing the feasibility of model deployment. Computational cost measures the computational complexity of the model during training and inference; higher values indicate a more complex model structure and higher demands on hardware computing power. The formulas for precision (P) and recall (R) are shown below:
[0150]
[0151]
[0152] In this model, TP, FP, and FN represent the number of true positive, false positive, and false negative samples, respectively. A true positive represents the number of predicted boxes with an IoU greater than 0.5, a false positive represents the number of predicted boxes with an IoU less than or equal to 0.5, and a false negative represents the number of true labels not detected. It can be seen that precision calculates the percentage of correctly predicted positive samples out of all positive samples, thus reflecting the accuracy of the prediction results; recall, on the other hand, is the percentage of correctly predicted positive samples out of all positive samples, used to measure the model's ability to identify positive samples.
[0153] Combining precision and recall, the average precision (AP) for class c targets is... c The formula for calculating the average precision mAP50 with an intersection-union ratio of 0.5 across all categories is as follows:
[0154]
[0155]
[0156] Where U represents the total number of categories in the dataset. For brevity, P and R are abbreviations for precision and recall, respectively, and AP50 and AP are abbreviations for mAP50 and mAP50:95, respectively.
[0157] Since small bounding boxes occupy only 2×2 to 5×5 pixels in the feature map, the penalty term of the native CIoU loss function of the YOLOv8n model is too strict for small objects, while the scale dynamic loss function L... SDIoU The weights of the IoU term and the distance penalty term can be dynamically adjusted, which makes the model focus more on precise overlap rather than coarse location for small targets. The model's total loss function L total for:
[0158]
[0159] Among them, L BCE The binary cross-entropy loss function is... For L BCE Weights; L SDIoU For the scaling dynamic loss function, For L SDIoU Weights; L DFL For the distribution focus loss function, For L DFL The weight.
[0160] (III) Comparative Experiment.
[0161] 1. Overall performance of the UC dataset.
[0162] Based on the above experimental environment and evaluation indicators, the experimental results of the PMDED-YOLO remote sensing target detection model proposed in this scheme on the UC dataset are shown in Table 1 and Table 2.
[0163] Table 1. Detection results of YOLOv8n model and this solution on the UC dataset.
[0164]
[0165] Table 2. Detection results of YOLOv8n model and this solution on the UC dataset for various categories.
[0166]
[0167] The comparative test results on the UC dataset show that the proposed model has improved overall performance, especially in recall (R), which is 13.4% higher than the baseline YOLOv8n model. Other metrics, such as precision (P), AP50, and AP, are improved by 11.6%, 12.5%, and 11.5%, respectively. In other words, the PMDED-YOLO remote sensing target detection model has significantly enhanced its ability to suppress false positives and reduce false negatives.
[0168] Among the three categories of target detection, the "other colored covered greenhouses" category mostly consists of complex field debris with irregular shapes, varying sizes, and discrete features, making its detection a weakness of the baseline model. For conventional targets such as "blue covered greenhouses" and "greenhouses," the model achieved stable accuracy improvements. Particularly noteworthy was the significant 19.7% increase in AP50 for the "other colored covered greenhouses" category, which is characterized by complex features and is the most difficult to detect. This effectively compensated for the baseline model's poor performance in recognizing irregular samples. The PMDED-YOLO remote sensing target detection model achieved a comprehensive leap in detection accuracy at the cost of only a slight increase in parameters and computational cost, fully validating its effectiveness and superiority in target detection tasks in complex field scenes.
[0169] To more intuitively see the differences in small target detection performance, Figure 6 The results of heatmap detection using the YOLOv8n model and this solution on partially enlarged images on the UC test set are shown. Figure 6 (a) in the image is a UC sample image with real annotations. Figure 6 Image (b) is an enlarged view of the labeled portion of the UC sample. Figure 6 (c) in the image shows the heatmap results of the YOLOv8n model on a partially enlarged image. Figure 6 (d) in the figure represents the heatmap result of the model of this scheme for a partially enlarged image. Figure 6 In (a) and (b), the cyan, yellow, and green boxes represent the true labels of three targets: a blue covered greenhouse, covered greenhouses of other colors, and a greenhouse, respectively. The red dashed box represents the image area to be enlarged. As can be seen from the first and second rows of images, the YOLOv8n model suffers from missing small samples when detecting targets with large size differences. The third row of images shows that the YOLOv8n model also has false positives. In contrast, the proposed solution demonstrates a higher level of accuracy, especially in the heat map of the fourth row, where the proposed solution correctly detects far more targets than the YOLOv8n model.
[0170] 2. Overall performance comparison of VisDrone 2019 test set.
[0171] This embodiment compares the proposed model with other models on the VisDrone2019 dataset test set. These other models include LMAD-YOLO, LDSNet, ST-YOLO, FBRT-YOLO, MTD-YOLO, LSOD-YOLO, U-ShapeNet, RSW-YOLO, ECP-YOLO, Faster R-CNN, DIDO-R50, SCENET, RE-YOLO, and DFPF-YOLO. In addition to a comprehensive comparison of the overall test set performance, a comparison was also made for each category on the VisDrone2019 test set. The experimental results are shown in Tables 3 and 4.
[0172] Table 3 Performance comparison of this scheme with other models on the VisDrone 2019 test set.
[0173]
[0174] Table 4. Performance comparison of this scheme with other models on the VisDrone 2019 test set for each category.
[0175]
[0176] On the VisDrone 2019 test set, our model outperformed other models in most core metrics. Specifically, our model achieved precision and recall of 50.1% and 39%, respectively, which are 11.8% and 10.1% higher than the YOLOv8n model. Except for a slightly lower recall (R) compared to ECP-YOLO, our overall recognition ability is significantly better than other models. This simultaneous improvement in these metrics directly indicates that our model correctly identifies an increased number of real targets, representing a significant improvement in our overall target detection capability. Compared to methods with similar parameter counts, ST-YOLO and RSW-YOLO, our proposed model significantly outperforms them in AP50 and AP metrics, improving by 5.1% and 4.4% compared to ST-YOLO, and by 7.4% and 5% compared to RSW-YOLO. Among models with similar or higher computational costs, such as MTD-YOLO, LSOD-YOLO, and U-ShapeNet, our proposed model improves AP50 by 3.6%, 1.3%, and 3.8%, respectively. When both parameter count and computational cost are higher than our proposed model, ECP-YOLO only slightly improves recall (R) by 0.5%, with other core metrics lower. Among models with lower parameter count and computational cost, such as LMAD-YOLO, LDSNet, and FBRT-YOLO, their overall performance is significantly worse than our proposed model. These data strongly demonstrate that our proposed model achieves superior detection accuracy with lower parameter count and computational cost.
[0177] Table 4 details the comparison results of AP50 and GFLOPs for different categories of target data in VisDrone2019. As shown in the table, PMDED-YOLO's AP50 is as high as 38.3%, which is 11.6% higher than the baseline YOLOv8n, achieving the best detection accuracy among various algorithms. Compared with other models, PMDED-YOLO performs significantly in 7 categories: pedestrians (PS), cars (CA), vans (VA), trucks (TU), tricycles (TC), trolley tricycles (AT), and motorcycles (MT). It also achieves good results in categories such as crowds (PP), bicycles (BC), and buses (BU). In addition, the computational cost of this solution is only 20.1 GFLOPs, which is far lower than that of highly complex detection models such as MTD-YOLO, LSOD-YOLO, Faster R-CNN, and DIDO-R50. While comprehensively improving the detection accuracy of small targets, this solution effectively controls the computational cost of the model, achieving a good balance between detection accuracy and inference cost, and is more suitable for real-time monitoring tasks of small targets in complex scenarios.
[0178] This embodiment selects four different urban scenes ( Figure 7 Each row in the image represents a group. Image characteristics include blurriness, occlusion, dense targets, and complex environments. Figure 7 Image (a) in the image is from the VisDrone 2019 test set. Figure 7 (b) in the image shows the heatmap results of the YOLOv8n model. Figure 7 (c) in the diagram represents the heatmap result of this scheme model. (Through...) Figure 7 The visualization results show that the YOLOv8n model can only generate sparse and limited-coverage activation heatmaps for targets with large scale and significant features. In contrast, the heatmaps output by our proposed model can accurately capture small-scale, densely packed, and difficult-to-identify targets in low-light conditions. It generates activation responses that fit the target contours for small targets, dense crowds, and occluded targets in various distant images. The density and target matching of our heatmaps are also significantly higher than those generated by the YOLOv8n model. These results fully demonstrate that our model has significantly optimized its feature extraction and weak feature perception capabilities for multi-scale dense targets in complex aerial photography scenarios, making it more suitable for the practical application needs of drone aerial target detection tasks.
[0179] 3. Overall performance comparison of the HazyDet test set.
[0180] To evaluate the detection capability of the PMDED-YOLO remote sensing target detection model under low visibility conditions, this embodiment conducted relevant experiments on the HazyDet test set. The performance comparison results of this model with other models on the HazyDet test set are shown in Tables 5 and 6. Other models include FBRT-YOLO, MTD-YOLO, DETR-R18, DCM-DETR, CS-YOLO, DREAM, YOLO-GML, MTF-NET, and HAE-Net.
[0181] Table 5. Performance comparison of this scheme with other models on the HazyDet test set.
[0182]
[0183] Table 6. Performance comparison of this scheme with other models on the HazyDet test set for each category.
[0184]
[0185] As shown in Table 5, the quantitative comparison results indicate that the proposed model significantly outperforms the YOLOv8n model in precision, recall, AP50, and AP, exceeding them by 2.6%, 8.5%, 8.8%, and 9.1%, respectively. Compared to models with fewer parameters, such as FBRT-YOLO, CS-YOLO, and DREAM, the proposed model achieves higher AP50 scores by 9.6%, 3%, and 6.4%, respectively, and higher AP scores by 8.6%, 3.6%, and 5.4%, respectively. Compared to the second-best performing MTF-NET, the proposed model improves AP50 and AP by 0.9% and 1.8%, respectively. Meanwhile, precision... The accuracy (0.846) and recall (0.705) are both higher than those of MTF-NET. These data demonstrate that the proposed model has a stronger ability to capture and locate weak features and small targets with a very small proportion in scenarios such as foggy weather and aerial photography. In terms of computational cost, the proposed model has only 9.93M parameters and 20.1 GLOPs of computation, which is far lower than heavy models such as DETR-R18, DCM-DETR, and YOLO-GML. Compared with lightweight networks such as MTD-YOLO and CS-YOLO, it only slightly increases the computational cost but achieves a significant leap in accuracy. These results fully demonstrate the comprehensive competitiveness of the proposed model architecture in small target detection tasks with high accuracy and moderate computational cost.
[0186] As shown in the comparative experimental results in Table 6, the proposed model is the best among all the compared methods in AP50 for both car and truck categories, and its performance for bus category is close to that of MTF-NET. In addition, the proposed model is significantly better than YOLOv8 and other models in terms of aerial target recognition and localization for small vehicles. With a computational scale of only 20.1 GFLOPs, it ensures the potential for lightweight deployment while maintaining high accuracy.
[0187] Figure 8 The HazyDet test suite showcases four typical scenarios captured by aerial photography in haze (corresponding to four rows): open-air parking lot, daytime city streets, low-light and obstructed areas of buildings at night, and main roads at night. Figure 8 (a) in the image is an image from the HazyDet test set. Figure 8 (b) in the image shows the heatmap results of the YOLOv8n model. Figure 8(c) shows the heatmap results of this proposed solution. Visual comparison of the heatmaps reveals that the YOLOv8n model shows a significant response to large, clearly defined targets with minimal fog obscuring. However, it often struggles to identify targets obscured by thin fog, small distant targets, or targets with low visibility in low-light conditions, resulting in false positives. This indicates that the YOLOv8n model's feature extraction capabilities need improvement. In contrast, the proposed model can accurately locate most targets even in images with poor scene conditions. Even when faced with heavy fog obscuring, small distant targets, or blurred objects under low-light conditions at night, it can generate thermally activated regions that closely match the target outline and significantly enhance response intensity. These results demonstrate that the proposed model possesses stronger robustness and multi-scale target perception capabilities under complex weather and lighting conditions, providing reliable technical support for target detection tasks in aerial remote sensing imagery.
[0188] (iv) Ablation experiment.
[0189] To test the performance of each module, ablation experiments were conducted on the VisDrone2019 and HazyDet datasets in this embodiment, and the results are shown in Tables 7 and 8. The PMDFU, MSCA, and DED-Head modules were gradually integrated into the baseline YOLOv8n model to more specifically examine the contribution of each module. Tables 7 and 8 illustrate how the individual addition or combination of each module affects the evaluation metrics, where × indicates that the module is not added, and √ indicates that the module is integrated into the baseline YOLOv8n model.
[0190] Table 7 Ablation experiments on the VisDrone2019 dataset
[0191]
[0192] Table 8 Ablation experiments on the HazyDet dataset
[0193]
[0194] To avoid being constrained by mutual interference between modules, this embodiment introduces three modules based on the baseline YOLOv8n model to test the improvement effect of the model, so as to quantify the basic performance advantages of the modules. The YOLOv8n model achieves precision of only 0.383 and 0.799 on the complex remote sensing scene VisDrone2019 dataset and the haze-degraded image HazyDet dataset, respectively, with recall rates of only 0.289 and 0.62. This indicates that the model's feature extraction ability for small, dense, and blurred remote sensing targets is insufficient, and its ability in accurate localization and classification needs to be strengthened.
[0195] After introducing the PMDFU module alone, Tables 7 and 8 clearly show a significant improvement in all evaluation metrics. On the VisDrone2019 dataset, the accuracy reached 0.455, an improvement of 7.2%, and on the HazyDet dataset, it improved to 0.839, an improvement of 4%. Furthermore, the PMDFU module increased AP50 by 8.2% and 7.5% respectively. Overall, the PMDFU module showed the most significant improvement in the AP50 metric, indicating that it not only enhanced the accuracy of remote sensing target detection but also improved the localization quality under the IoU threshold. Under interference from environments such as congestion, occlusion, low-light haze, etc., the PMDFU module systematically improved the model's response capability to fine-grained targets.
[0196] The MSCA module demonstrates a different approach, enhancing attention and feature selection. It consistently improves precision, recall, AP50, and AP metrics on both datasets. Especially in VisDrone2019, with its dense array of small targets and complex backgrounds, the MSCA module effectively suppresses background noise and highlights key regions, enabling the model to detect targets more accurately.
[0197] Finally, DED-Head optimized the traditional detection head, improving AP from 0.153 to 0.208 on the VisDrone2019 dataset and from 0.483 to 0.549 on the HazyDet dataset. This indicates that DED-Head improves both regression quality and classification accuracy. In other words, DED-Head prioritizes making the model more stable, improving the YOLOv8 model from the perspectives of feature selection and detection head modeling, respectively, along with the MSCA module.
[0198] The ablation experiment results of multi-module combination show that the fusion of PMDFU, MSCA and DED_Head modules has significant synergistic optimization capabilities. Their combination further breaks through the performance limit of a single module and achieves a step-by-step improvement in evaluation indicators, which fully verifies the architectural rationality and complementarity of the proposed improvement method. Here, this paper selects several representative combination modules for specific analysis.
[0199] Firstly, the combination of the PMDFU and MSCA modules clearly demonstrates the synergistic effect between multi-scale feature enhancement and channel attention filtering. On the VisDrone2019 dataset, PMDFU+MSCA improves AP50 to 0.351, while PMDFU alone only achieves 0.349. On the HazyDet dataset, PMDFU+MSCA reaches 0.764, slightly lower than the 0.768 of the PMDFU module alone, but still significantly better than the YOLOv8n model's 0.693. This indicates that the introduction of the MSCA module does not simply bring linear gains, but rather enhances target-related information and suppresses invalid background noise by recalibrating the channel responses, thus playing a stabilizing role in more complex imaging environments.
[0200] Secondly, the combination of the PMDFU module and DED-Head demonstrates the complementary relationship between feature enhancement and the detection head. On the VisDrone2019 dataset, the AP50 of PMDFU+DED-Head reached 0.380, the highest among all dual-module combinations, representing an 11.3% improvement over the baseline. On the HazyDet dataset, the AP50 reached 0.784, with a precision of 0.845, second only to the combined effect of all three modules. The recall of 0.707 was even higher than the PMDED-YOLO remote sensing target detection model. These results indicate that, based on features enhanced by the PMDFU module, introducing DED-Head can effectively alleviate the localization bias problem of densely packed small targets and improve the target feature discrimination challenge caused by blurred image boundaries.
[0201] On the VisDrone2019 dataset, the fusion of the three modules resulted in AP50 and AP values of 0.383 and 0.226, respectively, representing improvements of 11.6% and 7.3% compared to the baseline. This demonstrates that the combined performance of different modules can continuously unleash performance potential. Although the model's FPS (frames per second) decreased compared to the YOLOv8n model, it remained within an acceptable real-time range, and the performance improvement significantly outweighed the speed loss of the YOLOv8n model, reflecting a good trade-off between accuracy and efficiency. On the HazyDet dataset, the fusion of the three modules also showed a stable increase compared to the baseline, indicating that the proposed improvement strategy is also effective for degraded images in foggy conditions. In particular, the proposed model maintains high performance in both precision and recall, reducing false positives without sacrificing recall, which is crucial for safety monitoring, disaster identification, or low-visibility target detection in remote sensing scenarios.
[0202] In summary, this invention proposes a method for small target recognition in UAV aerial photography for urban scenarios. Specifically, it includes three modules: PMDFU, MSCA, and DED-Head. The PMDFU module is used in the backbone network as a core component, effectively enhancing the ability to extract multi-scale local features and perceive contextual information, thus comprehensively improving the overall detection performance of the model. The MSCA module extracts sparse attention in both spatial and channel aspects and efficiently fuses them, significantly improving the detection accuracy of small targets. Finally, DED-Head introduces differential convolution, GroupNorm, lightweight channel attention, learnable scale factor, and multi-detector head extension to form a detection head improvement scheme specifically for small targets. The detail enhancement convolution and lightweight channel attention modules work together to dynamically emphasize the active channels of small targets and suppress background noise, significantly enhancing the contours and internal textures of small targets. The learnable scale factor allows the model to automatically adjust the scaling of regression values based on the feature layers, effectively improving the localization accuracy of small targets.
[0203] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for small target recognition in drone aerial photography for urban scenarios, characterized in that, The steps include: acquiring remote sensing images and using the PMDED-YOLO remote sensing target detection model, which is an improvement on the YOLOv8n model, to identify targets in the remote sensing images; The PMDED-YOLO remote sensing target detection model introduces a multi-kernel perception module and a multi-kernel attention enhancement module; replaces the C2f module in the YOLOv8n model with a parallel multi-scale dilated receptive field module; and replaces the detection head in the YOLOv8n model with a detail enhancement detection head, which has four scales. The internal processing flow of the parallel multi-scale dilatation receptive field module is as follows: Input feature map X in Simultaneously, the input feature map X is fed into a multi-scale convolutional branch (MC) and a double depth dilation convolutional branch (DC); in the multi-scale convolutional branch (MC), the input feature map X passes through convolutional layers and depth convolutional layers before being combined with the input feature map X. in By adding element by element, we obtain the feature F. MC In the double-depth dilated convolution branch DC, after passing through deep convolutional layers with different dilation rates, the feature F is obtained. DC ;Feature F is learned through weights MC and feature F DC After fusion, channel calibration is performed, and finally the result is compared with the input feature X. in By adding elements one by one, we obtain the output feature map F. out .
2. The method for small target recognition in UAV aerial photography for urban scenarios according to claim 1, wherein the PMDED-YOLO remote sensing target detection model includes a backbone network, characterized in that, Multi-core sensing modules are represented by MKP; Parallel multiscale dilatational receptive field modules are represented by PMDFU; The backbone network includes MKP-1, MKP-2, PMDFU-1, MKP-3, PMDFU-2, MKP-4, PMDFU-3, MKP-5, PMDFU-4, and SPPF connected in sequence.
3. The method for small target recognition in UAV aerial photography for urban scenarios according to claim 2, wherein the PMDED-YOLO remote sensing target detection model further includes a neck network, characterized in that, Multi-core attention enhancement modules are denoted by MSCA; The hierarchical connection relationship within the neck network is as follows: Feature map F4 output by SPPF is input into the MSCA module to obtain feature map F5; Feature map F5 is upsampled and fused with feature map F3 output by PMDFU-3, and the fused feature map is input into PMDFU-5 to obtain feature map F6; Feature map F6 is upsampled and fused with feature map F2 output by PMDFU-2, and the fused feature map is input into PMDFU-6 to obtain feature map F7; Feature map F7 is upsampled and fused with feature map F1 output by PMDFU-1, and the fused feature map is input into PMDFU-7 to obtain feature map F8; Feature map F8 is fused with feature map F7 after passing through CBS-1, and the fused feature map is input into PMDFU-8 to obtain feature map F9; Feature map F9 is fused with feature map F6 after passing through CBS-2, and the fused feature map is input into PMDFU-9 to obtain feature map F10; Feature map F10 is fused with feature map F5 after passing through CBS-3, and the fused feature map is input into PMDFU-10 to obtain feature map F11.
4. The method for small target recognition in UAV aerial photography for urban scenarios according to claim 3, wherein the PMDED-YOLO remote sensing target detection model further includes a head network, characterized in that, The detail enhancement detection head is designated as DED-Head; The head network includes DED-Head-1, DED-Head-2, DED-Head-3, and DED-Head-4; the feature map F8 output by PMDFU-7 is input to DED-Head-1, the feature map F9 output by PMDFU-8 is input to DED-Head-2, the feature map F10 output by PMDFU-9 is input to DED-Head-3, and the feature map F11 output by PMDFU-10 is input to DED-Head-4.
5. The method for small target recognition in UAV aerial photography for urban scenarios according to claim 2, characterized in that, MKP consists of a series of convolutional layers, a 3×3 depth convolutional layer, a convolutional layer, a 5×5 depth convolutional layer, a convolutional layer, and a 7×7 depth convolutional layer. The output feature of MKP is the feature obtained by adding its input feature and the output feature of the 7×7 depth convolutional layer element by element.
6. The method for small target recognition in UAV aerial photography for urban scenarios according to claim 3, characterized in that, The internal processing flow of PMDFU is as follows: Input feature map X in Simultaneously feed in the multi-scale convolution branch MC and the double depth dilation convolution branch DC; Input feature map X in In the multi-scale convolution branch MC, the input features first pass through a convolutional layer, then are fed into parallel 1×1, 3×3, and 5×5 convolutional layers. The outputs of each scale are then summed and fused to obtain the multi-scale feature map E. After passing through the convolutional layer, the multi-scale feature map E is then combined with the input feature map X. in Element-wise addition yields the output feature F of the multi-scale convolution branch MC. MC ; Input feature map X in In the dual-depth dilated convolution branch DC, the output features F are obtained by sequentially passing through a 3×3 depth convolutional layer with dilation d=1, a 3×3 depth convolutional layer with dilation d=2, batch normalization, and the SiLU activation function. DC ; The fused feature F is obtained by summing the elements of the learnable weights. fuse : Among them, w A w B These are the learnable weights for the multi-scale convolution branch MC and the double depth dilation convolution branch DC, respectively. Fusion feature F fuse After channel calibration by CBS-4, it is then compared with the input feature X. in By performing element-wise addition, the feature map F output by the PMDFU module is finally obtained. out : Where SiLU is the SiLU activation function; BN is batch normalization; Conv 1×1 It is a 1×1 convolutional layer; This is an element-wise addition.
7. The method for small target recognition in UAV aerial photography for urban scenarios according to claim 3, characterized in that, MSCA's internal processing flow is as follows: Three parallel depthwise separable convolutional branches respectively process the input feature G in Perform depthwise convolution: Where k∈{3,5,7}, Y k For the output features corresponding to the depthwise separable convolution branches, k=3 represents the first depthwise separable convolution branch, k=5 represents the second depthwise separable convolution branch, and k=7 represents the third depthwise separable convolution branch; DWConv k×k A k×k depth convolutional layer; PWConv is a pointwise convolutional layer; BN is a normalization layer; SiLU is the SiLU activation function; The feature Y output by three parallel depthwise separable convolution branches k The feature Y is obtained by averaging and fusing: Feature Y passes through a channel attention unit to generate channel attention weights A. c : Where W1 and W2 are the weights of the fully connected layer, and GAP is the average pooling layer; Use the Sigmoid activation function; Feature Y is processed by a spatial attention unit to generate spatial attention weights A. s : Where Max is the max pooling layer; Concat is the fusion operation; Feature map S is obtained after CBS-5: in, This means that convolution transforms the channel dimension from 2C to C, where C is the input feature G. in The number of channels; For element-wise multiplication; MSCA's output characteristics G out for: Where α is the residual scaling factor.
8. The method for small target recognition in UAV aerial photography for urban scenarios according to claim 4, characterized in that, The DED-Head consists of CGS-1, IDEConv-1, LCA-1, IDEConv-2, and LCA-2 connected in sequence. Input feature X i Feature map T is obtained through CGS-1. out1 : Where SiLU is the SiLU activation function; GN is group normalization; Conv 1×1 It is a 1×1 convolutional layer; Feature map T out1 Feature map T is obtained through IDEConv-1. out2 : Among them, W fused W represents the convolution kernel weights. cdc W adc W vdc W hdc These are the weights for the center, diagonal, vertical, and horizontal difference convolutions, respectively; W fused For mixed bias; W 3×3 b represents the weights of a standard 3×3 convolution; cdc b adc b vdc b hdc These are the biases for the center, diagonal, vertical, and horizontal difference convolutions, respectively; b 3×3 The bias for a 3×3 standard convolution; For element-wise multiplication; Feature map T out2 Feature map T is obtained through LCA-1. out3 : in, The Sigmoid activation function is used; GAP is an average pooling layer. DED-Head's output characteristic T out4 for: Among them, IDEConv1 is the IDEConv-1 operation; LCA1 is the LCA-1 operation; IDEConv2 is the IDEConv-2 operation; and LCA2 is the LCA-2 operation.