Unmanned aerial vehicle remote sensing small target detection method based on expansion re-parameterization multidirectional feature pyramid network

By expanding the reparameterized multi-directional feature pyramid network (MD-DRIFPN), the difficult problem of small target detection in UAV remote sensing images is solved, and efficient and accurate small target detection is achieved, especially with significant improvements under complex background and low resolution conditions.

CN120726499APending Publication Date: 2025-09-30SHANDONG JIAOTONG UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510826233.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Small target detection in UAV remote sensing images faces the challenges of sparse target pixels, limited feature information, and complex and changing backgrounds. Existing methods have shortcomings in detection accuracy and efficiency.

Method used

The Dilated Reparameterized Multi-Directional Feature Pyramid Network (MD-DRIFPN) is adopted, including the Dilated Reparameterized Inverted Residual Module (DIMB), the Multi-Directional Scale-Aware Feature Pyramid Network (MD-SFPN) and the Lightweight Shared Convolutional Detection Head (LSDH), combined with the Focaler-CIoU loss function to enhance the accuracy of feature extraction, multi-scale fusion and bounding box regression.

Benefits of technology

The accuracy and efficiency of small target detection have been significantly improved, especially in complex backgrounds and low-resolution conditions. The AP50-95 index has been improved by 40.9%, the number of parameters has been reduced by 23.9%, and the bounding box regression accuracy has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726499A_ABST
    Figure CN120726499A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicle remote sensing small target detection, and particularly discloses an unmanned aerial vehicle remote sensing small target detection method based on an expansion re-parameterization multi-direction feature pyramid network, which comprises the following steps of: adopting an expansion re-parameterization inverse residual module as a backbone network, embedding 3 * 3 and 5 * 5 expansion convolution into inverse residual bottleneck, and obtaining an inverse residual bottleneck; the global-local coupling is dragged according to a 1: 1: 2 depth separation channel ratio, the global-local coupling is folded into single-branch convolution through re-parameterization in the reasoning stage, the receptive field is expanded, and the feature extraction capability is enhanced; a multi-direction scale perception feature pyramid network is used as a neck network. According to the unmanned aerial vehicle remote sensing small target detection method based on the expansion re-parameterization multi-direction feature pyramid network, double improvement of performance and efficiency is achieved, and compared with a traditional backbone network, the DIMB module remarkably enhances the feature representation capability of the small target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of UAV remote sensing small target detection, and in particular to a UAV remote sensing small target detection method based on an expanded reparameterized multi-directional feature pyramid network. Background Art

[0002] Small target detection in unmanned aerial vehicle (UAV) remote sensing images faces many challenges, including sparse target pixels, limited feature information, and complex and changing backgrounds. To address these challenges, this paper proposes a novel detection method called MD-DRIFPN (Multi-Directional Dilated Reparameterized Inverted Residual Feature Pyramid Network), which is specifically designed for the task of small target detection in UAV remote sensing images. The core innovations of this method include: (1) introducing the Dilated Reparameterized Inverted Residual Module (DIMB), which integrates multi-scale dilated convolution and dual attention mechanism to significantly expand the receptive field and enhance feature extraction capability; (2) designing the Multi-Directional Scale-Aware Feature Pyramid Network (MD-SFPN), which adopts bidirectional feature interaction and global heterogeneous convolution kernel selection mechanism to achieve effective multi-scale feature fusion; (3) proposing a lightweight shared convolutional detection head (LSDH), which optimizes detection accuracy through parameter sharing and label quality enhancement. In addition, this paper introduces the Focaler-IoU loss function to improve the bounding box regression accuracy and enhance the small target localization capability. Extensive experiments on the VisDrone and DIOR datasets demonstrate that MD-DRIFPN achieves significant improvements in small object detection metrics compared to existing state-of-the-art methods, reaching an AP50-95 score of 0.248, a 40.9% improvement over the YOLOv11s baseline while reducing the number of parameters by 23.9%. Ablation studies validate the effectiveness of the proposed components, particularly their robustness under complex backgrounds and low-resolution conditions. These results demonstrate that MD-DRIFPN provides an efficient and accurate solution for small object detection in UAV remote sensing. Summary of the Invention

[0003] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for detecting small targets in UAV remote sensing based on an expanded reparameterized multi-directional feature pyramid network, comprising the following steps:

[0004] The Dilated Reparameterized Inverted Residual Module (DIMB) is used as the backbone network. 3×3 and 5×5 dilated convolutions are embedded in the inverted residual bottleneck, with a 1:1:2 depth-wise separation channel ratio to drive global-local coupling. During the inference phase, the network is reparameterized and folded into a single-branch convolution, expanding the receptive field and enhancing feature extraction capabilities.

[0005] Using the Multi-Directional Scale-Aware Feature Pyramid Network (MD-SFPN) as the neck network, a bidirectional escalator pyramid is constructed through radial deconvolution and adaptive separable convolution kernel selection. Scale-invariant feature distribution gating is introduced to dynamically weight shallow details and deep semantics. A weakly supervised background suppression branch is designed to reduce the interference of ground texture and achieve multi-scale feature fusion.

[0006] The lightweight shared convolutional detection head (LSDH) is used to reduce detection head redundancy by leveraging inter-layer weight sharing and the Group Softmax strategy. The label-quality enhancement module is embedded in the category prediction end, and the swish-gated attention unit is combined to improve the discrimination of difficult samples and optimize detection accuracy.

[0007] The Focaler-CIoU loss function is introduced, and the Class-Balanced Focal Penalty is introduced based on the CIoU angle and distance constraints. It focuses on penalizing the errors of easily confused small targets, improving the bounding box regression accuracy, and enhancing the small target localization ability.

[0008] Preferably, the DIMB seamlessly combines "channel expansion-multi-branch convolution-attention-reparameterization" in an inverted residual bottleneck. The network first expands the number of input channels by about two times through 1×1 point-by-point convolution with BN and Swish activation, and is divided into three paths in a 1:1:2 ratio: branch A uses improved inverted residual+self-attention (iRMB), branches B1 / B2 respectively use 3×3 depth convolution with a dilation rate of 2 and 5×5 depth convolution with a dilation rate of 3. iRMB first uses depthwise separable convolution to capture local texture, then uses self-attention to aggregate long-range dependencies, and is supplemented by SE / CBAM for channel recalibration; the two dilated convolution branches extract context at different scales, and finally add them together to form branch B;

[0009] Subsequently, branches A and B are concat- ored in the channel dimension and compressed back to the original number of channels through 1×1 convolution to obtain fused features, which are then added element-by-element to the original input to form an inverse residual path. All multiple branches are retained during the training phase. During inference deployment, reparameterization technology is used to fold the 1×1 dilated convolution, depthwise convolution branches, and fused convolution weights into an equivalent single-core 3×3Depthwise+1×1Pointwise combination, and the computational graph is completely linearized.

[0010] Preferably, the MD-SFPN network starts from the three scale features P3, P4, and P5 output by the backbone, and first uses 1×1 convolution with BN and Swish activation to complete channel alignment, which are recorded as P3′, P4′, and P5′. Then, a top-down "semantic roll-up" pathway is established: the deep P5′ is upsampled by EUCB to obtain P4_up, and then another EUCB is used to continue upsampling to obtain P3_up. In parallel, the network also constructs a bottom-up "positioning downsampling" pathway: the shallow P3′ is first downsampled to P4_down by 3×3 depth-separable convolution, and P4_down is then downsampled in the same way to generate P5_down;

[0011] A Fusion node is set up at each scale to aggregate multi-source features. The Fusion node first weights each input feature channel by channel using a learnable channel weight, then performs Softmax normalization and sums the output. This allows the network to automatically measure the contributions of different scales and eliminates the redundant calculation of "Concat + 1×1 convolution".

[0012] Preferably, the overall process of the LSDH is "private Stem convolution → shared depth convolution stack → dual-branch prediction → label quality recalibration". The P3, P4, and P5 features input from the neck network first pass through a set of 1×1 convolution plus 3×3 convolution Stem layers to complete channel alignment and preliminary feature abstraction. Subsequently, the Stem outputs of the three scales enter a depth-separable convolution stack with fully shared parameters. The stack significantly compresses the detection head parameters while reusing the same filter across scales. The shared features are symmetrically bifurcated into classification branches and regression branches at each scale: the classification branch is connected to the Group Softmax classifier after the Swish-Gated attention unit, and the regression branch outputs a discrete bounding box distribution, which is then distributed through the Distributed Focal The Loss (DFL) decoder is reshaped into continuous bounding box coordinates, replacing the traditional four-dimensional offset regression to improve positioning accuracy. The outputs of the two branches are collaboratively recalibrated at the back end through the Label Quality Enhancement module (LQE) - this module dynamically weights the classification confidence based on positioning quality factors such as the IoU or GIoU of the predicted box. After non-maximum suppression, the prediction results of each scale are merged into the final detection box set.

[0013] Preferably, the Focaler-CIoU loss function is based on the basic IoU definition and modifies the traditional IoU calculation by rescaling the IoU values ​​using a normalization mechanism.

[0014] It has the following beneficial effects:

[0015] This method for detecting small targets in UAV remote sensing based on a dilated reparameterized multi-directional feature pyramid network, firstly, designs a DIMB backbone module, which effectively expands the network receptive field by integrating dilated reparameterized convolution and dual attention mechanism. During the training process, the module uses multi-scale dilated convolution to capture rich spatial context information, and in the inference stage, it converts it into a standard convolution through structural reparameterization technology, achieving a dual improvement in performance and efficiency. Compared with the traditional backbone network, the DIMB module significantly enhances the feature representation capability of small targets.

[0016] Secondly, the proposed MD-SFPN neck module achieves efficient multi-scale feature fusion under a lightweight design. Drawing on the design concept of BiFPN, this module introduces bidirectional feature interaction and adaptive weight fusion mechanism, optimizes the information flow of features of different scales, and cooperates with lightweight components such as CSP_MSCB and EUCB. While maintaining low computational cost, MD-SFPN significantly improves the detection ability of small targets, especially improving the detection effect of extremely small targets that account for less than 1% of the image.

[0017] Finally, the innovatively designed LSDH detection head improves detection accuracy while reducing model complexity through parameter sharing and label quality enhancement. Compared with traditional detection heads, LSDH not only has fewer parameters, but also achieves significant improvements in small object detection indicators. Combined with the optimization of the Focaler-CIoU loss function, it further improves the bounding box regression accuracy, especially in small object positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is the overall architecture diagram of yolov11 of the present invention;

[0019] Figure 2 This is the overall architecture of MD-DRIFPN of the present invention;

[0020] Figure 3 This is a structural diagram of the DIMB and related modules of the present invention;

[0021] Figure 4 Diagrams of the PAN-FPN structure and BiFPN structure of the present invention;

[0022] Figure 5 This is the EUCB structure diagram of the present invention;

[0023] Figure 6 This is a structural diagram of the MSCB of the present invention;

[0024] Figure 7 It is the structural diagram of MSDC of the present invention;

[0025] Figure 8 This is a structural diagram of the LSDH detection head of the present invention;

[0026] Figure 9 This is the structural diagram of the LQE module of the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0028] See also Figures 1-9 The present invention provides a technical solution: a method for detecting small targets in UAV remote sensing based on an expanded reparameterized multi-directional feature pyramid network, comprising the following steps:

[0029] The Dilated Reparameterized Inverted Residual Module (DIMB) is used as the backbone network. 3×3 and 5×5 dilated convolutions are embedded in the inverted residual bottleneck, with a 1:1:2 depth-wise separation channel ratio to drive global-local coupling. During the inference phase, the network is reparameterized and folded into a single-branch convolution, expanding the receptive field and enhancing feature extraction capabilities.

[0030] Using the Multi-Directional Scale-Aware Feature Pyramid Network (MD-SFPN) as the neck network, a bidirectional escalator pyramid is constructed through radial deconvolution and adaptive separable convolution kernel selection. Scale-invariant feature distribution gating is introduced to dynamically weight shallow details and deep semantics. A weakly supervised background suppression branch is designed to reduce the interference of ground texture and achieve multi-scale feature fusion.

[0031] The lightweight shared convolutional detection head (LSDH) is used to reduce detection head redundancy by leveraging inter-layer weight sharing and the Group Softmax strategy. The label-quality enhancement module is embedded in the category prediction end, and the swish-gated attention unit is combined to improve the discrimination of difficult samples and optimize detection accuracy.

[0032] The Focaler-CIoU loss function is introduced. Based on the CIoU angle and distance constraints, the Class-Balanced Focal Penalty is introduced to focus on penalizing errors of easily confused small objects, improve the bounding box regression accuracy, and enhance the small object localization capability.

[0033] Preferably, the DIMB seamlessly combines "channel expansion-multi-branch convolution-attention-reparameterization" in an inverted residual bottleneck. The network first expands the number of input channels by about two times through 1×1 point-by-point convolution with BN and Swish activation, and is divided into three paths in a 1:1:2 ratio: branch A uses improved inverted residual+self-attention (iRMB), branches B1 / B2 respectively use 3×3 depth convolution with a dilation rate of 2 and 5×5 depth convolution with a dilation rate of 3. iRMB first uses depthwise separable convolution to capture local texture, then uses self-attention to aggregate long-range dependencies, and is supplemented by SE / CBAM for channel recalibration; the two dilated convolution branches extract context at different scales, and finally add them together to form branch B;

[0034] Subsequently, branches A and B are concatenated in the channel dimension and compressed back to the original number of channels through a 1×1 convolution to obtain the fused features. These features are then element-wise added to the original input to form the inverse residual path. All multiple branches are retained during the training phase. During inference deployment, reparameterization technology is used to fold the 1×1 dilated convolution, depthwise convolution branches, and fused convolution weights into an equivalent single-core 3×3 depthwise + 1×1 pointwise combination, completely linearizing the computation graph.

[0035] The DIMB module can be represented as:

[0036]

[0037] in Indicates partial operations across stages, represents the inverse residual moving block function, It is the expansion reparameterization block operation;

[0038] iRMB is designed to efficiently utilize computing resources and achieve high accuracy while keeping the model lightweight. The main innovation of the iRMB module lies in the lightweight nature of the convolutional neural network (CNN) and the dynamic processing capabilities of the transformer model. This structure is particularly suitable for intensive prediction tasks on mobile devices because it aims to provide efficient performance in environments with limited computing power.

[0039] iRMB improves the handling of information flow through its inverted residual design, allowing it to capture and utilize long-range dependencies while keeping the model lightweight, which is critical for tasks such as image classification, object detection, and semantic segmentation. This design enables the model to run efficiently on resource-limited devices while maintaining or improving prediction accuracy;

[0040]

[0041] in is the self-attention operation, represents local convolution, It is squeeze-stimulate attention, is the projection operation, is the discard function;

[0042] The self-attention mechanism in iRMB can be expressed as:

[0043]

[0044] Where Q, K, and V are the query, key, and value projection of the input, and d is the dimension of the attention head;

[0045] The DRB module enhances the model's ability to handle larger receptive fields by adopting dilated convolutions, thereby improving the performance of complex or fine-grained tasks. This enables the model to better capture details and contextual information in the image, which is particularly beneficial for image processing tasks. It can be expressed as:

[0046]

[0047] in represents large kernel convolution, is batch normalization, Represents a kernel size k i and the expansion rate r i The dilated convolution, n is the number of parallel branches;

[0048] like Figure 3 As shown in the figure, by incorporating DRB into iRMB and integrating it into the C3k2 architecture, the model enhances the ability to capture fine details and contextual information from images, which is crucial for tasks that require accurate detection and recognition of objects. Integrating DRB in iRMB enables efficient attention mechanisms and spatial information processing, thereby promoting effective feature fusion within the baseline architecture, which ensures better integration of features at different scales to obtain a more comprehensive feature representation. This method enhances the model's ability to process features of different scales, which is crucial for object detection tasks because objects may appear in images at different sizes. By effectively processing features of different scales, the model improves the overall performance and accuracy of detecting objects of different sizes and complexities.

[0049] The MD-SFPN network starts with the three scale features P3, P4, and P5 output by the backbone. It first uses 1×1 convolution with BN and Swish activation to complete channel alignment, which are recorded as P3′, P4′, and P5′. Then it establishes a top-down "semantic up-convolution" pathway: the deep layer P5′ is upsampled by EUCB to obtain P4_up, and then further upsampled by another EUCB to obtain P3_up. In parallel, the network also constructs a bottom-up "positioning down-sampling" pathway: the shallow layer P3′ is first downsampled to P4_down through 3×3 depthwise separable convolution, and P4_down is then downsampled in the same way to generate P5_down;

[0050] A Fusion node is set up at each scale to aggregate multi-source features. The Fusion node first weights each input feature channel by channel using a learnable channel weight, then performs Softmax normalization and outputs the sum. This allows the network to automatically measure the contributions of different scales and eliminates the redundant calculation of "Concat + 1×1 convolution";

[0051] The MD-SFPN structure also improves the PAN-FPN structure by introducing new modules, including improving the C3k2 module into a multi-scale convolution module (MSCB) and improving the upsampling method into an efficient upconvolution module (EUCB);

[0052] The MD-SFPN structure is as follows Figure 4 As shown in b, it can be seen from the figure that, unlike the original PAN-FPN structure which tends to fuse features of the same scale, the MD-SFPN structure combines deep feature information with shallow feature information of the same level and high resolution, thereby retaining rich feature information. For example, in the Concat1 of the PAN-FPN structure, its input consists of the P4 layer of the same level and the P5 layer after upsampling. However, it ignores the importance of shallow positioning feature information in the P3 layer. In contrast, in the Fusion2 of the EMBSFPN structure, its input includes the P4 layer of the same level and the P5 layer after upsampling, as well as the P3 layer after downsampling. Therefore, in Fusion2, there are both shallow and high-resolution features of the same scale. There is also fusion between features, and fusion between multi-scale features of different resolution layers. In deeper network structures, such as Fusion4, information from four different layers can be fused at the same time, which greatly improves the performance of small target detection. The MD-SFPN structure fully utilizes the multi-scale information of different resolution layers in the above way, effectively improving the detection accuracy of the model. In addition, the MD-SFPN structure replaces the original channel splicing method Concat with the Fusion method. While reducing the number of parameters and the amount of calculation, it can also adaptively select weighted fusion according to the importance of features at different scales. In order to match this method, this paper performs a channel alignment operation before feature fusion, that is, using Figure 4Conv operation in b is used to reduce the number of channels;

[0053] EUCB uses efficient up-convolution blocks to upsample the feature map of the current stage step by step. This process can enhance the fusion of feature information between different levels and stages because it can make the dimension and resolution of the feature map of the current stage match the feature map of the next connection. The EUCB structure is as follows: Figure 5 As shown, from Figure 5 As can be seen from the figure, the EUCB structure first increases the scale of the input feature map by 2 times through upsampling operation, so that the model can capture more feature details. Secondly, 3×3 depth convolution is used to efficiently extract features and enhance expression ability without significantly increasing the amount of calculation. Then, BN and ReLU activation function operations are performed to speed up the network convergence and improve the nonlinear expression ability of the network. Finally, 1×1 convolution is performed to reduce the number of channels, thereby reducing the upsampled feature map to match the number of channels in the next stage.

[0054] The calculation formula for EUCB is as follows:

[0055] EUCB(x)=Conv 1×1 (ReLU(BN(DWC(Up(x)))))

[0056] like Figure 6 As shown in the figure, the structure of MSCB is similar to the inverted residual and linear bottleneck (IRB) structure in MobileNetV2. However, compared with the IRB structure, MSCB innovatively introduces the multi-scale depthwise convolution (MSDC) structure. The MSDC structure enables the model to perform depthwise convolution operations at multiple scales simultaneously, thereby extracting rich multi-scale features. In addition, MSCB also uses a channel shuffle operation, which enables the model to effectively integrate information at different scales, thereby greatly enhancing the expressive power of the feature map.

[0057] The MSDC structure introduces deep convolution kernels of sizes 5, 7, and 9. These convolution kernels of different sizes can adaptively match features of different scales and gradually obtain multi-scale feature information during the processing process, thereby improving the feature extraction ability of the model. The schematic diagram of the MSDC structure is shown in the figure. Figure 7 As shown in Figure 2, when the feature passes through the MSDC structure, the structure will first determine the most appropriate convolution kernel according to the scale of the feature, and then perform subsequent operations. Next, it will fuse the convolution results of these different scales, so that information-rich multi-scale features are finally formed.

[0058] The MSDC calculation formula is as follows:

[0059] MSCB(x)=BN(P2(CS(MSDC(σ(BN(P1(x))))))))

[0060] Where P represents the point-by-point convolution operation, σ represents the activation function, and the parallel MSDC(.) for different kernel sizes (K) can be expressed as MSDC(x)=∑ s∈K DWCB s (x) to indicate:

[0061] in:

[0062] DWCB s (x)=σ(BN(DWC s (x)))

[0063] Here, DWC s (·) is a depth-wise separable convolution with kernel size S, where BN and σ represent batch normalization and ReLU6 activation functions, respectively;

[0064] The resulting MD-SFPN significantly enhances its small target representation capabilities through multi-directional feature interaction and scale-aware fusion while maintaining its lightweight. Compared to the traditional PAFPN, this structure achieves more efficient feature fusion while maintaining similar parameter counts, demonstrating significant advantages in small target detection tasks in drone scenarios.

[0065] The overall process of LSDH is "private Stem convolution → shared depth convolution stack → dual-branch prediction → label quality recalibration". The P3, P4, and P5 features input from the neck network first pass through a set of 1×1 convolution plus 3×3 convolution Stem layers to complete channel alignment and preliminary feature abstraction. Subsequently, the Stem outputs of the three scales enter a depth-separable convolution stack with fully shared parameters. This stack significantly compresses the detection head parameters while reusing the same filter across scales. The shared features are symmetrically bifurcated into classification branches and regression branches at each scale: the classification branch is connected to the Group Softmax classifier after the Swish-Gated attention unit, and the regression branch outputs a discrete bounding box distribution, which is then distributed through the Distributed Focal The loss (DFL) decoder is reshaped into continuous bounding box coordinates, replacing traditional 4D offset regression to improve positioning accuracy. The outputs of the two branches are collaboratively recalibrated in the backend through the label quality enhancement module (LQE). This module dynamically weights the classification confidence based on positioning quality factors such as the predicted box's Intersection over Union (IoU) or GIoU. After non-maximum suppression, the prediction results at each scale are merged into the final detection box set.

[0066] The LSDH architecture introduces several key innovations to enhance small object detection capabilities. First, it uses a lightweight convolutional structure to process multi-scale feature maps while sharing parameters between different feature levels. The detection head processes the input through separate convolutional layers, followed by shared convolutional layers, creating a more efficient parameter usage pattern while maintaining detection accuracy.

[0067] The mathematical expression of feature processing in LSDH can be expressed as:

[0068]

[0069] where x i represents the input feature map of the i-th layer, n l is the number of detection layers, Conv i is a separate convolutional layer in layer i, and ShareConv is a shared convolutional module applied to all feature layers;

[0070] A key component of our design is the Label Quality Enhancement (LQE) module, which dynamically adjusts the contribution of each prediction based on the quality of the localization information. This innovation is particularly valuable for small object detection, where localization accuracy is often hampered by limited pixel information.

[0071] like Figure 9 , the LQE module operates as follows:

[0072]

[0073] in represents the predicted bounding box coordinates, represents the category score adjusted by the LQE module according to the positioning quality, P i is the final prediction output of the i-th layer;

[0074] During inference, the detection head processes these predictions through a distributed focal loss (DFL) decoder to produce the final bounding boxes:

[0075] B = dist2bbox(DFL(P box ),A,xywh=True)·S

[0076] Where B represents the decoded bounding box, P box It is the box prediction component, A represents the anchor point, and S represents the step value of each detection layer;

[0077] Our experiments show that the proposed LSDH detection head achieves significant improvements in small object detection metrics, especially in precision and recall for objects that occupy less than 1% of the image area in UAV-captured images.

[0078] The Focaler-CIoU loss function is based on the basic IoU definition and modifies the traditional IoU calculation by rescaling the IoU value using a normalization mechanism;

[0079] Focaler-IoU modifies the traditional IoU calculation by rescaling the IoU value using a normalization mechanism, which focuses the optimization process on a more efficient range. The mathematical expression of Focaler-IoU is defined as:

[0080]

[0081] Where d∈[0,1) represents the lower threshold parameter (default: d=0.0), u∈(d,1] represents the upper threshold parameter (default: u=0.95), and the function φ(·) ensures that the output value remains within the valid probability range. This transformation effectively remaps the IoU distribution, amplifying the gradient signal in the critical [d,u] interval while suppressing the range with less information;

[0082] The Focaler-IoU mechanism can be seamlessly integrated with various IoU variants to preserve their geometric advantages. For example, when combined with Complete IoU (CIoU), the composite loss function becomes:

[0083]

[0084] where ρ(b,b gt ) represents the predicted box b and the real box b gt The Euclidean distance between centroids, c represents the diagonal length of the minimum enclosing rectangle, v quantifies the aspect ratio consistency, and Δ serves as an adaptive trade-off parameter.

[0085] Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field and related fields without making creative efforts should fall within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described and explained in the present invention shall be implemented in accordance with conventional means in the field unless otherwise specified or limited.

Claims

1. A small target detection method for UAV remote sensing based on an expanded reparameterized multi-directional feature pyramid network, characterized by: The steps include: The dilated reparameterized inverted residual module (DIMB) is used as the backbone network. 3×3 and 5×5 dilated convolutions are embedded in the inverted residual bottleneck, with a 1:1:2 depth-wise separation channel ratio driving global-local coupling. During the inference phase, the network is reparameterized and folded into a single-branch convolution, expanding the receptive field and enhancing feature extraction capabilities. Using the Multi-Directional Scale-Aware Feature Pyramid Network (MD-SFPN) as the neck network, a bidirectional escalator pyramid is constructed through radial deconvolution and adaptive separable convolution kernel selection. Scale-invariant feature distribution gating is introduced to dynamically weight shallow details and deep semantics. A weakly supervised background suppression branch is designed to reduce the interference of ground texture and achieve multi-scale feature fusion. The lightweight shared convolutional detection head (LSDH) is used to reduce detection head redundancy by leveraging inter-layer weight sharing and the Group Softmax strategy. A label-quality enhancement module is embedded in the category prediction end, combined with the Swish-Gated attention unit to improve the discrimination of difficult samples and optimize detection accuracy. The Focaler-CIoU loss function is introduced, and the Class-BalancedFocal Penalty is introduced based on the CIoU angle and distance constraints. It focuses on penalizing the errors of easily confused small targets, improving the bounding box regression accuracy, and enhancing the small target positioning ability.

2. The method for detecting small targets in UAV remote sensing based on an expanded reparameterized multi-directional feature pyramid network according to claim x, characterized in that: The DIMB seamlessly combines "channel expansion - multi-branch convolution - attention - re-parameterization" in an inverted residual bottleneck. The network first expands the number of input channels by approximately two times through 1×1 point-by-point convolution with BN and Swish activation, and then splits it into three paths in a 1:1:2 ratio: Branch A uses the improved inverted residual + self-attention iRMB, branches B1 / B2 use 3×3 depth convolution with a dilation rate of 2 and 5×5 depth convolution with a dilation rate of 3 respectively. iRMB first uses depthwise separable convolution to capture local texture, then uses self-attention to aggregate long-range dependencies, and supplemented by SE / CBAM for channel recalibration; the two dilated convolution branches extract context at different scales and are finally added to form branch B; Subsequently, branches A and B are concat- ored in the channel dimension and compressed back to the original number of channels through 1×1 convolution to obtain fused features, which are then added element-by-element to the original input to form an inverse residual path. All multiple branches are retained during the training phase. During inference deployment, reparameterization technology is used to fold the 1×1 dilated convolution, depthwise convolution branches, and fused convolution weights into an equivalent single-core 3×3Depthwise+1×1Pointwise combination, and the computational graph is completely linearized.

3. The method for detecting small targets in UAV remote sensing based on an expanded reparameterized multi-directional feature pyramid network according to claim 1, characterized in that: The MD-SFPN network starts with the three scale features P3, P4, and P5 output by the backbone. It first uses 1×1 convolution with BN and Swish activation to complete channel alignment, which are recorded as P3′, P4′, and P5′. Then it establishes a top-down "semantic up-convolution" pathway: the deep layer P5′ is upsampled by EUCB to obtain P4_up, and then further upsampled by another EUCB to obtain P3_up. In parallel, the network also constructs a bottom-up "positioning down-sampling" pathway: the shallow layer P3′ is first downsampled to P4_down through 3×3 depthwise separable convolution, and P4_down is then downsampled in the same way to generate P5_down; A Fusion node is set up at each scale to aggregate multi-source features. The Fusion node first weights each input feature channel by channel using learnable channel weights, then performs Softmax normalization and sums the output. This allows the network to automatically measure the contributions of different scales and eliminates the redundant calculation of "Concat + 1×1 convolution".

4. The method for detecting small targets in UAV remote sensing based on an expanded reparameterized multi-directional feature pyramid network according to claim 1, characterized in that: The overall process of the LSDH is "private Stem convolution → shared depth convolution stack → dual-branch prediction → label quality recalibration". The P3, P4, and P5 features input from the neck network first pass through a set of 1×1 convolution plus 3×3 convolution Stem layers to complete channel alignment and preliminary feature abstraction. Subsequently, the Stem outputs of the three scales enter a depth-separable convolution stack with fully shared parameters. This stack significantly compresses the detection head parameters while reusing the same filter across scales. The shared features are symmetrically bifurcated into classification branches and regression branches at each scale: the classification branch is connected to the Group Softmax classifier after the Swish-Gated attention unit, and the regression branch outputs a discrete bounding box distribution, which is then distributed through the Distributed Focal The LossDFL decoder is reshaped into continuous bounding box coordinates, replacing the traditional four-dimensional offset regression to improve positioning accuracy. The outputs of the two branches are collaboratively recalibrated at the back end through the label quality enhancement module LQE - this module dynamically weights the classification confidence according to positioning quality factors such as the IoU or GIoU of the predicted box. After non-maximum suppression, the prediction results of each scale are merged into the final detection box set.

5. The method for detecting small targets in UAV remote sensing based on an expanded reparameterized multi-directional feature pyramid network according to claim 1, characterized in that: The Focaler-CIoU loss function is based on the basic IoU definition and modifies the traditional IoU calculation by rescaling the IoU values ​​using a normalization mechanism.

Citation Information

Cited By

  • Insulator ultraviolet corona discharge target detection method based on YOLO-SM

    CN120912997A

  • High-resolution remote sensing image road vehicle detection method based on feature fusion

    CN121033683A

  • River reach dike personnel intrusion intelligent identification method and device based on target detection

    CN121482722A

  • Substation metal expander top rushing detection method based on fine grit identification

    CN121564511A