Small target detection method in aerial photography by combining windmill convolution with weighted multi-branch fusion

By improving the YOLO11 network and combining it with the windmill convolution and WMFPN mechanism, the problem of low feature fusion efficiency in aerial photography small target detection is solved, and more efficient feature extraction and detection accuracy are achieved.

CN120510540BActive Publication Date: 2025-09-12SOUTHWEAT UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511004583.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-09-12
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Existing small target detection models in aerial images suffer from information redundancy and ineffective utilization in the feature fusion stage, resulting in reduced detection accuracy. In particular, it is difficult to extract effective features in complex backgrounds and dynamic environments.

Method used

The method of combining windmill convolution and weighted multi-branch fusion is adopted. By improving the YOLO11 network, the windmill convolution module and C3K2_HSD module are introduced to enhance the feature extraction capability, and the WMFPN mechanism is used for multi-scale feature fusion, and the learnable weights are used to improve the utilization of feature information.

Benefits of technology

It significantly improves the accuracy and robustness of small target detection in aerial photography, enhances the accuracy of feature extraction and detection precision, reduces computational costs, and adapts to the needs of multi-scale target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510540B_ABST
    Figure CN120510540B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting small aerial targets by combining windmill convolution with weighted multi-branch fusion, comprising the following steps: acquiring drone aerial images and constructing a training set and a validation set; constructing an aerial small target detection model based on an improved YOLO11; constructing an optimal aerial small target detection model based on the training set and the validation set; and inputting new drone aerial images into the optimal aerial small target detection model to obtain aerial small target detection results. The present invention utilizes a windmill convolution module and a C3K2_HSD module based on hidden state mixers and state-space duality to improve the backbone network of YOLO11, and designs a WMFPN mechanism based on learnable weights. This method can more effectively extract and utilize feature information of aerial small targets, improving overall detection accuracy. Experimental results demonstrate excellent detection performance in complex aerial scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of small target detection, and in particular relates to an aerial photography small target detection method combining windmill convolution with weighted multi-branch fusion. Background Art

[0002] With the rapid development and widespread application of drone technology, aerial imagery analysis is playing an increasingly important role in disaster relief, traffic monitoring, and other fields. Small object detection in aerial imagery, a key research area in computer vision, faces numerous technical challenges. First, in aerial photography, objects such as vehicles and pedestrians occupy a very small pixel fraction within the image, typically smaller than 32×32 pixels. This results in a significant loss of texture and edge information. This makes it difficult for traditional convolutional neural networks (CNNs), YOLO models, and Transformer models to extract discriminative features. Second, complex background interference, such as dense buildings and vegetation cover, combined with the low signal-to-noise ratio of small objects, further complicates feature extraction. Furthermore, changes in drone perspective, such as looking down or sideways, lead to significant differences in object scale, making a single-scale convolution kernel unsuitable for detecting multi-scale objects. Furthermore, dynamic environmental factors such as changing lighting, weather conditions, and occlusion introduce additional noise, reducing the robustness of feature representation and resulting in missed or under-detected objects.

[0003] Existing target detection models typically use multi-level feature fusion strategies, such as feature pyramid networks, to enhance the ability to represent targets at different scales. However, in the task of detecting small targets in aerial photography, the feature fusion process is often accompanied by the loss of effective information. On the one hand, as the network depth increases, the low-level detail features of small targets will gradually decay in the high-level feature maps, and conventional feature fusion methods have difficulty effectively retaining this key information. On the other hand, feature maps at different levels differ in spatial resolution and semantic information. Simple fusion operations may lead to feature misalignment, especially for small targets, where slight offsets can significantly affect detection accuracy. In addition, traditional feature fusion strategies usually treat all feature channels equally and fail to effectively suppress the interference of background noise, thereby reducing the significance of small target features.

[0004] Due to the unique characteristics of aerial photography, the target typically accounts for less than 0.1% of the pixels in the image and is often embedded in complex urban backgrounds, making it highly susceptible to interference from factors such as lighting variations and occlusions. These factors not only make it difficult for traditional convolutional neural networks to extract effective discriminative features, but also cause the loss of key features. Furthermore, existing detection networks suffer from redundant and inefficient feature information during the feature fusion stage, which not only increases computational costs but also significantly reduces detection accuracy. Summary of the Invention

[0005] In response to the above-mentioned deficiencies in the prior art, the aerial photography small target detection method provided by the present invention combines windmill convolution with weighted multi-branch fusion to solve the problem that the existing detection network has redundant and ineffective feature information utilization in the feature fusion stage, which significantly reduces the detection accuracy.

[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: a method for detecting small aerial targets by combining windmill convolution with weighted multi-branch fusion, comprising the following steps:

[0007] S1. Obtain drone aerial images and construct training and validation sets.

[0008] S2. Build an aerial small target detection model based on the improved YOLO11. The specific process is as follows: add the windmill convolution module to the backbone network of YOLO11, replace the C3K2 module of the backbone network with the C3K2_HSD module. The C3K2_HSD module is equipped with an HSM-SSD submodule based on the hidden state mixer and state space duality. The neck network uses the WMFPN mechanism to perform multi-scale feature fusion on the input feature map.

[0009] S3. Train the aerial small target detection model based on the training set. During the training process, use the validation set to tune the hyperparameters of the aerial small target detection model to obtain the optimal aerial small target detection model.

[0010] S4. Input the new UAV aerial image into the optimal aerial small target detection model to obtain the aerial small target detection result.

[0011] Furthermore: in S2, the aerial photography small target detection model includes a backbone network, a neck network, and a head network connected in sequence;

[0012] The backbone network includes a first windmill convolution module, a second windmill convolution module, a first C3K2_HSD module, a second C3K2_HSD module, a third C3K2_HSD module, a fourth C3K2_HSD module, an SPFF module and a C2PSA module connected in sequence;

[0013] The neck network includes first to fourth branches, the first branch includes a first Concat module, a first C3K2 module, a second Concat module, and a second C3K2 module connected in sequence, the second branch includes a third Concat module, a third C3K2 module, a fourth Concat module, and a fourth C3K2 module connected in sequence, the third branch includes a fifth Concat module, a fifth C3K2 module, a sixth Concat module, and a sixth C3K2 module connected in sequence, and the fourth branch includes a seventh Concat module and a seventh C3K2 module connected to each other;

[0014] The input of the first Concat module is connected to the output of the first C3K2_HSD module, the second C3K2_HSD module, and the third C3K2 module. The input of the second Concat module is also connected to the output of the third C3K2 module. The input of the third Concat module is connected to the output of the first C3K2_HSD module, the second C3K2_HSD module, the third C3K2_HSD module, and the fifth C3K2 module. The input of the fourth Concat module is also connected to the output of the first C3K2 module, the third C3K2 module, and the fifth C3K2 module. The input of the fifth Concat module is connected to the output of the second C3K2_HSD module, the third C3K2_HSD module, the C2PSA module, and the seventh C3K2 module. The input of the sixth Concat module is also connected to the output of the seventh C3K2 module. The input of the seventh Concat module is connected to the output of the third C3K2_HSD module and the C2PSA module.

[0015] The head network includes the first Detect module to the third Detect module, the input of the first Detect module is connected to the output of the second C3K2 module, the input of the second Detect module is connected to the output of the fourth C3K2 module, and the input of the third Detect module is connected to the output of the sixth C3K2 module.

[0016] Furthermore: In S2, the workflow of the windmill convolution module is as follows:

[0017] A1. For the input tensor of the pinwheel convolution module, perform asymmetric padding on the input tensor and perform staggered convolution on the first layer of the pinwheel convolution module to obtain the first to fourth tensors;

[0018] A2. Concatenate the first to fourth tensors and normalize the concatenated tensors to obtain the output of the pinwheel convolution module.

[0019] Further: In A1, the first tensor is obtained , the second tensor , the third tensor and the fourth tensor The specific expression is:

[0020]

[0021]

[0022]

[0023]

[0024] Where, is the convolution operator, is batch normalization, is the SiLU activation function, To fill in the parameters The result of asymmetric padding of the input tensor, To fill in the parameters The result of asymmetric padding of the input tensor, To fill in the parameters The result of asymmetric padding of the input tensor, To fill in the parameters The result of asymmetric padding of the input tensor, for The first convolution kernel, for The second convolution kernel, for The third convolution kernel, for The fourth convolution kernel, 、 and are the height, width, and number of channels of the output tensor after interleaved convolution, , , , 、 and are the height, width, and number of channels of the input tensor, respectively, and s is the convolution stride.

[0025] Further: In A2, the output of the windmill convolution module The specific expression is:

[0026]

[0027] Where, is the concatenated tensor, for The convolution kernel, 、 and are the height, width, and number of channels of the output of the pinwheel convolution module, respectively. , .

[0028] Further: In S2, the C3K2_HSD module includes a Split sub-module, a stacked HSM-SSD structure layer and a Concat sub-module connected in sequence, the stacked HSM-SSD structure layer includes several HSM-SSD sub-modules connected in sequence, and the Concat sub-module is connected to the output of the Split sub-module and all HSM-SSD sub-modules.

[0029] Furthermore, the workflow of the HSM-SSD submodule is as follows:

[0030] B1. For the input of the HSM-SSD submodule, use the importance weights to perform a weighted linear combination of the input states to obtain a shared global hidden state;

[0031] B2. Perform channel mixing based on the shared global hidden state, including gating and output projection, to obtain the output of the HSM-SSD submodule.

[0032] Furthermore: In B1, the expression for calculating the shared global hidden state is:

[0033]

[0034] Where, is the importance weight, For use vector, T is the transpose symbol, is the Hadamard product, B is the input state, It is the input of HSM-SSD submodule. is hidden state, is the input projection matrix.

[0035] Further: In B2, the output of the HSM-SSD submodule is obtained The specific expression is:

[0036]

[0037] Where, is the channel mixing of the gating function, C i is the projection matrix from state to output, is the output projection matrix, is a learnable matrix, is the activation function.

[0038] Furthermore, in S2, the neck network uses the WMFPN mechanism to perform multi-scale feature fusion on the input feature map as follows:

[0039] Align the input feature maps in the spatial dimension, perform channel weighting on the input feature images through the channel weight vector, and obtain the weighted fusion output;

[0040]

[0041] Where, is the kth input feature map, , , K is the total number of input feature maps, For Hadamard, is a matrix, N is the batch size, C k , H and W are the number of channels, height and width of the k-th input feature map respectively, is the k-th normalized channel weight;

[0042]

[0043] Where, is the kth learnable channel weight vector, A positive number less than 1 to prevent numerical instability caused by division by zero. is the j-th learnable channel weight vector, , , C is the total number of channels, and .

[0044] The beneficial effects of the present invention are:

[0045] (1) The present invention proposes a method for detecting small aerial targets by combining windmill convolution with weighted multi-branch fusion, which can more effectively extract and utilize the feature information of small aerial targets and improve the overall detection accuracy and robustness. Based on the YOLO11 algorithm as the basic network framework, from the perspective of strengthening feature semantic perception and efficiently extracting feature information of small aerial targets, the windmill convolution module and the C3K2_HSD module based on hidden state mixer and state space duality are used to improve the backbone network of the original network, enhance the underlying feature extraction capability, expand the receptive field and optimize the feature extraction of the target. On the basis of almost no increase in computational cost, the global dependency relationship is more effectively captured, the accuracy of feature extraction is further improved, and the accuracy of target detection is improved.

[0046] (2) The present invention designs a WMFPN mechanism based on learnable weights and adopts a dynamic weight allocation strategy to adaptively fuse features at different levels, significantly improving the effective utilization of feature information. It can fuse shallow features more effectively and in multiple dimensions, and transmit key low-level spatial information with higher quality, providing richer high-level fusion gradient information for detection, further enhancing the accuracy of target detection.

[0047] (3) Experimental results show that the proposed method exhibits excellent detection performance in complex aerial photography scenarios, especially in terms of small target detection accuracy and anti-interference ability. At the same time, it achieves an acceptable increase in computational complexity while maintaining a low number of parameters. Overall, the proposed method achieves a good balance between detection accuracy and computational efficiency, demonstrating its superiority and practical value in the task of small target detection in aerial photography. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1This is a flow chart of the aerial photography small target detection method combining windmill convolution and weighted multi-branch fusion of the present invention.

[0049] Figure 2 This is the improved network structure diagram.

[0050] Figure 3 This is the workflow of the pinwheel convolution module.

[0051] Figure 4 This is the structural diagram of the C3K2_HSD module.

[0052] Figure 5 Schematic diagram of the traditional feature fusion algorithm structure.

[0053] Figure 6 Schematic diagram of the WMFPN mechanism structure designed for the present invention. DETAILED DESCRIPTION

[0054] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0055] like Figure 1 As shown, in one embodiment of the present invention, a method for detecting small aerial targets by combining windmill convolution with weighted multi-branch fusion includes the following steps:

[0056] S1. Obtain drone aerial images and construct training and validation sets.

[0057] S2. Build an aerial small target detection model based on the improved YOLO11. The specific process is as follows: add the windmill convolution module (PConv) to the backbone network (Bottleneck) of YOLO11, improve the C3K2 feature extraction structure in the original network, propose the C3K2_HSD module, replace the C3K2 module of the backbone network with the C3K2_HSD module. The C3K2_HSD module is equipped with an HSM-SSD submodule based on the hidden state mixer and state space duality. The neck network uses the WMFPN mechanism to perform multi-scale feature fusion on the input feature map.

[0058] S3. Train the aerial small target detection model based on the training set. During the training process, use the validation set to tune the hyperparameters of the aerial small target detection model to obtain the optimal aerial small target detection model.

[0059] S4. Input the new UAV aerial image into the optimal aerial small target detection model to obtain the aerial small target detection result.

[0060] like Figure 2 As shown in S2, the aerial small target detection model includes a backbone network, a neck network (Neck), and a head network (Head) connected in sequence;

[0061] The backbone network includes a first windmill convolution module, a second windmill convolution module, a first C3K2_HSD module, a second C3K2_HSD module, a third C3K2_HSD module, a fourth C3K2_HSD module, an SPFF module and a C2PSA module connected in sequence;

[0062] The neck network includes first to fourth branches, the first branch includes a first Concat module, a first C3K2 module, a second Concat module, and a second C3K2 module connected in sequence, the second branch includes a third Concat module, a third C3K2 module, a fourth Concat module, and a fourth C3K2 module connected in sequence, the third branch includes a fifth Concat module, a fifth C3K2 module, a sixth Concat module, and a sixth C3K2 module connected in sequence, and the fourth branch includes a seventh Concat module and a seventh C3K2 module connected to each other;

[0063] The input of the first Concat module is connected to the output of the first C3K2_HSD module, the second C3K2_HSD module, and the third C3K2 module. The input of the second Concat module is also connected to the output of the third C3K2 module. The input of the third Concat module is connected to the output of the first C3K2_HSD module, the second C3K2_HSD module, the third C3K2_HSD module, and the fifth C3K2 module. The input of the fourth Concat module is also connected to the output of the first C3K2 module, the third C3K2 module, and the fifth C3K2 module. The input of the fifth Concat module is connected to the output of the second C3K2_HSD module, the third C3K2_HSD module, the C2PSA module, and the seventh C3K2 module. The input of the sixth Concat module is also connected to the output of the seventh C3K2 module. The input of the seventh Concat module is connected to the output of the third C3K2_HSD module and the C2PSA module.

[0064] The head network includes the first Detect module to the third Detect module, the input of the first Detect module is connected to the output of the second C3K2 module, the input of the second Detect module is connected to the output of the fourth C3K2 module, and the input of the third Detect module is connected to the output of the sixth C3K2 module.

[0065] In this embodiment, the present invention uses the YOLO11 algorithm as the basic network framework, focusing on optimizing feature extraction and feature fusion mechanisms. In terms of feature extraction, the windmill convolution module is innovatively introduced. Through its unique windmill-shaped convolution kernel structure, the model's feature capture capability for small aerial targets is effectively enhanced, and the receptive field is expanded to obtain richer contextual information. At the same time, the HSM-SSD module based on the hidden state mixer and state space duality is introduced to improve the C3K2 feature extraction structure in the original network, and the C3K2_HSD module is proposed to improve the backbone network part of the original network in a targeted manner and enhance the underlying feature extraction capability. Together with the windmill convolution module, it can more effectively capture global dependencies without increasing the computational cost, further improve the accuracy of feature extraction, and improve the accuracy of target detection.

[0066] In S2, the workflow of the pinwheel convolution module is as follows:

[0067] A1. For the input tensor of the pinwheel convolution module, perform asymmetric padding on the input tensor and perform staggered convolution on the first layer of the pinwheel convolution module to obtain the first to fourth tensors;

[0068] A2. Concatenate the first to fourth tensors and normalize the concatenated tensors to obtain the output of the pinwheel convolution module.

[0069] In this embodiment, the pinwheel convolution module is mainly used to address the narrow receptive field of the traditional convolution layer during feature extraction and the insufficient alignment of the Gaussian spatial distribution of pixels. The backbone network uses the pinwheel convolution module to enhance feature extraction and significantly increase the receptive field. Unlike standard convolution, the pinwheel convolution module uses asymmetric padding to create horizontal and vertical convolution kernels for different areas of the image. The convolution kernel spreads outward, such as Figure 3 As shown, the input tensor for the windmill convolution module is , 、 and are the height, width, and number of channels of the input tensor, respectively. To improve the stability and speed of training, the windmill convolution module applies batch normalization and SiLU activation function after each convolution to calculate the first tensor. , the second tensor , the third tensor and the fourth tensor .

[0070] In A1, we get the first tensor , the second tensor , the third tensor and the fourth tensor The specific expression is:

[0071]

[0072]

[0073]

[0074]

[0075] Where, is the convolution operator, is batch normalization, is the SiLU activation function, To fill in the parameters The result of asymmetric padding of the input tensor, To fill in the parameters The result of asymmetric padding of the input tensor, To fill in the parameters The result of asymmetric padding of the input tensor, To fill in the parameters The result of asymmetric padding of the input tensor, for The first convolution kernel, for The second convolution kernel, for The third convolution kernel, for The fourth convolution kernel, 、 and are the height, width, and number of channels of the output tensor after interleaved convolution, , , , s is the convolution step size.

[0076] In this embodiment, the padding parameter represents the number of padding pixels in the left, right, top and bottom directions, such as the padding parameter This means padding 1 pixel to the left and 3 pixels to the bottom.

[0077] The results of the first layer of interleaved convolution are concatenated, and the calculated concatenated tensor is output:

[0078]

[0079] Where, For the splicing operation. Through the convolution kernel Normalize the concatenated tensor without padding. The height and width of the output are adjusted to the preset values. and , making the pinwheel convolution module interchangeable with the standard convolution (Conv), and can be used as a channel attention mechanism to analyze the contribution of different convolution directions.

[0080] In A2, the output of the pinwheel convolution module The specific expression is:

[0081]

[0082] Where, is the concatenated tensor, for The convolution kernel, 、 and are the height, width, and number of channels of the output of the pinwheel convolution module, respectively. , .

[0083] The effectiveness of the receptive field in target detection gradually decreases outward, similar to a Gaussian distribution, and since the smaller the target, the more concentrated its features are, this highlights the importance of the central features. In this embodiment, the receptive field of the windmill convolution module is 25, and the number of convolutions decreases from the center outward, similar to a Gaussian distribution. The windmill convolution module uses grouped convolution to significantly increase the receptive field while minimizing the number of parameters. The parameters of the standard convolution are The specific expression is:

[0084]

[0085] Where k is the convolution kernel size, bias is the bias term in the convolution operation, and False is the removal of the bias term in the convolution operation.

[0086] Compared with the standard convolution, the parameters of the windmill convolution module are reduced by 22.2%, and the receptive field is increased by 177%, which is more conducive to the feature extraction of small targets in aerial photography. The specific expression is:

[0087]

[0088] Compared with standard convolution, the parameters of the windmill convolution module are reduced by 22.2%, and the receptive field is increased by 177%, which is more conducive to the feature extraction of small targets in aerial photography.

[0089] In S2, the C3K2_HSD module includes a Split submodule, a stacked HSM-SSD structure layer, and a Concat submodule connected in sequence. The stacked HSM-SSD structure layer includes several HSM-SSD submodules connected in sequence. The Concat submodule is connected to the outputs of the Split submodule and all HSM-SSD submodules.

[0090] In this embodiment, the present invention targets the feature extraction C3K2 module in the original YOLO11 network. Its structure is more suitable for local feature extraction, but it is weak in modeling global relationships and has difficulty in modeling the relationship between scattered objects. Therefore, the HSM-SSD module based on the hidden state mixer and state space duality is introduced to optimize the C3K2 module and propose the C3K2_HSD feature extraction module. The structure is as follows Figure 4 As shown in the figure, it can more effectively capture global dependencies without increasing computational costs, further improving the accuracy of feature extraction and the precision of target detection.

[0091] The specific workflow of the HSM-SSD submodule is as follows:

[0092] B1. For the input of the HSM-SSD submodule, use the importance weights to perform a weighted linear combination of the input states to obtain a shared global hidden state;

[0093] B2. Perform channel mixing based on the shared global hidden state, including gating and output projection, to obtain the output of the HSM-SSD submodule.

[0094] In this embodiment, the HSM-SSD submodule introduces a hidden state mixer (HSM) based on state space duality (SSD) to optimize computational efficiency, thereby reducing computational complexity and improving model performance. As can be seen from non-causal SSD (NC-SSD), by using importance weights on the input state, Perform a weighted linear combination to obtain a shared global hidden state h. At the same time, the output of each input is generated by using its corresponding mapped hidden state. If the projected input is represented by , we can get the expression for computing the shared global hidden state.

[0095] In B1, the expression for calculating the shared global hidden state is:

[0096]

[0097] Where, is the importance weight, For use vector, T is the transpose symbol, is the Hadamard product, B is the input state, It is the input of HSM-SSD submodule. is a matrix, , is hidden state, is the input projection matrix, By calculating the hidden state , which is then used as a linear projection to the hidden state, thus reducing the computational complexity from Reduce to , where N is the number of samples, L is the sequence length, and D is the number of channels. This optimization relies on the fact that N is much smaller than the number of channels D.

[0098] In this embodiment, the core idea of ​​the present invention is to use the shared global hidden state h to perform channel mixing, including gating and output projection, and operate directly on h, such as Figure 4 The HSM part can be used to obtain the output expression of the HSM-SSD submodule.

[0099] In B2, the output of the HSM-SSD submodule is obtained The specific expression is:

[0100]

[0101] Where, is the channel mixing of the gating function, C i is the projection matrix from state to output, , is the output projection matrix, , is a learnable matrix, is the activation function.

[0102] In this embodiment, by first calculating , which is then fed into a gating function, using HSM to apply gating and projection directly to the hidden state, and finally using Project the updated hidden state to generate the final output Therefore, the total complexity of capturing the global context in the HSM-SSD submodule becomes , which can be ignored as N becomes smaller.

[0103] In aerial photography scenes, due to the high shooting height, wide field of view, and small target, the target is easily affected by factors such as lighting changes and occlusion in a complex background, resulting in insufficient or lost feature information extraction. At the same time, during the feature fusion stage of the detection network, the redundancy and ineffective use of target feature information will increase the computational cost and reduce the detection accuracy. To solve this problem, the present invention designs a WMFPN mechanism to perform multi-scale feature fusion on the feature map of drone aerial images. The structure of the traditional feature fusion algorithm is as follows: Figure 5As shown in the figure, a large amount of information may be lost during the calculation process, and the interaction and fusion between layer information are poor. The WMFPN mechanism designed by the present invention can not only fuse shallow features more effectively and in multiple dimensions, but also transmit key low-level spatial information with higher quality, providing richer high-level fusion gradient information for detection, and further enhancing the target detection accuracy. The WMFPN mechanism structure designed by the present invention is shown in the figure. Figure 6 As shown, Figure 6 Weighted represents weighted feature fusion.

[0104] In S2, the neck network uses the WMFPN mechanism to perform multi-scale feature fusion on the input feature map as follows:

[0105] Align the input feature maps in the spatial dimension, perform channel weighting on the input feature images through the channel weight vector, and obtain the weighted fusion output;

[0106]

[0107] Where, is the kth input feature map, , , K is the total number of input feature maps, For Hadamard, is a matrix, N is the batch size, C k , H and W are the number of channels, height and width of the k-th input feature map respectively, is the k-th normalized channel weight;

[0108]

[0109] Where, is the kth learnable channel weight vector, A positive number less than 1 to prevent numerical instability caused by division by zero. is the j-th learnable channel weight vector, , , C is the total number of channels, and , the output of weighted fusion .

[0110] In this embodiment, the WMFPN mechanism introduces a learnable channel weight vector, assigning a trainable fusion coefficient to each input feature channel, explicitly modeling its importance and effectively improving the information expression and discriminative capabilities of multi-source features. Experiments show that this fusion strategy not only enhances the network's responsiveness to key features, but also possesses good flexibility and scalability, making it suitable for a variety of object detection and semantic segmentation tasks.

[0111] In S2, the neck network sends the weighted fused feature map to the head network through the WMFPN mechanism, and the head network outputs the final detection result.

[0112] In this embodiment, in order to verify the effectiveness of the aerial photography small target detection model based on the improved YOLO11 proposed in the present invention, the following experimental analysis is provided.

[0113] This example systematically trains and evaluates the VisDrone2019 dataset. This dataset, composed of images captured from drones, presents challenges such as small objects, severe occlusion, and complex backgrounds. It is widely used to evaluate the robustness of object detection algorithms in complex scenes. We use the official training set (train) for model training and the validation set (val) for model selection. Finally, performance is evaluated on the test set.

[0114] (1) Experimental environment:

[0115] This experiment was built on the deep learning framework Pytorch, and the Ubuntu system was used for related tests. The specific experimental configuration is shown in Table 1. The experiment in this embodiment uses the YOLO11s network model as a benchmark for improvement and training, and no pre-trained weights are used in the training process.

[0116] Table 1 Experimental environment configuration

[0117]

[0118] To ensure smooth convergence on both datasets and fully demonstrate the superiority of our method on different datasets, we set different training rounds for the two datasets while keeping other training hyperparameters consistent. We used the stochastic gradient descent (SGD) optimizer for 200 epochs, with a momentum of 0.937, a weight decay of 0.0005, a batch size of 16, and an initial learning rate of 0.01.

[0119] (2) Evaluation indicators:

[0120] To verify the performance of the model, this paper uses three evaluation metrics: precision, recall, and mean average precision (mAP). These metrics are calculated based on the confusion matrix. As shown in Table 2, TP indicates that the original data is a positive sample and is also a positive sample after the model prediction. FN indicates that the original data is a negative sample and is also a negative sample after the model prediction.

[0121] Table 2 Confusion Matrix

[0122]

[0123] Precision (P) is an important metric used to evaluate model performance in machine learning and deep learning. It is primarily used to measure the accuracy of the model's predictions for positive samples. Specifically, precision indicates how many of the results predicted by the model as positive samples are actually correct. The mathematical definition is as follows:

[0124]

[0125] Recall R, also known as sensitivity or True Positive Rate (TPR), is a metric that measures the model's coverage of positive examples. Specifically, recall represents the proportion of objects that the model correctly predicts as positive among all objects that are actually positive. The formula is as follows:

[0126]

[0127] mAP@50 is a commonly used evaluation metric in the field of object detection. It represents the average precision (AP) of a model across all categories at an Intersection over Union (IoU) threshold of 0.5. It comprehensively considers the model's accuracy in both object location detection and category prediction. A predicted box is considered valid only when its Intersection over Union (IoU) with the ground-truth box is ≥ 0.5.

[0128]

[0129] Among them, N d is the number of categories, is the AP of the i-th category, and the expression of the average precision AP is:

[0130]

[0131] in, is the changing function of precision P and recall R;

[0132] IoU measures the ratio of the intersection area and the union area of ​​the predicted box and the true box. A and 𝐵 represent the true box and the predicted box respectively. The expression is:

[0133]

[0134] (3) Performance comparison:

[0135] The experiment compared different versions of the model, including YOLOv5s, YOLOv8s, YOLOv9s, YOLOv10s, RT-DETR-R18, and our proposed method (Ours). As shown in Table 3, our proposed method (Ours) achieved the highest values ​​in precision (P), recall (R), and mAP@50, at 58.2%, 45.8%, and 48.3%, respectively, outperforming the other models overall.

[0136] Table 3 Comparative experiment

[0137]

[0138] 3.1) Quantitative analysis:

[0139] Comparing different models using five metrics—precision, recall, mAP@50, number of parameters (params), and number of GFLOPs—the proposed method achieved the highest precision, recall, and mAP@50, while maintaining a moderate increase in computational overhead. The parameter count was reduced by 8.7% compared to the original YOLO11s network, significantly lower than that of the original RT-DETR-R18 network. Furthermore, the mAP@50 was higher at the same computational overhead. Given the significant improvement achieved by the proposed method and its relatively low parameter count, the increased computational overhead is considered acceptable.

[0140] 3.2) Qualitative analysis:

[0141] To more intuitively demonstrate the comparative test results, the final detection images of the method of the present invention are compared with those of other advanced methods. It can be seen that when comparing the detection results of different models on the VisDrone2019 dataset, the detection box of the method of the present invention is closest to the true label, indicating that it has significant advantages in positioning accuracy and target recognition accuracy. In contrast, although other model methods can also detect targets, there are certain deviations in some samples, especially in complex backgrounds or when targets overlap. This further verifies that the method of the present invention can more accurately identify and locate targets, greatly reducing the false detection and missed detection rates.

[0142] In summary, the present invention addresses the difficulties in extracting features of small targets in aerial images, the low efficiency of conventional feature fusion, and the large amount of information lost during the calculation process, resulting in low final detection accuracy. A method for detecting small aerial targets that combines windmill convolution with multi-branch fusion is proposed. By introducing windmill convolution modules and C3K2_HSD modules with stronger semantic modeling capabilities and larger receptive fields to enhance the backbone network's feature extraction capabilities, and designing a WMFPN mechanism with learnable weights to improve feature fusion quality, the network's ability to detect small aerial targets is significantly enhanced.

[0143] In comparative experiments with various mainstream object detection models, the proposed method achieved optimal performance in precision (58.2%), recall (45.8%), and mAP@50 (48.3%), while maintaining a low parameter count and achieving an acceptable increase in computational complexity. Overall, the proposed method strikes a good balance between detection accuracy and computational efficiency, demonstrating its superiority and practical value in the task of detecting small objects in aerial photography.

[0144] In the description of the present invention, it should be understood that the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and cannot be understood as indicating or implying the relative importance or the number of technical features implicitly specified. Therefore, the features defined by "first", "second", and "third" may explicitly or implicitly include one or more of such features.

Claims

1. Aerial photography small target detection method combining windmill convolution and weighted multi-branch fusion, characterized by: The following steps are involved: S1. Obtain drone aerial images and construct training and validation sets. S2. Build an aerial small target detection model based on the improved YOLO11. The specific process is as follows: add the windmill convolution module to the backbone network of YOLO11, replace the C3K2 module of the backbone network with the C3K2_HSD module. The C3K2_HSD module is equipped with an HSM-SSD submodule based on the hidden state mixer and state space duality. The neck network uses the WMFPN mechanism to perform multi-scale feature fusion on the input feature map. S3. Train the aerial small target detection model based on the training set. During the training process, use the validation set to tune the hyperparameters of the aerial small target detection model to obtain the optimal aerial small target detection model. S4, input the new UAV aerial image into the optimal aerial small target detection model to obtain the aerial small target detection result; In S2, the workflow of the pinwheel convolution module is as follows: A1. For the input tensor of the pinwheel convolution module, perform asymmetric padding on the input tensor and perform staggered convolution on the first layer of the pinwheel convolution module to obtain the first to fourth tensors; A2. Concatenate the first to fourth tensors and normalize the concatenated tensors to obtain the output of the pinwheel convolution module. The specific workflow of the HSM-SSD submodule is as follows: B1. For the input of the HSM-SSD submodule, use the importance weights to perform a weighted linear combination of the input states to obtain a shared global hidden state; B2. Perform channel mixing based on the shared global hidden state, including gating and output projection, to obtain the output of the HSM-SSD submodule; In B1, the expression for calculating the shared global hidden state is: Where, is the importance weight, For use vector, T is the transpose symbol, is the Hadamard product, B is the input state, It is the input of HSM-SSD submodule. is hidden state, is the input projection matrix; In S2, the neck network uses the WMFPN mechanism to perform multi-scale feature fusion on the input feature map as follows: Align the input feature maps in the spatial dimension, perform channel weighting on the input feature images through the channel weight vector, and obtain the weighted fusion output; Where, For the k Input feature map, , , K is the total number of input feature maps, For Hadamard, is a matrix, N is the batch size, C k 、 H and W Respectively k The number of channels, height and width of the input feature map, For the k Normalized channel weights; Where, For the k Learnable channel weight vector, A positive number less than 1 to prevent numerical instability caused by division by zero. For the j Learnable channel weight vector, , , C is the total number of channels, and .

2. The aerial photography small target detection method combining windmill convolution and weighted multi-branch fusion according to claim 1 is characterized in that: In S2, the aerial small target detection model includes a backbone network, a neck network, and a head network connected sequentially; The backbone network includes a first windmill convolution module, a second windmill convolution module, a first C3K2_HSD module, a second C3K2_HSD module, a third C3K2_HSD module, a fourth C3K2_HSD module, an SPFF module and a C2PSA module connected in sequence; The neck network includes first to fourth branches, the first branch includes a first Concat module, a first C3K2 module, a second Concat module, and a second C3K2 module connected in sequence, the second branch includes a third Concat module, a third C3K2 module, a fourth Concat module, and a fourth C3K2 module connected in sequence, the third branch includes a fifth Concat module, a fifth C3K2 module, a sixth Concat module, and a sixth C3K2 module connected in sequence, and the fourth branch includes a seventh Concat module and a seventh C3K2 module connected to each other; The input of the first Concat module is connected to the output of the first C3K2_HSD module, the second C3K2_HSD module, and the third C3K2 module. The input of the second Concat module is also connected to the output of the third C3K2 module. The input of the third Concat module is connected to the output of the first C3K2_HSD module, the second C3K2_HSD module, the third C3K2_HSD module, and the fifth C3K2 module. The input of the fourth Concat module is also connected to the output of the first C3K2 module, the third C3K2 module, and the fifth C3K2 module. The input of the fifth Concat module is connected to the output of the second C3K2_HSD module, the third C3K2_HSD module, the C2PSA module, and the seventh C3K2 module. The input of the sixth Concat module is also connected to the output of the seventh C3K2 module. The input of the seventh Concat module is connected to the output of the third C3K2_HSD module and the C2PSA module. The head network includes the first Detect module to the third Detect module, the input of the first Detect module is connected to the output of the second C3K2 module, the input of the second Detect module is connected to the output of the fourth C3K2 module, and the input of the third Detect module is connected to the output of the sixth C3K2 module.

3. The aerial photography small target detection method combining windmill convolution and weighted multi-branch fusion according to claim 1 is characterized in that: In A1, we get the first tensor , the second tensor , the third tensor and the fourth tensor The specific expression is: Where, is the convolution operator, is batch normalization, is the SiLU activation function, To fill in the parameters The result of asymmetric padding of the input tensor, To fill in the parameters The result of asymmetric padding of the input tensor, To fill in the parameters The result of asymmetric padding of the input tensor, To fill in the parameters The result of asymmetric padding of the input tensor, for The first convolution kernel, for The second convolution kernel, for The third convolution kernel, for The fourth convolution kernel, 、 and are the height, width, and number of channels of the output tensor after interleaved convolution, , , , 、 and are the height, width, and number of channels of the input tensor, respectively. s is the convolution stride.

4. The aerial photography small target detection method combining windmill convolution and weighted multi-branch fusion according to claim 3 is characterized in that: In A2, the output of the pinwheel convolution module The specific expression is: Where, is the concatenated tensor, for The convolution kernel, 、 and are the height, width, and number of channels of the output of the pinwheel convolution module, respectively. , .

5. The aerial photography small target detection method combining windmill convolution and weighted multi-branch fusion according to claim 2 is characterized in that: In S2, the C3K2_HSD module includes a Split submodule, a stacked HSM-SSD structure layer, and a Concat submodule connected in sequence. The stacked HSM-SSD structure layer includes several HSM-SSD submodules connected in sequence. The Concat submodule is connected to the outputs of the Split submodule and all HSM-SSD submodules.

6. The aerial photography small target detection method combining windmill convolution and weighted multi-branch fusion according to claim 1 is characterized in that: In B2, the output of the HSM-SSD submodule is obtained The specific expression is: Where, is the channel mixing of the gating function, C i is the projection matrix from state to output, is the output projection matrix, is a learnable matrix, is the activation function.

Citation Information

Patent Citations

  • High-precision lightweight unmanned aerial vehicle image target detection algorithm based on double prediction heads

    CN119445414A

  • Remote sensing image small target detection method and device

    CN120339594A