Infrared unmanned aerial vehicle target detection method based on efficient convolution and feature empowerment fusion

Through the infrared UAV target detection method of dual-branch hybrid convolution, multi-scale feature empowerment fusion and three-branch attention guidance module, the accuracy and real-time problems of UAV detection in complex background and low signal-to-noise ratio environment are solved, and efficient and accurate UAV target recognition is achieved.

CN120635757APending Publication Date: 2025-09-12YUNNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510923589.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing drone detection methods have limited performance in complex backgrounds and low signal-to-noise ratio environments. In particular, traditional methods have difficulty effectively detecting drone targets at night or in dim light conditions. Although deep learning methods have advantages, they consume large computing resources and cannot meet real-time and high efficiency requirements.

Method used

An infrared UAV target detection method based on efficient convolution and feature weighted fusion is adopted. Through two-branch hybrid convolution, multi-scale feature weighted fusion and three-branch attention guidance module, the feature expression ability and noise filtering effect are improved, the feature propagation path is optimized, and the detection accuracy and real-time performance are enhanced.

Benefits of technology

It significantly improves the accuracy and real-time performance of infrared UAV target detection, effectively filters noise, reduces missed detections, adapts to complex backgrounds and low signal-to-noise ratio environments, and reduces computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635757A_ABST
    Figure CN120635757A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and provides an infrared unmanned aerial vehicle target detection method based on efficient convolution and feature empowerment fusion, and the method comprises the steps: designing a dual-branch hybrid convolution in a feature extraction stage, and enabling the feature information of an infrared small target to be extracted more efficiently; a three-branch attention guiding module is designed in a feature fusion stage, and attention enhancement is performed on input features from two dimensions of space and channel. A three-branch parallel structure is adopted to carry out differentiation processing on input features, so that noise is effectively filtered, and target detail information is reserved. Besides, in order to further solve the problems of missing detection and the like, a multi-scale feature weighting fusion module is designed at a newly added detection head, the features are weighted in a self-adaptive mode, a feature propagation path is optimized, and therefore the detection precision of the model is improved. A large number of experimental results show that the provided method can effectively detect the small target of the unmanned aerial vehicle in real time, and advanced performance is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and in particular relates to an infrared UAV target detection method based on efficient convolution and feature weighting fusion. Background Art

[0002] With advances in computer, sensor, and communications technologies, drones have become a crucial vehicle for the convergence of these advanced technologies, significantly enhancing their performance and functionality. Currently, drones, due to their high maneuverability, compact size, and adaptability to terrain and environments, are rapidly gaining popularity in a wide range of applications, including military reconnaissance, aerial mapping, logistics and transportation, agricultural plant protection, disaster relief, and power inspection. Furthermore, their relatively affordable price, ease of operation, user-friendly interaction, and lightweight portability have led to explosive growth in the consumer drone market, attracting a large number of drone enthusiasts. However, the emergence of this emerging market has also brought with it a host of safety concerns, including incidents endangering public safety due to users' lack of awareness of safety and legal issues.

[0003] Therefore, research on drone countermeasures has been increasing in recent years. The foundation of a drone countermeasure system is the ability to successfully detect and identify drone targets, and then subsequently process them. In the field of drone detection, current researchers have found that due to their small size, drones lack distinct features in visible light images. Furthermore, when flying over complex terrain, drones are easily obscured by background obstacles, leading to missed or false detections. In dimly lit conditions, such as at night, some modified drones can almost blend in with the surroundings, resulting in a high number of missed detections. Given these factors, infrared thermal imaging is currently the primary method for drone detection. Drones generate high levels of thermal radiation during flight, which results in localized highlights in infrared images, particularly noticeable in low-contrast and high-noise environments. This characteristic makes infrared technology feasible for detecting small drone targets. It also overcomes the limitations of visible light in nighttime detection, facilitating the development of all-weather detection systems. Building on the principles of these two approaches to small drone target detection, some research has attempted to combine visible light detection with infrared detection. However, this integrated detection technology consumes more computing resources and increases model parameters, which limits the actual deployment and application of the algorithm model. In summary, infrared thermal imaging technology has become one of the mainstream technologies in the field of drone detection.

[0004] Existing research methods can be broadly categorized as traditional approaches and deep learning-based methods. Traditional methods primarily rely on image processing and mathematical modeling, offering some adaptability to specific scenarios but limited performance in complex backgrounds and high-noise environments. These methods primarily include filtering-based methods, methods based on the human visual system, and low-rank methods. While these traditional methods achieved some success in early research and applications, their limitations are becoming apparent in the face of increasingly complex scene changes and diverse object morphologies.

[0005] In contrast, deep learning-based methods improve detection performance by automatically learning multi-level features from images, effectively overcoming the limitations of traditional methods. Deep learning methods demonstrate enhanced robustness and generalization, particularly in complex backgrounds and low signal-to-noise ratio environments. Currently, researchers are primarily improving the network architecture of advanced deep learning models, such as convolutional neural networks (CNNs), transformers, and generative adversarial networks (GANs), to adapt them to the task of infrared small target detection. For example, multi-scale feature fusion strategies are employed to enhance small target detection capabilities by integrating feature information from different levels. Furthermore, the introduction of attention mechanisms to further enhance the model's focus on the target area is also a mainstream improvement approach. Furthermore, some research has enhanced the model's ability to perceive small infrared targets by incorporating prior knowledge into deep learning frameworks, adjusting loss functions, and considering target shape, resulting in even better performance in complex backgrounds and low signal-to-noise ratio environments.

[0006] Overall, deep learning methods have demonstrated significant advantages in the field of infrared small target detection and are gradually replacing traditional methods as the mainstream approach. Current research will further explore lighter and more efficient network structures to improve the real-time performance and accuracy of detection while reducing computing resource consumption, thereby better adapting to practical application scenarios. Summary of the Invention

[0007] To solve the above technical problems, the present invention provides an infrared UAV target detection method based on efficient convolution and feature weighted fusion to solve the problems in the prior art. The technical solution adopted by the present invention is:

[0008] Infrared UAV target detection method based on efficient convolution and feature weighted fusion, including:

[0009] Step S0: Use an infrared camera to capture the target image to obtain a dataset of small infrared drone targets and preprocess the image;

[0010] Step S1: label the data set and then send it to the network backbone for feature extraction;

[0011] Step S2: The extracted multi-scale features are fed into the neck network for fusion, and the output layer is selected as the input feature of the detection head;

[0012] Step S3: The detection head performs target discrimination on the input features;

[0013] Step S4: Based on the prediction results, calculate the positioning loss and confidence loss to obtain the total loss; train the model by optimizing the total loss;

[0014] Step S5: Repeat steps S0-S4 to make the network weights converge to the optimum and obtain the best weight file, thereby realizing the detection of small drone targets.

[0015] Furthermore, in step S0, the image is preprocessed including: adjusting the image to a uniform size, normalizing the pixel values, and data augmentation to improve the generalization ability of the model.

[0016] Furthermore, step S1 includes: extracting multi-scale feature information from the input infrared image through the Conv module, C3 module and DBHConv module in the Backbone of the network;

[0017] The DBHConv module adopts a dual-path structure design. In the first branch, the input features are first evenly divided into two parts X1 and X2 along the channel dimension. X1 is downsampled by 3×3 convolution to mine the local detail features of small targets; X2 undergoes a 3×3 maximum pooling operation to retain the information of the highlight area and focus on the possible target area; the two are spliced ​​and output F1; in the second branch, the complete input feature X is downsampled by a 3×3 convolution, and then the branch uses a 1×1 depth-separable convolution to model the information between channels to capture a more discriminative feature distribution, and uses jump connections to further fuse the features of the previous step and output it as F2; ​​finally, the output features F1 and F2 of the two trunk branches are spliced ​​to form an enhanced feature representation F out :

[0018]

[0019] F1=Concat(Conv(X1),Conv(Poll(X2))),

[0020] F2=Concat(DWConv(Conv(X)),Conv(X)),

[0021] F out =Concat(F1,F2).

[0022] Furthermore, step S2 includes: the fused features are processed by a three-branch attention guidance module before being sent to the detection head, which extracts features through three parallel branches of different depths and uses an attention mechanism to focus on key areas in the image;

[0023] The first branch first applies a 1×1 convolution to the input features for channel adjustment, and then enters the Bottleneck structure, which includes a 1×1 convolution for channel compression and a 3×3 convolution for local area feature extraction. The structure then passes through the CASATT module to enhance the target area features in both channel and spatial dimensions. Finally, the branch output is added to the current 1×1 convolution input features through a residual connection, and deep features are extracted through this iterative process.

[0024] The second branch first adjusts the channel dimension through a 1×1 convolution, then extracts local features through a 3×3 convolution, then enhances the response of important areas through the CASATT module, and finally adjusts the number of channels through a 1×1 convolution;

[0025] The third branch uses a single 1×1 convolution;

[0026] The three-branch calculation process is shown as follows:

[0027]

[0028] X3=X′,

[0029] X out =Concat(X1,Concat(X2,X3)).

[0030] Finally, the outputs of the three branches are spliced ​​in the channel dimension to achieve the purpose of fusing different depths and multi-semantic information.

[0031] Furthermore, step S3 includes: combining the newly added detection head and the multi-scale feature weighted fusion module to perform target discrimination on its input features, wherein the multi-scale feature weighted fusion module is used to dynamically adjust the contribution of each scale feature to the final fusion feature through a learnable weight parameter, wherein X i Represents the features of the i-th scale; assign a learnable weight W to each feature i , and constrain the sum of the weights to 1 through weighted normalization; each feature is weighted according to the normalized weight, and then bilinear interpolation is used to upsample the smaller-sized features to match the resolution of the largest-scale features. After completing the scale alignment, all features are spliced ​​according to the channel dimension, which is expressed as follows:

[0032] X={X1,X2,...,X n ,},

[0033]

[0034] X′ i =α i ·X i ,

[0035] X″ i =Interp(X′ i ,size=(H max ,W max )),

[0036] X out =Concat(X″1,X″2,...,X″ n ).

[0037] Furthermore, in step S4, the loss function is as follows:

[0038] Loss=λ1L box +λ2L obj .

[0039] The present invention has the following beneficial effects:

[0040] The present invention discloses real-time target detection based on the characteristics of small targets of drones in infrared images; a dual-branch hybrid convolution is designed in the Backbone stage, which significantly improves the feature expression ability of small targets and enhances the noise filtering effect through the dual-branch structure and hybrid convolution operation. A three-branch attention guidance module is designed in the Neck stage to enhance the attention of input features from two dimensions, namely, space and channel. The input features are differentiated by adopting a three-branch parallel structure, thereby effectively filtering the noise and retaining the target detail information. In addition, in order to further solve the problems of missed detection, a multi-scale feature weighting fusion module is designed at the newly added detection head to weight the features in an adaptive manner, optimize the feature propagation path, and thus improve the detection accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flow chart of infrared UAV target detection according to an embodiment of the present invention;

[0042] Figure 2 This is a diagram of the overall architecture of infrared UAV target detection according to an embodiment of the present invention;

[0043] Figure 3 Schematic diagram of a dual-branch hybrid convolutional architecture for infrared UAV target detection according to an embodiment of the present invention;

[0044] Figure 4This is a schematic diagram of the structure of a three-branch attention guidance module for infrared UAV target detection according to an embodiment of the present invention;

[0045] Figure 5 This is a schematic diagram of the CASATT attention mechanism structure for infrared UAV target detection according to an embodiment of the present invention;

[0046] Figure 6 This is a schematic diagram of the structure of a multi-scale feature weighted fusion module for infrared UAV target detection according to an embodiment of the present invention;

[0047] Figure 7 This is a comparison chart of target detection indicators of various comparison methods for infrared UAV target detection according to an embodiment of the present invention;

[0048] Figure 8 This is a comparison chart of model performance indicators of various comparative methods for infrared UAV target detection according to an embodiment of the present invention;

[0049] FIG9 is a diagram illustrating the detection effect of infrared drone target detection in multiple scenarios according to an embodiment of the present invention; FIG9 is a diagram illustrating the detection effect of a non-dedicated infrared small target detection method in multiple scenarios according to an embodiment of the present invention; FIG9 is a diagram illustrating the detection effect of a dedicated infrared small target detection method in multiple scenarios according to an embodiment of the present invention; FIG9 is a diagram illustrating the detection effect of a non-dedicated infrared small target detection method in multiple scenarios according to an embodiment of the present invention; FIG9 is a diagram illustrating

[0050] Figure 10 This is a diagram showing the ablation experiment results of the innovative module for infrared UAV target detection according to an embodiment of the present invention. DETAILED DESCRIPTION

[0051] The following is a combination of the embodiments of the present invention Figure 1-Figure 2 , the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.

[0052] During the specific implementation of the present invention, an infrared camera is used to capture drone target images and construct an infrared small target dataset. The algorithm model is trained on the dataset to accurately identify the positioning of drone targets within the area, providing technical support for subsequent countermeasures.

[0053] The present invention provides an infrared UAV target detection method based on efficient convolution and feature weighting fusion. Considering that in actual application scenarios, small infrared UAV targets are prone to missed detection and false detection in complex environments due to their small size and low contrast, detection efficiency must also be taken into account. A dual-branch hybrid convolution (DBHConv) is designed in the Backbone stage. Through the dual-branch structure and hybrid convolution operation, the feature expression ability of small targets is significantly improved, and the noise filtering effect is enhanced; then a three-branch attention guidance module (TBAG) is designed in the Neck stage. Combining the spatial and channel attention mechanisms, the input features are differentiated by adopting a three-branch parallel structure, thereby effectively filtering noise and retaining target detail information; finally, in order to further solve the problem of missed detection, a multi-scale feature weighting fusion module (MSFWF) is designed at the newly added detection head. The features are weighted in an adaptive manner, the feature propagation path is optimized, thereby improving the detection accuracy of the model and meeting the real-time detection requirements. The present invention provides a new method for regional countermeasures against illegal UAV flights.

[0054] Specifically, the infrared UAV target detection method based on efficient convolution and feature weighted fusion includes the following steps:

[0055] Step S0: Use professional equipment such as infrared cameras to obtain an infrared drone small target dataset and preprocess the images; the image preprocessing includes: adjusting the image to a uniform size, normalizing the pixel values, and data enhancement (such as random cropping, flipping, etc.) to improve the generalization ability of the model.

[0056] Step S1: Label the dataset including information such as the target bounding box coordinates, and then feed it into the network backbone for feature extraction:

[0057] The input infrared image is passed through the Conv module, C3 module, and DBHConv module in the backbone of the network to extract multi-scale feature information. The DBHConv module adopts a dual-path structure design. In the first branch, the input features are first evenly divided into two parts, X1 and X2, along the channel dimension. X1 is downsampled through a 3x3 convolution to extract local detailed features of small targets; while X2 undergoes a 3x3 max pooling operation to retain information in highlight areas and focus on possible target areas. This parallel structure improves the effectiveness of downsampling while effectively alleviating the problem of information loss. The two are then spliced ​​together to output F1. In the second branch, the complete input feature X is directly downsampled through a 3x3 convolution to supplement the feature information not involved in the convolution operation in the previous branch. Subsequently, this branch models the information between channels through 1x1 depthwise separable convolution (DWConv) to capture a more discriminative feature distribution, and further fuses the features of the previous step using skip connections, outputting them as F2. Finally, the output features F1 and F2 of the two trunk branches are concatenated to form an enhanced feature representation F out The present invention replaces some general convolutions in the backbone network (Backbone stage) with a dual-branch hybrid convolution module to obtain multi-scale features. Figure 2 In the present invention, the module structure is as shown. Figure 3 As shown, it is expressed as:

[0058]

[0059] F1=Concat(Conv(X1),Conv(Pool(X2))),

[0060] F2=Concat(DWConv(Conv(X)),Conv(X)),

[0061] F out =Concat(F1,F2).

[0062] in Indicates that the input feature X is divided into two parts along the channel dimension. Pool represents the maximum pooling effect. X1 and X2 are the two parts of the input feature X divided equally along the channel dimension. Conv is a convolution operation, Concat is a feature concatenation operation, and DWConv is a depthwise separable convolution.

[0063] Step S2: The extracted multi-scale features are fed into the neck network for fusion, and the appropriate output layer is selected as the input feature of the detection head:

[0064] Before being fed into the detection head, the fused features are processed by a three-branch attention-guided module. This module extracts features through three parallel branches of varying depths and uses an attention mechanism to focus on key regions in the image. The first branch focuses on deep semantic information. It first applies a 1×1 convolution to the input features for channel rescaling. The features then enter a bottleneck structure consisting of a 1×1 convolution for channel compression and a 3×3 convolution for local region feature extraction. The structure then passes through a CASATT module to enhance the target region features in both the channel and spatial dimensions. Finally, the branch output is added to the input features that have undergone the previous 1×1 convolution via a residual connection. This iterative process further extracts deep features. The second branch is relatively simple in structure and focuses on capturing features of medium depth. It also first adjusts the channel dimension through a 1×1 convolution, then extracts local features through a 3×3 convolution. The CASATT module then enhances the response of important regions, and finally adjusts the number of channels through a 1×1 convolution. This branch has a lower convolution depth than the first branch. The third branch adopts a minimalist design, using a single 1×1 convolution as a lightweight path. This is primarily used to preserve the original input information and avoid information loss during multiple attention guidance and deep convolutions. The three branches form a complementary feature expression structure with deep and shallow layers, enhancing feature expression capabilities and model generalization capabilities across different scenarios.

[0065] The feature maps extracted by the present invention are fused, and the feature expression ability is enhanced through the neck network (Neck stage). The three-branch attention guidance module is placed in front of the detection heads of head2 and head3, as shown in the figure. Figure 2 As shown, the network is enhanced to pay attention to the target area in the complex background. In the present invention, the module structure is as follows Figure 4 As shown, it is expressed as:

[0066]

[0067]

[0068] X3=X′,

[0069] X out =Concat(X1,Concat(X2,X3)).

[0070] Where X is the input feature, is a 1×1 convolution operation, is a 3×3 convolution operation, is an element-by-element addition operation, CASATT represents the spatial and channel attention mechanism, X i=1,2,3 It represents the features after processing the i-th branch, and Concat is the feature concatenation operation.

[0071] Among them, CASAt represents the spatial and channel attention mechanism, and the specific structure is as follows Figure 5 As shown, the spatial attention mechanism is denoted as SAtt and the channel attention mechanism is denoted as CAtt, where SAtt and CAtt can be expressed as:

[0072]

[0073] in represents element-wise multiplication, σ represents the Sigmoid activation function, and APP stands for adaptive average pooling. The three branches of the three-branch attention guidance module process the input feature X and output X1, X2, and X3, respectively, which are then concatenated according to the channel dimension. ReLU is the activation function, and BN stands for batch normalization.

[0074] Finally, the outputs of the three branches are spliced ​​in the channel dimension to achieve the purpose of fusing different depths and multi-semantic information;

[0075] Step S3: The detection head performs target discrimination on the input features: Combined with the newly added detection head and the multi-scale feature weighted fusion module, the input features are used to perform target discrimination:

[0076] The multi-scale feature weighted fusion module uses learnable weight parameters to dynamically adjust the contribution of each scale feature to the final fusion feature, where X i Represents the feature of the i-th scale. In order to achieve the empowerment of features of different scales, we assign a learnable weight W to each feature i , and weighted normalization is used to constrain the sum of the weights to 1 to ensure that different features contribute reasonably during the fusion process. Secondly, due to the inconsistent spatial dimensions of features at different scales, direct weighted fusion may result in information loss. Therefore, each feature is first weighted according to the normalized weight, and then bilinear interpolation is used to upsample the smaller features to match the resolution of the largest feature. After completing the scale alignment, all features are spliced ​​according to the channel dimension.

[0077] The new detection heads include Figure 2 As shown in head1, based on the characteristics of the data set that mainly consists of small targets, a new detection head is added to reduce the target missed detection rate, and the feature path is replanned for it. The different scale features of the Backbone stage and the Neck stage are adaptively weighted and fused through the feature weighting fusion module. Figure 6 As shown, it can be expressed as:

[0078] X={X1,X2,...,X n ,},

[0079]

[0080] X′ i =α i ·X i ,

[0081] X″ i =Interp(X′ i ,size=(H max ,W max )),

[0082] X out =Concat(X″1,X″2,...,X″ n ).

[0083] Where X is the multi-scale feature set of the input, W i To initialize the feature weights, θ is a small constant to prevent division by zero errors, and α i is the normalized weight coefficient, and Interp is the bilinear interpolation method for upsampling.

[0084] The present invention selects the appropriate feature layer in the network and outputs it to the detection head, such as Figure 2 As shown in Figure 2, target prediction is performed, which includes bounding box prediction and confidence prediction. Since there is only one category, predicting it as a target includes category prediction.

[0085] S4: Based on the prediction results, calculate the positioning loss and confidence loss to get the total loss. The model is trained by optimizing the total loss. The loss function is as follows:

[0086] Loss=λ1L box +λ2L obj .

[0087] S5: Repeat the process from S0 to S4 to make the network weight converge to the optimal value and obtain the best weight file, thereby realizing the detection of small drone targets.

[0088] In this implementation case, two datasets were selected: the RISTCS and NUDT-SIRST datasets. The RISTCS dataset is a real-world drone target dataset with 11,545 images, each with a resolution of 640×512 and containing only one target. In contrast, the NUDT-SIRST dataset contains 1,327 images, each with a resolution of 512×512. Its characteristic is that the number of targets in each image is not fixed. To better train and validate the model, both datasets are divided into training and test sets in a ratio of approximately 8:2. When evaluating model performance, the evaluation metrics used include Precision (P), Recall (R), F1-Score (F1), Average Precision (AP), number of parameters (Parameters), computational effort (GFLOPs), and frame rate (FPS), comprehensively measuring detection accuracy, efficiency, and model complexity. For comparison, we selected two-stage detection methods: Faster R-CNN and Cascade R-CNN; single-stage detection methods: YOLOv5, YOLOv10, YOLOv11, YOLOv12, FFCA-YOLO, Sparse R-CNN, and TOOD; and several advanced methods designed for infrared small target detection: DNANet, AGPCNet, RPCANet, ABCNet, and L2SKNet. Regarding the experimental setup, our method was trained for 80 epochs on the RISTCS dataset and 1000 epochs on the NUDT-SIRST dataset, with a batch size of 40 and the SGD optimizer for training. All experiments were conducted on an NVIDIA GeForce RTX 3090 graphics card, using Python 3.8.18, CUDA 12.1, and PyTorch 2.1.0. We strictly follow the principle of fairness and use the default parameters in the original code of each comparison algorithm for training, and ensure that all algorithms are trained to the optimal state to fully tap their performance potential. Figure 7 As shown in the above two data sets, its indicators are almost the best, with F1 values ​​reaching 86.72% and 98.63%. Figure 8 ,The method of the present invention achieves a real-time detection efficiency of 74.63 frames per second on the RISTCS dataset, while ,taking into account both model parameters and computational ,complexity.

[0089] At the same time, we selected the detection effect diagrams in different scenes from the RISTCS dataset and demonstrated them with various comparison methods, as shown in Figure 9. The blue box indicates the original target position, the yellow box indicates missed detection or wrong detection, and the red box indicates correct detection. In many complex real scenes, the method proposed in this invention achieved the best results, which is mutually confirmed with the above target detection indicators. At the same time, we also conducted ablation experiments on the three modules proposed in this invention, such as Figure 10 As shown, to further verify its innovation.

[0090] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various deformations, modifications, and substitutions made to the technical solutions of the present invention by ordinary technicians in this field should fall within the scope of protection determined by the claims of the present invention.

Claims

1. Infrared UAV target detection method based on efficient convolution and feature weighting fusion, characterized by: include: Step S0: Use an infrared camera to capture the target image to obtain a dataset of small infrared drone targets and preprocess the image; Step S1: label the data set and then send it to the network backbone for feature extraction; Step S2: The extracted multi-scale features are fed into the neck network for fusion, and the output layer is selected as the input feature of the detection head; Step S3: The detection head performs target discrimination on the input features; Step S4: Based on the prediction results, calculate the positioning loss and confidence loss to obtain the total loss; train the model by optimizing the total loss; Step S5: Repeat steps S0-S4 to make the network weights converge to the optimum and obtain the best weight file, thereby realizing the detection of small drone targets.

2. The infrared UAV target detection method based on efficient convolution and feature weighted fusion according to claim 1 is characterized in that: In step S0, the image is preprocessed, including: adjusting the image to a uniform size, normalizing the pixel values, and data augmentation to improve the generalization ability of the model.

3. The infrared UAV target detection method based on efficient convolution and feature weighted fusion according to claim 1 is characterized in that: Step S1 includes: extracting multi-scale feature information from the input infrared image through the Conv module, C3 module and DBHConv module in the Backbone of the network; The DBHConv module adopts a dual-path structure design. In the first branch, the input features are first evenly divided into two parts X1 and X2 along the channel dimension. X1 is downsampled by 3×3 convolution to mine the local detail features of small targets; X2 undergoes a 3×3 maximum pooling operation to retain the information of the highlight area and focus on the possible target area; the two are spliced ​​and output F1; in the second branch, the complete input feature X is downsampled by a 3×3 convolution, and then the branch uses a 1×1 depth-separable convolution to model the information between channels to capture a more discriminative feature distribution, and uses jump connections to further fuse the features of the previous step and output it as F2; ​​finally, the output features F1 and F2 of the two trunk branches are spliced ​​to form an enhanced feature representation F out : F1=Concat(Conv(X1),Conv(Pool(X2))), F2=Concat(DWConv(Conv(X)),Conv(X)), F out =Concat(F1,F2) in Indicates that the input feature X is divided into two parts along the channel dimension, Pool represents the maximum pooling effect; X1 and X2 are respectively the input feature X is divided into two parts along the channel dimension, Conv is the convolution operation, Concat is the feature splicing operation, and DWConv is the depth-wise separable convolution.

4. The infrared UAV target detection method based on efficient convolution and feature weighted fusion according to claim 1 is characterized in that: Step S2 includes: before the fused features are fed into the detection head, they are processed by a three-branch attention guidance module, which extracts features through three parallel branches of different depths and uses an attention mechanism to focus on key areas in the image; The first branch first applies a 1×1 convolution to the input features for channel adjustment, and then enters the Bottleneck structure, which includes a 1×1 convolution for channel compression and a 3×3 convolution for local area feature extraction. The structure then passes through the CASATT module to enhance the target area features in both channel and spatial dimensions. Finally, the branch output is added to the current 1×1 convolution input features through a residual connection, and deep features are extracted through this iterative process. The second branch first adjusts the channel dimension through a 1×1 convolution, then extracts local features through a 3×3 convolution, then enhances the response of important areas through the CASATT module, and finally adjusts the number of channels through a 1×1 convolution; The third branch uses a single 1×1 convolution; The three-branch calculation process is shown as follows: X3=X′, X out =Concat(X1,Concat(X2,X3)) Where X is the input feature, is a 1×1 convolution operation, is a 3×3 convolution operation, ⊕ is an element-by-element addition operation, CASATT represents the spatial and channel attention mechanism, X i=1,2,3 It represents the features after processing the i-th branch, and Concat is the feature concatenation operation. Finally, the outputs of the three branches are spliced ​​in the channel dimension to achieve the purpose of fusing different depths and multi-semantic information.

5. The infrared UAV target detection method based on efficient convolution and feature weighted fusion according to claim 1 is characterized in that: Step S3 includes: combining the newly added detection head and the multi-scale feature weighted fusion module to perform target discrimination on its input features, wherein the multi-scale feature weighted fusion module is used to dynamically adjust the contribution of each scale feature to the final fusion feature through a learnable weight parameter, wherein X i Represents the features of the i-th scale; assign a learnable weight W to each feature i , and constrain the sum of the weights to 1 through weighted normalization; each feature is weighted according to the normalized weight, and then bilinear interpolation is used to upsample the smaller-sized features to match the resolution of the largest-scale features. After completing the scale alignment, all features are spliced ​​according to the channel dimension, which is expressed as follows: X={X1,X2,...,X n ,}, X′ i =a i ·X i , X″ i =Interp(X′ i ,size=(H max ,W max )), X out =Concat(X″1,X″2,...,X″ n )。 6. The infrared UAV target detection method based on efficient convolution and feature weighted fusion according to claim 1 is characterized in that: In step S4, the loss function is as follows: Loss=λ1L box +λ2L obj 。