Infrared small target detection method and system based on unmanned aerial vehicle aerial photography scene, and medium

By employing multi-scale feature extraction and feature fusion methods, the challenge of small target detection in UAV infrared images is solved, improving detection accuracy and robustness, and making it suitable for infrared small target detection on UAV platforms.

CN120953572APending Publication Date: 2025-11-14SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510822104.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing target detection algorithms have weak ability to represent small target features in UAV infrared images, insufficient background noise suppression, low detection accuracy, and poor real-time performance, especially in complex scenarios.

Method used

Multi-scale feature extraction is performed using a feature extraction module, and feature enhancement is performed using a feature fusion module. The target category and bounding box coordinates are obtained through a decoding and prediction module. Feature extraction is performed using the Stem, Conv, HGBlock, C2f, and DAF modules, and feature fusion is performed using the Conv, Upsample, DTAB, and DGCST2 modules.

Benefits of technology

It significantly improves the robustness and accuracy of infrared small target detection, especially in complex environments, and is suitable for resource-constrained UAV platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953572A_ABST
    Figure CN120953572A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared small target detection method and system based on an unmanned aerial vehicle aerial photography scene and a medium, and the method comprises the steps: carrying out the feature extraction of an input image through a feature extraction module, and obtaining a multi-scale feature map; performing feature enhancement processing on the multi-scale feature map based on a feature fusion module to obtain a quality enhanced feature map; and performing decoding prediction on the quality enhancement feature map through a decoding prediction module to obtain the category and bounding box coordinates of the target. According to the method, through combination of multi-scale feature extraction, feature fusion enhancement and a decoding prediction method, accurate prediction of a target category and bounding box coordinates thereof is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an infrared small target detection method, system, and medium based on drone aerial photography scenarios. Background Technology

[0002] In the field of UAV remote sensing monitoring and intelligent visual analysis, with the continuous advancement of UAV platforms and infrared imaging technology, their image acquisition and target recognition capabilities in complex environments such as disaster early warning, urban governance, traffic control, and agricultural inspection have been significantly enhanced. Compared with traditional ground monitoring methods, UAVs have advantages such as flexible deployment, wide coverage, and low operating costs, making them particularly suitable for real-time perception and information acquisition in large-scale dynamic scenarios. In recent years, the integration of infrared imaging technology has further expanded the application boundaries of UAVs in nighttime, low-light, or adverse weather conditions, enabling effective identification and continuous tracking of heat source targets. However, limited by factors such as long high-altitude shooting distance, limited sensor resolution, and complex imaging environments, targets in infrared images acquired by UAVs often appear as small targets with blurred edges and low contrast, and are significantly affected by background noise and thermal interference, leading to a substantial increase in the difficulty of target detection.

[0003] Current mainstream target detection methods mainly include single-stage detectors based on convolutional neural networks (such as the YOLO series) and end-to-end detectors based on the Transformer architecture (such as the DETR series). YOLO-like models, with their advantages of lightweight structure, fast inference speed, and convenient deployment, perform excellently in most real-time detection tasks. However, they still face many challenges when processing small infrared targets: on the one hand, their deep feature maps have low spatial resolution, making it difficult to effectively preserve the detailed information of small targets; on the other hand, infrared images themselves lack color and texture features, making small targets more easily obscured by complex backgrounds, leading to missed or false detections. Although YOLO introduces multi-scale feature fusion mechanisms (such as FPN and PAN), its detection performance is still insufficient when dealing with extreme scale changes or dense small target scenes. In contrast, the DETR series models achieve end-to-end prediction of targets through global modeling and query mechanisms, avoiding the complexity brought by anchor box design and NMS post-processing. Improved models such as Deformable-DETR and RT-DETR have improved convergence speed and detection accuracy to some extent. However, in infrared small target detection tasks, this type of model still reveals several bottlenecks: First, its network parameters are large and computational overhead is high, making it difficult to adapt to resource-constrained UAV embedded platforms; second, the Transformer structure has limitations in fine-grained feature extraction and is weak in recognizing small targets with low signal-to-noise ratios; in addition, the query initialization strategy is prone to matching bias when facing complex scenes with dense targets, severe occlusion, or low contrast, which in turn affects the detection stability and accuracy.

[0004] In summary, existing target detection algorithms used in UAV infrared imagery generally suffer from problems such as weak representation of small target features, poor background suppression, and low model inference efficiency, making it difficult to meet the comprehensive requirements of detection accuracy, real-time performance, and robustness in practical deployments. Especially in typical scenarios such as drastic changes in target scale, complex lighting conditions, or multiple overlapping targets, the overall performance of the detection system still faces severe challenges.

[0005] Therefore, existing technologies still need to be improved and enhanced. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide an infrared small target detection method, system and medium based on UAV aerial photography scenarios, which addresses the above-mentioned defects of the prior art. The aim is to solve the problems faced by existing target detection algorithms when processing infrared images taken from high altitudes and long distances, such as weak small target feature expression ability, insufficient background noise suppression, low detection accuracy and poor real-time performance.

[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0008] In a first aspect, the present invention provides an infrared small target detection method based on UAV aerial photography scenarios, wherein the method includes:

[0009] The feature extraction module is used to extract features from the input image to obtain a multi-scale feature map;

[0010] The feature fusion module is used to perform feature enhancement processing on the multi-scale feature map to obtain a quality-enhanced feature map.

[0011] The enhanced feature map is decoded and predicted by the decoding and prediction module to obtain the target's category and bounding box coordinates.

[0012] In one implementation, the feature extraction module includes a Stem module, a Conv module, an HGBlock module, a C2f module, and a DAF module. The step of using the feature extraction module to extract features from the input image to obtain a multi-scale feature map includes:

[0013] The Stem module is used to perform preliminary feature extraction on the input image to obtain an initial feature map;

[0014] Based on the Conv module, HGBlock module, C2f module and DAF module, feature extraction is performed on the initial feature map to obtain the first multi-scale feature map;

[0015] Based on the Conv module, HGBlock module, C2f module and DAF module, feature extraction is performed on the first multi-scale feature map to obtain the second multi-scale feature map;

[0016] Based on the Conv module, HGBlock module, C2f module and DAF module, feature extraction is performed on the second multi-scale feature map to obtain the third multi-scale feature map.

[0017] In one implementation, the step of extracting features from the initial feature map based on the Conv module, HGBlock module, C2f module, and DAF module to obtain a first multi-scale feature map includes:

[0018] The initial feature map is extracted sequentially using the C2f module, the Conv module, and the C2f module to obtain the first branch feature map;

[0019] The initial feature map is then processed sequentially using the HGBlock module, the DWConv module, and the HGBlock module to obtain the second branch feature map.

[0020] Using the DAF module, feature extraction is performed on the first branch feature map and the second branch feature map to obtain the first multi-scale feature map.

[0021] In one implementation, the step of extracting features from the first multi-scale feature map based on the Conv module, HGBlock module, C2f module, and DAF module to obtain the second multi-scale feature map includes:

[0022] The first multi-scale feature map is extracted by sequentially using the Conv module and the C2f module to obtain the third branch feature map;

[0023] The initial feature map is then extracted sequentially using the DWConv module and three HGBlock modules to obtain the fourth branch feature map.

[0024] Using the DAF module, feature extraction is performed on the third branch feature map and the fourth branch feature map to obtain the second multi-scale feature map.

[0025] In one implementation, the step of extracting features from the second multi-scale feature map based on the Conv module, HGBlock module, C2f module, and DAF module to obtain the third multi-scale feature map includes:

[0026] The Conv module and C2f module are used sequentially to extract features from the second multi-scale feature map to obtain the fifth branch feature map;

[0027] The second multi-scale feature map is then extracted sequentially using the DWConv module and the HGBlock module to obtain the sixth branch feature map.

[0028] Using the DAF module, feature extraction is performed on the fifth branch feature map and the sixth branch feature map to obtain the third multi-scale feature map.

[0029] In one implementation, the feature fusion module performs feature enhancement processing on the multi-scale feature map to obtain a quality-enhanced feature map, wherein the feature fusion module includes a Conv module, an Upsample module, a DTAB module, and a DGCST2 module, comprising:

[0030] The multi-scale feature map is processed using the Conv module, Upsample module, DTAB module, and DGCST2 module to obtain the first quality-enhanced feature map.

[0031] Based on the first quality-enhanced feature map, the Conv module and the DGCST2 module are used for convolution and fusion processing to obtain the second quality-enhanced feature map.

[0032] Based on the second quality enhancement feature map, the third quality enhancement feature map is obtained by performing convolution and fusion processing using the Conv module and the DGCST2 module.

[0033] The first quality enhancement feature map, the second quality enhancement feature map, and the third quality enhancement feature map are fused to obtain a quality enhancement feature map.

[0034] In one implementation, the processing of the multi-scale feature map based on the Conv module, Upsample module, DTAB module, and DGCST2 module to obtain a first quality-enhanced feature map includes:

[0035] The third multi-scale feature map is processed sequentially using the Conv module, DTAB module, and Conv module to obtain a preliminary quality-enhanced feature map.

[0036] Based on the preliminary quality enhancement feature map and the second multi-scale feature map, the Upsample module and Conv module are used for upsampling and convolution processing, and a Concat operation is performed to obtain the first concatenated feature map.

[0037] Based on the DGCST2 module and the Conv module, the first concatenated feature map is fused and convolved to obtain a convolutional fused feature map.

[0038] Based on the convolutional fusion feature map and the first multi-scale feature map, the Upsample module and Conv module are used to perform upsampling and convolution processing, and then a Concat operation is performed to obtain the second concatenated feature map.

[0039] The second stitched feature map is fused using the DGCST2 module to obtain the first quality-enhanced feature map.

[0040] In one implementation, the step of performing upsampling and convolution processing using the Upsample and Conv modules based on the preliminary quality-enhanced feature map and the second multi-scale feature map, followed by a Concat operation to obtain a first concatenated feature map, includes:

[0041] The Upsample module is used to upsample the preliminary quality enhancement feature map to obtain a sampled preliminary quality enhancement feature map;

[0042] The second multi-scale feature map is convolved using the Conv module to obtain the convolved second multi-scale feature map.

[0043] The sampled preliminary quality-enhanced feature map and the convolutional second multi-scale feature map are concatenated to obtain the first concatenated feature map.

[0044] In one implementation, the step of performing upsampling and convolution processing using the Upsample and Conv modules based on the convolutionally fused feature map and the first multi-scale feature map, followed by a Concat operation to obtain the second concatenated feature map, includes:

[0045] The Upsample module is used to upsample the convolutional fusion feature map to obtain the sampled convolutional fusion feature map;

[0046] Using the Conv module, the first multi-scale feature map is convolved to obtain the convolved first multi-scale feature map;

[0047] The sampled convolutional fusion feature map and the first convolutional multi-scale feature map are concatenated to obtain the second concatenated feature map.

[0048] In one implementation, the step of obtaining a second quality-enhanced feature map by performing convolution and fusion processing using the Conv module and the DGCST2 module based on the first quality-enhanced feature map includes:

[0049] Based on the Conv module, the first quality enhancement feature map is convolved to obtain the convolved first quality enhancement feature map;

[0050] The convolutional fusion feature map and the first convolutional quality enhancement feature map are concatenated to obtain the third concatenated feature map;

[0051] The third stitched feature map is fused using the DGCST2 module to obtain the second quality-enhanced feature map.

[0052] In one implementation, the step of obtaining a third quality-enhanced feature map by performing convolution and fusion processing using the Conv module and the DGCST2 module based on the second quality-enhanced feature map includes:

[0053] Based on the Conv module, the second quality enhancement feature map is convolved to obtain the convolutional second quality enhancement feature map;

[0054] The second quality enhancement feature map of the convolution is concatenated with the preliminary quality enhancement feature map to obtain the fourth concatenated feature map;

[0055] The fourth stitched feature map is fused using the DGCST2 module to obtain the third quality-enhanced feature map.

[0056] In one implementation, the step of decoding and predicting the quality-enhanced feature map through a decoding and prediction module to obtain the target's category and bounding box coordinates includes:

[0057] The quality-enhanced feature map is processed through an embedded query selection mechanism to obtain a position-sensitive query vector;

[0058] The query vector is decoded and predicted using a decoder and a prediction head to obtain the target's category and bounding box coordinates.

[0059] Secondly, embodiments of the present invention also provide an infrared small target detection system based on UAV aerial photography scenarios, wherein the system includes:

[0060] The multi-scale feature map acquisition module is used to extract features from the input image using the feature extraction module to obtain a multi-scale feature map;

[0061] The quality enhancement feature map acquisition module is used to process the multi-scale feature map based on the feature fusion module to obtain the quality enhancement feature map;

[0062] The category and bounding box coordinate acquisition module is used to decode and predict the quality-enhanced feature map through the decoding and prediction module to obtain the category and bounding box coordinates of the target.

[0063] Thirdly, embodiments of the present invention also provide a terminal, wherein the terminal includes a memory, a processor, and an infrared small target detection program based on a drone aerial photography scene stored in the memory and capable of running on the processor. When the processor executes the infrared small target detection program based on a drone aerial photography scene, it implements the steps of the infrared small target detection method based on a drone aerial photography scene of any of the above-mentioned solutions.

[0064] Fourthly, embodiments of the present invention also provide a computer-readable storage medium, wherein the computer-readable storage medium stores an infrared small target detection program based on a drone aerial photography scene, and when the infrared small target detection program based on a drone aerial photography scene is executed by a processor, it implements the steps of the infrared small target detection method based on a drone aerial photography scene as described in any of the above schemes.

[0065] Beneficial Effects: This invention provides an infrared small target detection method based on UAV aerial photography scenes. Compared with existing technologies, this invention first utilizes a feature extraction module to extract features from the input image, obtaining a multi-scale feature map. This step processes the input image using the feature extraction module, capturing detailed information at different levels through multi-scale analysis, resulting in a multi-scale feature map containing semantic information from low to high levels. This multi-level information representation helps improve the recognition ability of targets of different sizes, especially in complex backgrounds or under occlusion. Next, a feature fusion module performs feature enhancement processing on the multi-scale feature map to obtain a quality-enhanced feature map. This process integrates various features... Information from feature maps at different scales effectively enhances key features and suppresses noise, generating higher-quality feature maps. This step is particularly important for improving the model's performance when dealing with small targets or fine-grained features, as high-quality feature representations can more accurately describe the characteristics and contextual information of the target. Then, the enhanced feature map is decoded and predicted by the decoding and prediction module to obtain the target's category and bounding box coordinates. This module transforms the abstract features learned by the deep learning model into concrete, interpretable results, i.e., determining the specific location of the target in the image and its category. This not only improves the accuracy of localization but also enhances the precision of classification, making the final detection results more reliable. Overall, the method proposed in this invention combines multi-scale feature extraction, feature fusion, and accurate decoding and prediction to form an efficient target detection process. This method can significantly improve the detection performance of targets of various sizes, especially robustness and accuracy in complex environments, while also improving the detection effect of small targets, providing strong support for practical applications. Attached Figure Description

[0066] Figure 1This is a flowchart illustrating a specific implementation method for an infrared small target detection method based on UAV aerial photography scenarios provided in an embodiment of the present invention.

[0067] Figure 2 This is a flowchart of a preferred embodiment of the infrared small target detection method based on UAV aerial photography scenarios provided by the present invention.

[0068] Figure 3 This is a schematic diagram of the HGBlock module in the infrared small target detection method based on UAV aerial photography scenarios provided in an embodiment of the present invention.

[0069] Figure 4 This is a schematic diagram of the C2f module in the infrared small target detection method based on UAV aerial photography scenarios provided in an embodiment of the present invention.

[0070] Figure 5 This is a schematic diagram of the DAF module in the infrared small target detection method based on UAV aerial photography scenarios provided in an embodiment of the present invention.

[0071] Figure 6 This is a schematic diagram of the DTAB module in the infrared small target detection method based on UAV aerial photography scenarios provided in an embodiment of the present invention.

[0072] Figure 7 This is a schematic diagram of the DGCST2 module in the infrared small target detection method based on UAV aerial photography scenarios provided in the embodiments of the present invention.

[0073] Figure 8 This diagram illustrates the components of the original MPDIoU loss function in the infrared small target detection method based on UAV aerial photography scenarios provided in this embodiment of the invention.

[0074] Figure 9 This is a schematic diagram of the infrared small target detection system based on UAV aerial photography scenarios provided in an embodiment of the present invention.

[0075] Figure 10 This is a block diagram illustrating the internal structure of a terminal provided in an embodiment of the present invention. Detailed Implementation

[0076] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0077] In the field of UAV remote sensing monitoring and intelligent visual analysis, with the continuous advancement of UAV platforms and infrared imaging technology, their image acquisition and target recognition capabilities in complex environments such as disaster early warning, urban governance, traffic control, and agricultural inspection have been significantly enhanced. Compared with traditional ground monitoring methods, UAVs have advantages such as flexible deployment, wide coverage, and low operating costs, making them particularly suitable for real-time perception and information acquisition in large-scale dynamic scenarios. In recent years, the integration of infrared imaging technology has further expanded the application boundaries of UAVs in nighttime, low-light, or adverse weather conditions, enabling effective identification and continuous tracking of heat source targets. However, limited by factors such as long high-altitude shooting distance, limited sensor resolution, and complex imaging environments, targets in infrared images acquired by UAVs often appear as small targets with blurred edges and low contrast, and are significantly affected by background noise and thermal interference, leading to a substantial increase in the difficulty of target detection.

[0078] Current mainstream target detection methods mainly include single-stage detectors based on convolutional neural networks (such as the YOLO series) and end-to-end detectors based on the Transformer architecture (such as the DETR series). YOLO-like models, with their advantages of lightweight structure, fast inference speed, and convenient deployment, perform excellently in most real-time detection tasks. However, they still face many challenges when processing small infrared targets: on the one hand, their deep feature maps have low spatial resolution, making it difficult to effectively preserve the detailed information of small targets; on the other hand, infrared images themselves lack color and texture features, making small targets more easily obscured by complex backgrounds, leading to missed or false detections. Although YOLO introduces multi-scale feature fusion mechanisms (such as FPN and PAN), its detection performance is still insufficient when dealing with extreme scale changes or dense small target scenes. In contrast, the DETR series models achieve end-to-end prediction of targets through global modeling and query mechanisms, avoiding the complexity brought by anchor box design and NMS post-processing. Improved models such as Deformable-DETR and RT-DETR have improved convergence speed and detection accuracy to some extent. However, in infrared small target detection tasks, this type of model still reveals several bottlenecks: First, its network parameters are large and computational overhead is high, making it difficult to adapt to resource-constrained UAV embedded platforms; second, the Transformer structure has limitations in fine-grained feature extraction and is weak in recognizing small targets with low signal-to-noise ratios; in addition, the query initialization strategy is prone to matching bias when facing complex scenes with dense targets, severe occlusion, or low contrast, which in turn affects the detection stability and accuracy.

[0079] In summary, existing target detection algorithms used in UAV infrared imagery generally suffer from problems such as weak representation of small target features, poor background suppression, and low model inference efficiency, making it difficult to meet the comprehensive requirements of detection accuracy, real-time performance, and robustness in practical deployments. Especially in typical scenarios such as drastic changes in target scale, complex lighting conditions, or multiple overlapping targets, the overall performance of the detection system still faces severe challenges.

[0080] To address the aforementioned issues, this embodiment provides an infrared small target detection method based on UAV aerial photography scenarios. Specifically, this embodiment first utilizes a feature extraction module to extract features from the input image, obtaining a multi-scale feature map. This step processes the input image using the feature extraction module, capturing detailed information at different levels through multi-scale analysis, resulting in a multi-scale feature map containing semantic information from low to high levels. This multi-level information representation helps improve the recognition ability of targets of different sizes, especially in complex backgrounds or under occlusion. Next, a feature fusion module performs feature enhancement processing on the multi-scale feature map to obtain a quality-enhanced feature map. This process comprehensively... By incorporating information from feature maps at different scales, key features can be effectively enhanced and noise suppressed, resulting in higher-quality feature maps. This step is particularly important for improving the model's performance when dealing with small targets or fine-grained features, as high-quality feature representations can more accurately describe the characteristics and contextual information of the target. Then, the enhanced feature map is decoded and predicted by a decoding prediction module to obtain the target's category and bounding box coordinates. This module transforms the abstract features learned by the deep learning model into concrete, interpretable results, i.e., determining the specific location of the target in the image and its category. This not only improves the accuracy of localization but also enhances the precision of classification, making the final detection results more reliable. Overall, the method proposed in this invention forms an efficient target detection process by combining multi-scale feature extraction, feature fusion, and accurate decoding prediction. This method can significantly improve the detection performance of targets of various sizes, especially robustness and accuracy in complex environments, while also improving the detection effect of small targets, providing strong support for practical applications.

[0081] This embodiment provides an infrared small target detection method based on UAV aerial photography scenarios, which can be applied to smart terminals, such as... Figure 1 As shown, the specific steps include the following:

[0082] Step S100: Use the feature extraction module to extract features from the input image to obtain a multi-scale feature map.

[0083] In this embodiment, the input image is first efficiently and comprehensively extracted using a feature extraction module to obtain a multi-scale feature map with rich semantic information and spatial details. This feature extraction module consists of several key components, including the Stem module (backbone module), the Conv module (Convolution module: the process of weighting input data using a sliding "filter" or "convolution kernel" in image processing and neural networks), the HGBlock module (High-performance GPU Block: a convolutional feature extraction module that implements efficient computation on a GPU), the C2f module (Coarse-to-Fine combined feature extraction module: a feature extraction module that extracts features from coarse to fine granularity), and the DAF module (Dynamic Alignment Fusion: a technique for fusing information by dynamically adjusting alignment strategies in multimodal, multi-layer feature, or multi-scale information processing). Each module works collaboratively, fully leveraging its respective advantages. The feature extraction module employs a dual-path parallel structure. The Stem module performs initial high-dimensional feature mapping on the original input image, providing high-quality initial features for subsequent processing. The Conv module further extracts local features through stacked convolutional layers and enhances the model's non-linear expressive power. The shallow path uses HGBlock to extract edge structure information from multiple receptive fields, while the deep path uses the C2f module to model compact semantic representations. The multi-branch features output from each main stage are aligned and fused by the DAF module, effectively improving feature consistency and discriminative power. This structural design not only enhances the hierarchy and diversity of feature extraction but also significantly improves the model's ability to identify and locate targets in complex scenes, thereby improving the overall robustness and generalization performance of the system.

[0084] Specifically, step S100 includes the following steps:

[0085] Step S101: Use the Stem module to perform preliminary feature extraction on the input image to obtain an initial feature map;

[0086] Step S102: Based on the Conv module, HGBlock module, C2f module and DAF module, perform feature extraction on the initial feature map to obtain the first multi-scale feature map;

[0087] Step S103: Based on the Conv module, HGBlock module, C2f module and DAF module, perform feature extraction on the first multi-scale feature map to obtain the second multi-scale feature map;

[0088] Step S104: Based on the Conv module, HGBlock module, C2f module and DAF module, perform feature extraction on the second multi-scale feature map to obtain the third multi-scale feature map.

[0089] In one implementation, such as Figure 2 As shown, firstly, the Stem module (backbone module) performs preliminary feature extraction on the input image to obtain an initial feature map. The Stem module, as the starting point of the entire process, is responsible for processing the original input image, transforming it into an initial feature map with rich semantic information and spatial details through a series of efficient calculations. This not only provides a high-quality foundation for subsequent feature extraction but also enhances the model's understanding of complex scenes. Next, the C2f module, Conv module, and C2f module are used sequentially to extract features from the initial feature map, obtaining the first branch feature map. The C2f module extracts features from the image progressively from coarse to fine, enabling the model to capture information at different levels. The Conv module further enhances the extraction of local features by stacking convolutional layers and improves the model's ability to express nonlinear features. This combination not only effectively improves the diversity and accuracy of features but also enhances the model's accuracy in target recognition and localization. Then, based on the HGBlock module, DWConv module (Depthly Separable Convolutional Module: a technique to reduce the number of parameters and computational cost), and then the HGBlock module again, feature extraction is performed on the initial feature map to obtain the second branch feature map. Using the HGBlock module for shallow paths helps extract edge structure information from multiple receptive fields, while the application of the DWConv module reduces the computational burden while maintaining the efficiency of feature extraction. This method is particularly suitable for capturing detailed information at different scales in an image, thereby improving the richness of the feature map and the adaptability of the model. Next, using the DAF module, feature extraction is performed on the first branch feature map and the second branch feature map to obtain the first multi-scale feature map. The DAF module dynamically adjusts the alignment strategy to fuse multimodal or multi-scale information, ensuring the consistency and discriminativeness of features from different sources. This step significantly improves the overall effect and robustness of feature extraction.

[0090] Subsequently, the Conv and C2f modules are used sequentially to extract features from the first multi-scale feature map, resulting in the third branch feature map. This process enhances the non-linear expressive power of the features and the extraction of multi-level features, helping the model better understand complex scene information. Next, based on the DWConv module and three HGBlock modules, features are extracted from the initial feature map to obtain the fourth branch feature map. This design fully utilizes the advantages of depthwise separable convolution and high-performance GPU blocks, reducing computational costs while ensuring the quality and efficiency of feature extraction. Then, the DAF module is used to extract features from the third and fourth branch feature maps to obtain the second multi-scale feature map. This step further optimizes the consistency and discriminative power of the features, laying the foundation for subsequent feature processing.

[0091] Next, the Conv and C2f modules are used sequentially to extract features from the second multi-scale feature map, resulting in the fifth branch feature map. This stage further refines the details of the feature map, enhancing the model's ability to recognize targets in complex scenes. Subsequently, the DWConv and HGBlock modules are used sequentially to extract features from the second multi-scale feature map, resulting in the sixth branch feature map. This is done to capture more subtle feature information while maintaining low computational overhead. Then, the DAF module is used to extract features from the fifth and sixth branch feature maps, resulting in the third multi-scale feature map. Through this multi-level, multi-path feature extraction and fusion, the entire system not only significantly improves the comprehensiveness and accuracy of feature extraction but also significantly enhances the model's robustness and generalization performance in various application scenarios.

[0092] In one implementation, the HGBlock module was proposed by the Baidu PaddlePaddle Vision team, and its source code repository is located at: https: / / github.com / PaddlePaddle / PaddleClas / blob / release / 2.5.2 / docs / zh_CN / models / ImageNet1k / PP-HGNet.md. For example... Figure 3As shown, the HGBlock module extracts feature representations from different receptive fields through multiple parallel convolutional branches, including multi-level convolutional kernels ranging from 3×3 to 11×11 in size, to simultaneously capture local texture and broad-range contextual information. After the outputs at each scale are stacked and fused in the spatial dimension, the channel dimension is compressed via 1×1 convolution, and channel attention enhancement is introduced through the ESE (Efficient Squeeze-and-Excitation) mechanism, enabling the network to automatically focus on feature response regions with high target discrimination. This approach enhances the modeling ability of small target edges and morphological features, making it particularly suitable for infrared image environments with blurred target boundaries and irregular shapes.

[0093] In one implementation, the C2f module is a feature extraction module within the UltraLytics open-source vision algorithm framework, and its source code repository address is:

[0094] https: / / github.com / ultralytics / ultralytics / blob / main / docs / en / models / yolov8.md. (For example...) Figure 4 As shown, the C2f module is an efficient multi-branch feature fusion structure designed to enhance feature expressiveness while maintaining low computational overhead. Input features are first adjusted for channel dimensions using a convolution (kernel size 1, stride 1, padding 0), then divided into multiple sub-channels by a split operation, and each sub-channel is fed into several parallel Bottleneck sub-modules for feature transformation. Each Bottleneck module consists of two cascaded convolutions (kernel size 1, stride 1, padding 0), where residual connections (shortcuts) can be enabled or disabled as needed. When shortcut = True, the input features are directly added to the output through an identity mapping, preserving cross-level information; when shortcut = False, the transformation is completed only through convolution stacking. This design allows for flexible selection of whether to introduce residuals based on feature complexity, enhancing the model's expressiveness and convergence stability. The output features of each sub-path are then concatenated along the channel dimension and finally integrated into a unified output feature through another convolution (kernel size 1, stride 1, padding 0). This module introduces structural diversity through multi-branch parallel computation, improves feature flow efficiency through flexible residual paths, and maintains low computational complexity, making it suitable for the dual requirements of efficient semantic integration and edge preservation in infrared image small target detection tasks.

[0095] In one implementation, such as Figure 5As shown, the core design goal of the DAF module is to achieve adaptive fusion and weight balancing between feature maps of different scales or from different sources. The DAF module receives two input feature maps (feature maps from different sources). Figure 1 and characteristics Figure 2 The two feature maps are aligned using 1×1 convolutions, then concatenated along the channel dimension and fed into an attention branch consisting of a 1×1 convolution and a sigmoid activation function to generate a set of fusion guiding weights. These guiding weights are then split into two parts along the channel dimension, corresponding to two separate feature maps. Each feature map is then multiplied channel-wise with its corresponding guiding weight to achieve response adjustment based on local region awareness. Each feature map is further multiplied with a set of learnable global fusion coefficients (param1, param2) to further adjust the feature contribution ratio at the overall scale. Finally, the two weighted feature maps are fused using a weighted summation method, and the fused feature map is output through a 1×1 convolution. This mechanism effectively mitigates fusion degradation caused by scale differences, semantic inconsistencies, or spatial misalignments, improving feature consistency and expressive strength.

[0096] In summary, the feature extraction module achieves collaborative representation of local and semantic information through shallow and deep dual-branch modeling using the HGBlock and C2f modules. A dynamic alignment fusion strategy addresses the information imbalance problem in multi-scale target representation encountered by traditional static fusion. This module is lightweight and flexible in its design, improving the overall accuracy and robustness of small target detection while reducing redundant computational overhead. It is suitable for deployment in resource-constrained UAV edge devices and possesses high engineering application value and scalability.

[0097] Step S200: Perform feature enhancement processing on the multi-scale feature map based on the feature fusion module to obtain a quality-enhanced feature map.

[0098] In this embodiment, firstly, feature enhancement processing is performed on the multi-scale feature maps based on the feature fusion module to obtain more discriminative and robust quality-enhanced feature maps. The multi-scale feature maps include a first multi-scale feature map, a second multi-scale feature map, and a third multi-scale feature map. The feature fusion module includes a Conv module, an Upsample module, a DTAB module (Dilated Transformer Attention Block: an attention module that introduces a dilation mechanism into the Transformer structure to expand the receptive field and enhance feature capture capability), and a DGCST2 module (Depth-wise Group Convolution and ShuffleTransformer v2: this module improves model computational efficiency and feature representation capability by fusing deep grouped convolution and shuffle transformer structures). The Conv module primarily extracts and optimizes local feature information, enhancing feature representation by stacking convolutional layers with different receptive fields. The Upsample module increases the spatial dimension of low-resolution feature maps, making them consistent in spatial size with high-resolution feature maps, thus achieving cross-layer information complementarity. The DTAB module introduces an expansion mechanism on top of the traditional Transformer structure, effectively expanding the model's receptive field, enabling the network to more efficiently capture long-distance dependencies and multi-scale contextual information. Simultaneously, the attention mechanism helps focus on key regions, improving feature selectivity. The DGCST2 module combines the computational efficiency of depthwise groupable convolutions with the global modeling capability of shuffling transformers, reducing model computational complexity while further enhancing feature interaction and information flow between channels, improving overall feature representation and inference efficiency. These modules work collaboratively, fully utilizing local details and global semantic information, significantly enhancing the model's ability to perceive and recognize multi-scale targets, providing high-quality feature representations for subsequent tasks.

[0099] Specifically, step S200 includes the following steps:

[0100] Step S201: Based on the Conv module, Upsample module, DTAB module and DGCST2 module, process the multi-scale feature map to obtain the first quality-enhanced feature map;

[0101] Step S202: Based on the first quality enhancement feature map, perform convolution and fusion processing using the Conv module and the DGCST2 module to obtain the second quality enhancement feature map;

[0102] Step S203: Based on the second quality enhancement feature map, perform convolution and fusion processing using the Conv module and the DGCST2 module to obtain the third quality enhancement feature map;

[0103] Step S204: Perform a fusion operation on the first quality enhancement feature map, the second quality enhancement feature map, and the third quality enhancement feature map to obtain a quality enhancement feature map.

[0104] In one implementation, this invention designs a feature fusion module to perform feature enhancement processing on multi-scale feature maps, aiming to obtain more discriminative and robust quality-enhanced feature maps. For example... Figure 2As shown, firstly, the third multi-scale feature map is processed sequentially using the Conv module, DTAB module, and Conv module to extract and optimize local feature information. The receptive field of the model is expanded through the dilation transformer attention mechanism to capture long-distance dependencies and multi-scale contextual information, thereby obtaining a preliminary quality-enhanced feature map. This step fully utilizes the local feature extraction capabilities of the convolutional layer and the global feature capture and attention focusing functions of the DTAB module, laying a solid foundation for subsequent processing. Next, the Upsample module is used to upsample the preliminary quality-enhanced feature map, increasing the spatial dimension of the low-resolution feature map to achieve spatial size consistency with the high-resolution feature map, thus completing cross-level information complementarity. Subsequently, the Conv module is used to convolve the second multi-scale feature map, further enhancing its feature representation capability, resulting in a convolutional second multi-scale feature map. The sampled preliminary quality-enhanced feature map and the convolutional second multi-scale feature map are then concatenated (using the concatenation module) to form the first concatenated feature map. This process helps integrate feature information from different levels, enhancing the model's ability to perceive multi-scale targets. Then, the first concatenated feature map is fused and convolved using the DGCST2 and Conv modules to generate a convolutional fused feature map. This step combines the efficient computational characteristics of depthwise groupable convolution with the global modeling capability of the shuffle transformer, enhancing feature interaction and information flow between channels. Next, the convolutional fused feature map is upsampled using the Upsample module to further increase its spatial dimension, resulting in a sampled convolutional fused feature map. Then, the first multi-scale feature map is convolved using the Conv module to obtain a convolutional first multi-scale feature map, which is then concatenated with the sampled convolutional fused feature map to form a second concatenated feature map. This operation further enriches the feature representation. Subsequently, the second concatenated feature map is fused using the DGCST2 module to obtain a first quality-enhanced feature map. This step not only improves the quality of the feature map but also significantly improves the model's inference efficiency. Next, the first quality-enhanced feature map is convolved to obtain a convolutional first quality-enhanced feature map, which is then concatenated with the convolutional fusion feature map to obtain a third concatenated feature map. This process promotes effective fusion between features at different levels. Following this, the DGCST2 module is used to fuse the third concatenated feature map to generate a second quality-enhanced feature map. Then, the second quality-enhanced feature map is convolved using the Conv module to obtain a convolutional second quality-enhanced feature map, which is then concatenated with the initial quality-enhanced feature map to form a fourth concatenated feature map. Finally, the DGCST2 module is used to fuse the fourth concatenated feature map to obtain the third quality-enhanced feature map.The entire process, through multiple feature fusion and convolution processes, fully mines and integrates local details and global semantic information in the input data, significantly enhancing the model's ability to perceive and recognize multi-scale targets, providing high-quality feature representations, and providing strong support for the successful execution of subsequent tasks.

[0105] In one implementation, such as Figure 6 As shown, to address the limitations of existing technologies in global perception and low efficiency in spatial context modeling when dealing with complex backgrounds and small targets, this invention utilizes the DTAB module proposed in the paper "Rethinking Transformer-Based Blind-Spot Network for Self-Supervised Image Denoising" (AAAI2025). This module is a feature interaction structure combining dilated convolution and sparse attention mechanisms, aiming to effectively improve the network's feature modeling ability for small targets and complex regions in infrared images while maintaining computational efficiency. Specifically, the DTAB module consists of four sub-modules in sequence: Dilated Channel Self-Attention Module (DilatedG-CSA), the first Dilated Feedforward Network Module (DilatedFFN), the Dilated Mask Window Attention Module (DilatedM-WSA), and the second Dilated Feedforward Network Module (DilatedFFN). The input features first enter the Dilated Channel Self-Attention (D-CSA) module. This module first undergoes layer normalization (LayerNorm), then models the channel responses under different receptive fields through three parallel 3×3 dilated convolution branches. The fused features are used to enhance the channel attention representation at multiple scales and are fused with the main branch through a weighted mechanism. Subsequently, the output features enter the first Dilated Feedforward (FFN) module. This module consists of layer normalization (LayerNorm) and two 3×3 dilated convolutions, with a nonlinear activation function (GeLU) introduced in between to further enhance the inter-channel feature representation and nonlinear transformation capabilities. After the above modeling, the features are fed into the Dilated Masked Window Attention (M-WSA) module. This module performs three 1×1 convolution operations after layer normalization (LayerNorm) to generate the query (Q), key (K), and value (V) processes, and performs sparse attention interactions within a local window. Specifically, the original attention weights are obtained by matrix multiplication and scaling of the query (Q) and key (K) for all pixels within the window. (where d is the feature dimension), and then a predefined binary mask matrix is ​​introduced. The elements in matrix M are defined as follows:

[0106] In the formula, (x i ,y i ) and (x j ,y j The coordinates of the query pixel i and the key pixel j within the window are represented by (i,j) and (j,j), respectively. For example, the coordinates of the top-left corner pixel of the window are (0,0), and M(i,j) represents the element value in the i-th row and j-th column of the mask matrix M. After superimposing the mask matrix M onto the original attention weights, the final attention distribution is generated through SoftMax normalization. That is, only key / value pixel interaction weights that satisfy both horizontal and vertical relative distances are retained, while the mask values ​​for other positions are -∞, approaching zero after SoftMax. Through the above-mentioned synergistic mechanism of local dense feature extraction and global sparse interaction, the network's ability to model image structural features can be enhanced, and noise interference can be suppressed. Finally, the fused features are compressed and integrated again through a second dilated feedforward network module, outputting a high-quality semantic representation for subsequent detection and decoding stages. The DTAB module, through multi-stage dilated convolution enhancement and attention modeling strategies, significantly improves the recognition ability of small targets, blurred boundary regions, and low-contrast regions in infrared images while ensuring low computational overhead.

[0107] In one implementation, such as Figure 7As shown, to further reduce the computational burden of the model and improve the consistency of expression and cross-scale perception among different feature maps, this invention designs the DGCST2 module. The DGCST2 module is a highly efficient feature fusion structure that integrates lightweight convolution and cross-channel interaction mechanisms to achieve compressed integration of multi-scale semantic information and enhanced spatial consistency. The DGCST2 module consists of two parts: a main structure and a nested DGSM module (Depthwise Group Shuffle Module). Each sub-path achieves collaborative modeling of different feature channels through proportional branching, concatenation, and convolutional fusion. First, the input features are divided into two parts according to channel dimension. 75% of the channels are directly processed by a 1×1 convolution for dimensionality transformation, serving as the main branch; the remaining 25% of the channels are fed into the DGSM sub-module for structured modeling. The DGSM module first divides this sub-channel into two again. One path retains the identity mapping, while the other path extracts fine-grained spatial structure features through 3×3 group convolution and introduces channel rearrangement operations to promote information interaction between different channels. Subsequently, the two sub-channels undergo channel concatenation and are integrated through 1×1 convolutions. The output of the DGSM module is concatenated with the output of the backbone path along the channel dimension to form a unified intermediate feature map. This fused feature is further input into the ConvFFN submodule (ConvolutionalFeed-Forward Network Submodule), which consists of two layers of 1×1 convolutions and introduces residual connections to improve feature consistency and deep expressive power, thereby generating the final feature representation. The DGCST2 module effectively improves the expressive independence and semantic interaction capabilities between feature channels through channel partitioning, multi-scale path modeling, grouped convolutions, and channel rearrangement operations. While maintaining low computational cost, it improves the fusion effect of multi-layer features and is one of the important components in this invention for achieving efficient feature interaction and fusion.

[0108] Step S300: Decode and predict the quality-enhanced feature map using the decoding prediction module to obtain the target category and bounding box coordinates.

[0109] In this embodiment, by decoding and predicting the quality-enhanced feature maps, the target's category information and bounding box coordinates can be effectively obtained. Specifically, the decoding and prediction module first performs layer-by-layer parsing on the three obtained quality-enhanced feature maps (i.e., the first, second, and third quality-enhanced feature maps) to gradually restore the target semantic information contained in the high-dimensional features. Simultaneously, in the prediction of bounding box coordinates, the decoder combines multi-scale feature information and a position-sensitive mechanism, enabling more accurate target location and reducing false positives and false negatives. Compared to traditional methods, this approach not only enhances the model's adaptability to complex scenes (such as occlusion and lighting changes) but also improves detection efficiency and overall performance, exhibiting higher robustness and practicality.

[0110] Specifically, step S300 includes the following steps:

[0111] Step S301: The quality enhancement feature map is processed through an embedded query selection mechanism to obtain a position-sensitive query vector;

[0112] Step S302: Decode and predict the query vector using a decoder and a prediction head to obtain the target's category and bounding box coordinates.

[0113] In one implementation, such as Figure 2 As shown, the three fused feature maps—the first, second, and third quality-enhanced feature maps—are fed into the decoding and prediction module. First, these enhanced feature maps are processed using an Inner-MPDIoU-aware Query Selection mechanism to generate location-sensitive query vectors. This process not only strengthens the model's understanding of target features at different scales and resolutions but also significantly improves the ability to accurately capture the target's location through location-sensitive processing. The location-sensitive query vectors can more effectively represent the relationship between various parts of the image and the target, thus laying a solid foundation for subsequent target classification and localization. Next, these query vectors are fed into the decoder and prediction head for further decoding and prediction, ultimately obtaining the target's category and bounding box coordinates. This two-stage approach, first extracting and transforming features through a carefully designed location-sensitive query selection mechanism and then utilizing an efficient decoder and prediction head architecture to complete the final recognition task, not only greatly improves the accuracy and efficiency of target detection but also enhances the system's robustness in challenging scenarios such as complex backgrounds, occlusion, or small targets. In this way, the entire system can achieve rapid response while ensuring high accuracy, making real-time target detection possible.

[0114] Throughout the process, a loss function was also designed. Regarding the design of the model's loss function, this invention introduces a novel IoU loss function—Inner-MPDIoU (Inner-region Maximum Probability Distance IoU: Inner-MPDIoU is an improved IoU evaluation metric that enhances the accuracy of object detection or segmentation by focusing on the maximum probability distance within the target's intra-region). This loss function integrates MPDIoU geometric sensitivity optimization with the Inner-IoU dynamic bounding box scaling mechanism, aiming to improve the accuracy and convergence efficiency of object detection in complex scenes. Its formula is defined as:

[0115]

[0116] L Inner-MPDIoU =L MPDIoU +IoU-IoU inner As shown in the formula above, and These represent the coordinates of the top-left and bottom-right corners of the actual auxiliary bounding box, respectively. and These represent the coordinates of the top-left and bottom-right corners of the predicted bounding box, respectively. `inter` represents the area of ​​the intersection between the predicted bounding box and the ground truth bounding box. `IoU` represents the area of ​​the original predicted bounding box B. pre With the original true frame B gt The ratio of the intersection region to the union region can be expressed as: The dynamic scaling factor `ratio` is used to generate auxiliary bounding boxes on top of the original bounding boxes. For high IoU samples (IoU > 0.5), `ratio < 1` is used to shrink the auxiliary boxes to accelerate convergence, while for low IoU samples (IoU ≤ 0.5), `ratio ≥ 1` is used to expand the auxiliary boxes to enhance regression robustness. The width and height of the generated true auxiliary boxes after scaling by `ratio` are represented as follows: Where w gt ,h gt These represent the width and height of the ground truth bounding box, and the width and height of the predicted auxiliary bounding box, respectively. As mentioned above, the union can be represented as the area of ​​the union region of the predicted auxiliary box and the ground truth auxiliary box, IoU. inner This is the intersection-over-union ratio (IoU) between the auxiliary predicted bounding boxes and the auxiliary ground truth bounding boxes. L MPDIoU This is represented as the original MPDIoU loss mechanism, where λ is an adjustable hyperparameter (default λ = 0.5) used to dynamically adjust corner constraints. and , where w and h are the squared Euclidean distances between the diagonal points of the predicted bounding box and the ground truth bounding box, respectively, and w and h represent the width and height of the input image, respectively. This constitutes an adjustable corner distance penalty term; please refer to the diagram for details. Figure 8 .

[0117] As in formula L Inner-MPDIoU =L MPDIoU +IoU-IoU inner As shown, in calculating the model loss function L Inner-MPDIoU At that time, the initial L MPDIoU The geometric sensitivity of the loss mechanism is preserved, where the intersection-union term IoU is... inner This is used to measure the degree of overlap between the auxiliary predicted bounding box and the auxiliary ground truth bounding box, while the corner distance penalty term... This is used to explicitly optimize the distance between the top-left and bottom-right corners of the predicted bounding box and the ground truth bounding box, thus avoiding the failure problem of traditional IoU when the aspect ratio is the same but the dimensions are different. Furthermore, the dynamic corner constraint term is adjusted through the hyperparameter λ to further enhance the stability of corner alignment. Finally, the initial L... MPDIoU Loss mechanism input L Inner-MPDIoU =L MPDIoU +IoU-IoU inner In this context, the model loss function can be further expressed as: By optimizing the loss function as described above, a balance between accuracy and efficiency can be achieved in real-time detection scenarios, providing a better solution for target localization in complex scenarios.

[0118] Furthermore, this invention also conducted corresponding experiments. In the model testing phase, the publicly available HIT-UAV dataset was used for model comparison experiments and performance verification. The proposed method (MDDI-RTDETR model) was compared with several other advanced object detection models, including the YOLO series models, the YOLO series algorithms (YOLOV8, YOLOV9, YOLOV10, YOLOV11) widely used in the field of computer vision object detection, the RT-DETR model in "DETRs Beat YOLOs on Real-time Object Detection" (CVPR2024), and the improved model PHSI-RTDETR for infrared small target detection ("PHSI-RTDETR: A Lightweight Infrared Small Target Detection Algorithm Based on UAV Aerial Photography"). The object detection evaluation metrics included precision, recall, average precision (mAP), and inference time. Precision measures the percentage of real targets predicted as positive by the model, reflecting the model's discrimination accuracy. Recall measures the proportion of real targets correctly detected by the model, reflecting the model's recall capability. mAP50 and mAP50:95 are commonly used comprehensive performance indicators, representing the average precision at an Intersection over Union (IoU) threshold of 0.5 and averaging from 0.5 to 0.95, respectively; higher values ​​indicate better overall model detection performance. Furthermore, inference time (in milliseconds, ms) measures the time required for the model to process a single image and is an important indicator for evaluating the model's real-time performance. These indicators collectively constitute a comprehensive evaluation of the target detection model in terms of accuracy, robustness, and efficiency. As shown in Table 1 below, compared with various advanced infrared small target detection models, the MDDI-RTDETR model proposed in this invention performs excellently in terms of detection accuracy and inference speed, outperforming current mainstream real-time infrared small target detection models.

[0119] Table 1. Model Comparison Test Results

[0120] Model Name Accuracy / % Recall rate / % mAP50 / % mAP50:95 / % Inference time (ms) YOLOV8 80.0 71.5 76.2 51.7 7.4 YOLOV9 84.0 72.9 79.1 51.2 8.1 YOLOV10 93.1 73.8 81.6 52.9 7.5 YOLOV11 83.2 73.0 80.7 52.0 7.7 RT-DETR 86.5 76.9 79.8 50.8 8.9 PHSI-RTDETR 89.9 76.1 82.6 51.6 7.9 MDDI-RTDETR 86.5 78.4 83.0 54.0 7.2

[0121] In summary, the MDDI-RTDETR model proposed in this invention achieves end-to-end infrared small target detection while balancing detection speed and accuracy. It is particularly suitable for resource-constrained UAV platforms to perform small target detection tasks in infrared images. It has significant improvements in accuracy, robustness, and deployment adaptability compared to existing mainstream detection algorithms (YOLO series, DETR series), and has good engineering application prospects and promotion value.

[0122] In summary, this embodiment first utilizes a feature extraction module to extract features from the input image, obtaining a multi-scale feature map. This step processes the input image using the feature extraction module, capturing detailed information at different levels through multi-scale analysis, resulting in a multi-scale feature map containing semantic information from low to high levels. This multi-level information representation helps improve the recognition ability of targets of different sizes, especially in complex backgrounds or in the presence of occlusion. Next, the feature fusion module performs feature enhancement processing on the multi-scale feature map to obtain a quality-enhanced feature map. This process integrates information from feature maps of different scales, effectively strengthening key features and suppressing noise, thereby generating a higher-quality feature map. This step is particularly important for improving the model's performance when facing small targets or fine-grained features, because high-quality feature representation can more accurately describe the characteristics and contextual information of the target. Then, the decoding and prediction module decodes and predicts the quality-enhanced feature map to obtain the target's category and bounding box coordinates. This module transforms the abstract features learned by the deep learning model into concrete and interpretable results, that is, determining the specific location of the target in the image and its category. This not only improves the accuracy of localization but also enhances the precision of classification, making the final detection results more reliable. Overall, the method proposed in this invention forms an efficient target detection process by combining three stages: multi-scale feature extraction, feature fusion, and accurate decoding prediction. This method can significantly improve the detection performance of targets of various sizes, especially the robustness and accuracy in complex environments, while also improving the detection effect of small targets, providing strong support for practical applications.

[0123] like Figure 9As shown in the illustration, this embodiment also provides an infrared small target detection system based on UAV aerial photography scenarios. The system includes: a multi-scale feature map acquisition module 10, a quality enhancement feature map acquisition module 20, and a category and bounding box coordinate acquisition module 30. Specifically, the multi-scale feature map acquisition module 10 is used to extract features from the input image using a feature extraction module to obtain a multi-scale feature map. The quality enhancement feature map acquisition module 20 is used to process the multi-scale feature map based on a feature fusion module to obtain a quality enhancement feature map. The category and bounding box coordinate acquisition module 30 is used to decode and predict the quality enhancement feature map using a decoding and prediction module to obtain the target's category and bounding box coordinates.

[0124] In one implementation, the multi-scale feature map acquisition module 10 includes:

[0125] The initial feature map acquisition unit is used to perform preliminary feature extraction on the input image using the Stem module to obtain an initial feature map.

[0126] The first multi-scale feature map acquisition unit is used to extract features from the initial feature map based on the Conv module, HGBlock module, C2f module and DAF module to obtain the first multi-scale feature map;

[0127] The second multi-scale feature map acquisition unit is used to extract features from the first multi-scale feature map based on the Conv module, HGBlock module, C2f module and DAF module to obtain the second multi-scale feature map.

[0128] The third multi-scale feature map acquisition unit is used to extract features from the second multi-scale feature map based on the Conv module, HGBlock module, C2f module and DAF module to obtain the third multi-scale feature map.

[0129] In one implementation, the first multi-scale feature map acquisition unit includes:

[0130] The first branch feature map acquisition subunit is used to sequentially use the C2f module, the Conv module and the C2f module to extract features from the initial feature map to obtain the first branch feature map.

[0131] The second branch feature map acquisition subunit is used to extract features from the initial feature map based on the HGBlock module, the DWConv module and the HGBlock module in sequence to obtain the second branch feature map.

[0132] The first multi-scale feature map acquisition subunit is used to extract features from the first branch feature map and the second branch feature map using the DAF module to obtain the first multi-scale feature map.

[0133] In one implementation, the second multi-scale feature map acquisition unit includes:

[0134] The third branch feature map acquisition subunit is used to extract features from the first multi-scale feature map by sequentially using the Conv module and the C2f module to obtain the third branch feature map.

[0135] The fourth branch feature map acquisition subunit is used to extract features from the initial feature map based on the DWConv module and the three HGBlock modules in sequence to obtain the fourth branch feature map.

[0136] The second multi-scale feature map acquisition subunit is used to extract features from the third branch feature map and the fourth branch feature map using the DAF module to obtain the second multi-scale feature map.

[0137] In one implementation, the third multi-scale feature map acquisition unit includes:

[0138] The fifth branch feature map acquisition subunit is used to extract features from the second multi-scale feature map by sequentially using the Conv module and the C2f module to obtain the fifth branch feature map.

[0139] The sixth branch feature map acquisition subunit is used to extract features from the second multi-scale feature map based on the DWConv module and the HGBlock module in sequence to obtain the sixth branch feature map.

[0140] The third multi-scale feature map acquisition subunit is used to extract features from the fifth branch feature map and the sixth branch feature map using the DAF module to obtain the third multi-scale feature map.

[0141] In one implementation, the quality enhancement feature map acquisition module 20 includes:

[0142] The first quality-enhanced feature map acquisition unit is used to process the multi-scale feature map based on the Conv module, Upsample module, DTAB module and DGCST2 module to obtain the first quality-enhanced feature map;

[0143] The second quality enhancement feature map acquisition unit is used to perform convolution and fusion processing on the first quality enhancement feature map using the Conv module and the DGCST2 module to obtain the second quality enhancement feature map.

[0144] The third quality enhancement feature map acquisition unit is used to perform convolution and fusion processing on the second quality enhancement feature map using the Conv module and the DGCST2 module to obtain the third quality enhancement feature map.

[0145] The quality enhancement feature map acquisition unit is used to perform a fusion operation on the first quality enhancement feature map, the second quality enhancement feature map, and the third quality enhancement feature map to obtain a quality enhancement feature map.

[0146] In one implementation, the first quality-enhanced feature map acquisition unit includes:

[0147] The preliminary quality enhancement feature map acquisition subunit is used to process the third multi-scale feature map sequentially based on the Conv module, DTAB module, and Conv module to obtain the preliminary quality enhancement feature map;

[0148] The first concatenated feature map obtains a first sub-unit, which is used to perform upsampling and convolution processing using the Upsample module and Conv module based on the preliminary quality enhancement feature map and the second multi-scale feature map, and to perform a Concat operation to obtain the first concatenated feature map;

[0149] The convolutional fusion feature map acquisition sub-unit is used to perform fusion and convolution processing on the first concatenated feature map based on the DGCST2 module and the Conv module to obtain the convolutional fusion feature map.

[0150] The second concatenated feature map obtains the first sub-unit, which is used to perform upsampling and convolution processing using the Upsample module and Conv module based on the convolutional fusion feature map and the first multi-scale feature map, and to perform a Concat operation to obtain the second concatenated feature map.

[0151] The first quality enhancement feature map acquisition subunit is used to fuse the second spliced ​​feature map using the DGCST2 module to obtain the first quality enhancement feature map.

[0152] In one implementation, obtaining the first sub-unit from the first concatenated feature map includes:

[0153] The sampled preliminary quality enhancement feature map acquisition subunit is used to upsample the preliminary quality enhancement feature map using the Upsample module to obtain the sampled preliminary quality enhancement feature map;

[0154] The second multi-scale feature map acquisition sub-unit is used to perform convolution processing on the second multi-scale feature map using the Conv module to obtain the convolutional second multi-scale feature map.

[0155] The first concatenated feature map obtains a second sub-unit, which is used to concatenate the sampled preliminary quality enhancement feature map and the convolutional second multi-scale feature map to obtain the first concatenated feature map.

[0156] In one implementation, the second concatenated feature map obtaining the first sub-unit includes:

[0157] The sampled convolutional fusion feature map acquisition sub-unit is used to upsample the convolutional fusion feature map using the Upsample module to obtain the sampled convolutional fusion feature map;

[0158] The first multi-scale feature map acquisition sub-unit is used to perform convolution processing on the first multi-scale feature map using the Conv module to obtain the first multi-scale feature map.

[0159] The second concatenated feature map obtains the second sub-unit, which is used to concatenate the sampled convolutional fusion feature map and the first convolutional multi-scale feature map to obtain the second concatenated feature map.

[0160] In one implementation, the second quality-enhanced feature map acquisition unit includes:

[0161] The first quality enhancement feature map acquisition sub-unit is used to perform convolution processing on the first quality enhancement feature map based on the Conv module to obtain the first quality enhancement feature map by convolution.

[0162] The third concatenated feature map acquisition subunit is used to concat the convolutional fusion feature map with the convolutional first quality enhancement feature map to obtain the third concatenated feature map;

[0163] The second quality enhancement feature map acquisition subunit is used to perform fusion processing on the third stitched feature map using the DGCST2 module to obtain the second quality enhancement feature map.

[0164] In one implementation, the third quality-enhanced feature map acquisition unit includes:

[0165] The convolutional second quality enhancement feature map acquisition sub-unit is used to perform convolution processing on the second quality enhancement feature map based on the Conv module to obtain the convolutional second quality enhancement feature map;

[0166] The fourth concatenated feature map acquisition subunit is used to concatenate the second quality enhancement feature map of the convolution with the preliminary quality enhancement feature map to obtain the fourth concatenated feature map;

[0167] The third quality enhancement feature map acquisition subunit is used to fuse the fourth stitched feature map using the DGCST2 module to obtain the third quality enhancement feature map.

[0168] In one implementation, the category and bounding box coordinate acquisition module 30 includes:

[0169] The query vector acquisition unit is used to process the quality-enhanced feature map through an embedded query selection mechanism to obtain a position-sensitive query vector;

[0170] The category and bounding box coordinate acquisition unit is used to decode and predict the query vector through a decoder and a prediction head to obtain the category and bounding box coordinates of the target.

[0171] The working principle of each module in the infrared small target detection system based on UAV aerial photography scenarios in this embodiment is the same as that of each step in the above method embodiment, and will not be repeated here.

[0172] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 10 As shown. The terminal may include one or more processors 100 ( Figure 10 (Only one is shown in the image), a memory 101, and a computer program 102 stored in the memory 101 and executable on one or more processors 100, such as an infrared small target detection program based on a drone aerial photography scene. When one or more processors 100 execute the computer program 102, they can implement the various steps in the embodiment of the infrared small target detection method based on a drone aerial photography scene. Alternatively, when one or more processors 100 execute the computer program 102, they can implement the functions of each module / unit in the embodiment of the infrared small target detection method based on a drone aerial photography scene, which is not limited here.

[0173] In one embodiment, the processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0174] In one embodiment, memory 101 may be an internal storage unit of an electronic device, such as a hard drive or RAM. Memory 101 may also be an external storage device of the electronic device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc. Furthermore, memory 101 may include both internal and external storage units. Memory 101 is used to store computer programs and other programs and data required by the terminal. Memory 101 can also be used to temporarily store data that has been output or will be output.

[0175] Those skilled in the art will understand that Figure 10 The schematic diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal to which the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0176] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the methods described above. Any references to memory, storage, operational databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual operating data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting small infrared targets based on drone aerial photography scenarios, characterized in that, The method includes: The feature extraction module is used to extract features from the input image to obtain a multi-scale feature map; The feature fusion module is used to perform feature enhancement processing on the multi-scale feature map to obtain a quality-enhanced feature map. The enhanced feature map is decoded and predicted by the decoding and prediction module to obtain the target's category and bounding box coordinates.

2. The infrared small target detection method based on UAV aerial photography scenes according to claim 1, characterized in that, The feature extraction module includes a Stem module, a Conv module, an HGBlock module, a C2f module, and a DAF module. The feature extraction module extracts features from the input image to obtain a multi-scale feature map, including: The Stem module is used to perform preliminary feature extraction on the input image to obtain an initial feature map; Based on the Conv module, HGBlock module, C2f module and DAF module, feature extraction is performed on the initial feature map to obtain the first multi-scale feature map; Based on the Conv module, HGBlock module, C2f module and DAF module, feature extraction is performed on the first multi-scale feature map to obtain the second multi-scale feature map; Based on the Conv module, HGBlock module, C2f module and DAF module, feature extraction is performed on the second multi-scale feature map to obtain the third multi-scale feature map.

3. The infrared small target detection method based on UAV aerial photography scenes according to claim 2, characterized in that, The first multi-scale feature map is obtained by extracting features from the initial feature map using the Conv module, HGBlock module, C2f module, and DAF module, including: The initial feature map is extracted sequentially using the C2f module, the Conv module, and the C2f module to obtain the first branch feature map; The initial feature map is then processed sequentially using the HGBlock module, the DWConv module, and the HGBlock module to obtain the second branch feature map. Using the DAF module, feature extraction is performed on the first branch feature map and the second branch feature map to obtain the first multi-scale feature map.

4. The infrared small target detection method based on UAV aerial photography scenarios according to claim 3, characterized in that, The second multi-scale feature map is obtained by extracting features from the first multi-scale feature map based on the Conv module, HGBlock module, C2f module, and DAF module, including: The first multi-scale feature map is extracted by sequentially using the Conv module and the C2f module to obtain the third branch feature map; The initial feature map is then extracted sequentially using the DWConv module and three HGBlock modules to obtain the fourth branch feature map. Using the DAF module, feature extraction is performed on the third branch feature map and the fourth branch feature map to obtain the second multi-scale feature map.

5. The infrared small target detection method based on UAV aerial photography scenes according to claim 4, characterized in that, The third multi-scale feature map is obtained by extracting features from the second multi-scale feature map based on the Conv module, HGBlock module, C2f module, and DAF module, including: The Conv module and C2f module are used sequentially to extract features from the second multi-scale feature map to obtain the fifth branch feature map; The second multi-scale feature map is then extracted sequentially using the DWConv module and the HGBlock module to obtain the sixth branch feature map. Using the DAF module, feature extraction is performed on the fifth branch feature map and the sixth branch feature map to obtain the third multi-scale feature map.

6. The infrared small target detection method based on UAV aerial photography scenes according to claim 5, characterized in that, The feature fusion module performs feature enhancement processing on the multi-scale feature map to obtain a quality-enhanced feature map. The feature fusion module includes a Conv module, an Upsample module, a DTAB module, and a DGCST2 module. The multi-scale feature map is processed using the Conv module, Upsample module, DTAB module, and DGCST2 module to obtain the first quality-enhanced feature map. Based on the first quality-enhanced feature map, the Conv module and the DGCST2 module are used for convolution and fusion processing to obtain the second quality-enhanced feature map. Based on the second quality enhancement feature map, the third quality enhancement feature map is obtained by performing convolution and fusion processing using the Conv module and the DGCST2 module. The first quality enhancement feature map, the second quality enhancement feature map, and the third quality enhancement feature map are fused to obtain a quality enhancement feature map.

7. The infrared small target detection method based on UAV aerial photography scene according to claim 6, characterized in that, The first quality-enhanced feature map is obtained by processing the multi-scale feature map using the Conv module, Upsample module, DTAB module, and DGCST2 module, including: The third multi-scale feature map is processed sequentially using the Conv module, DTAB module, and Conv module to obtain a preliminary quality-enhanced feature map. Based on the preliminary quality enhancement feature map and the second multi-scale feature map, the Upsample module and Conv module are used for upsampling and convolution processing, and a Concat operation is performed to obtain the first concatenated feature map. Based on the DGCST2 module and the Conv module, the first concatenated feature map is fused and convolved to obtain a convolutional fused feature map. Based on the convolutional fusion feature map and the first multi-scale feature map, the Upsample module and Conv module are used to perform upsampling and convolution processing, and then a Concat operation is performed to obtain the second concatenated feature map. The second stitched feature map is fused using the DGCST2 module to obtain the first quality-enhanced feature map.

8. The infrared small target detection method based on UAV aerial photography scene according to claim 7, characterized in that, Based on the preliminary quality-enhanced feature map and the second multi-scale feature map, upsampling and convolution processing are performed using the Upsample and Conv modules, followed by a Concat operation to obtain the first concatenated feature map, including: The Upsample module is used to upsample the preliminary quality enhancement feature map to obtain a sampled preliminary quality enhancement feature map; The second multi-scale feature map is convolved using the Conv module to obtain the convolved second multi-scale feature map. The sampled preliminary quality-enhanced feature map and the convolutional second multi-scale feature map are concatenated to obtain the first concatenated feature map.

9. The infrared small target detection method based on UAV aerial photography scene according to claim 7, characterized in that, The second concatenated feature map is obtained by performing upsampling and convolution processing using the Upsample and Conv modules based on the convolutional fused feature map and the first multi-scale feature map, followed by a Concat operation. The Upsample module is used to upsample the convolutional fusion feature map to obtain the sampled convolutional fusion feature map; Using the Conv module, the first multi-scale feature map is convolved to obtain the convolved first multi-scale feature map; The sampled convolutional fusion feature map and the first convolutional multi-scale feature map are concatenated to obtain the second concatenated feature map.

10. The infrared small target detection method based on UAV aerial photography scenes according to claim 6, characterized in that, The process of obtaining a second quality-enhanced feature map based on the first quality-enhanced feature map, using the Conv module and the DGCST2 module for convolution and fusion processing, includes: Based on the Conv module, the first quality enhancement feature map is convolved to obtain the convolved first quality enhancement feature map; The convolutional fusion feature map and the first convolutional quality enhancement feature map are concatenated to obtain the third concatenated feature map; The third stitched feature map is fused using the DGCST2 module to obtain the second quality-enhanced feature map.

11. The infrared small target detection method based on UAV aerial photography scene according to claim 7, characterized in that, The process of obtaining a third quality-enhanced feature map based on the second quality-enhanced feature map involves convolution and fusion using the Conv module and the DGCST2 module, including: Based on the Conv module, the second quality enhancement feature map is convolved to obtain the convolutional second quality enhancement feature map; The second quality enhancement feature map of the convolution is concatenated with the preliminary quality enhancement feature map to obtain the fourth concatenated feature map; The fourth stitched feature map is fused using the DGCST2 module to obtain the third quality-enhanced feature map.

12. The infrared small target detection method based on UAV aerial photography scenes according to claim 1, characterized in that, The step of decoding and predicting the quality-enhanced feature map through the decoding and prediction module to obtain the target's category and bounding box coordinates includes: The quality-enhanced feature map is processed through an embedded query selection mechanism to obtain a position-sensitive query vector; The query vector is decoded and predicted using a decoder and a prediction head to obtain the target's category and bounding box coordinates.

13. An infrared small target detection system based on UAV aerial photography scenarios, characterized in that, The system includes: The multi-scale feature map acquisition module is used to extract features from the input image using the feature extraction module to obtain a multi-scale feature map; The quality enhancement feature map acquisition module is used to process the multi-scale feature map based on the feature fusion module to obtain the quality enhancement feature map; The category and bounding box coordinate acquisition module is used to decode and predict the quality-enhanced feature map through the decoding and prediction module to obtain the category and bounding box coordinates of the target.

14. A terminal, characterized in that, The terminal includes a memory, a processor, and an infrared small target detection program based on a drone aerial photography scene stored in the memory and executable on the processor. When the processor executes the infrared small target detection program based on a drone aerial photography scene, it implements the steps of the infrared small target detection method based on a drone aerial photography scene as described in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an infrared small target detection program based on a drone aerial photography scene. When the infrared small target detection program based on a drone aerial photography scene is executed by a processor, it implements the steps of the infrared small target detection method based on a drone aerial photography scene as described in any one of claims 1-12.