An infrared image unmanned aerial vehicle target detection method and device based on complementary design, and a medium
By adding a low-level feature detail enhancement module and a high-level feature semantic preservation module to the YOLOv8 network, the problems of low-level features being easily affected by background interference and high-level feature semantic information loss in small target detection in infrared images are solved, and robust detection under complex backgrounds is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 713TH RES INST OF CHINA STATE SHIPBUILDING CORP LTD
- Filing Date
- 2026-03-26
- Publication Date
- 2026-06-26
AI Technical Summary
Existing UAV small target detection methods lack modeling of image characteristics in infrared images, resulting in low-level features being easily affected by background interference and unstable structural representation, while high-level features lose semantic information, making it difficult to accurately detect small targets in complex backgrounds.
In the backbone and neck network of the YOLOv8 network, a low-level feature detail enhancement module and a high-level feature semantic preservation module are added respectively. The low-level features are enhanced by random perturbation convolution and frequency domain reconstruction. The edge and structural information of small targets are explicitly extracted and enhanced in the high-level features by combining low-pass branch and multi-directional gradient filter group. Column direction attention and context self-calibration are also introduced.
It significantly improves the accuracy and robustness of small target detection in infrared images, enabling robust detection under complex backgrounds and weak textures, preserving the edge structure information of small targets, suppressing background interference, and improving positioning accuracy.
Smart Images

Figure CN122289650A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of target detection technology, specifically relating to an infrared image-based UAV target detection method, device, and medium based on complementary design. Background Technology
[0002] In infrared images, small targets such as UAVs are typically small in size and have low contrast, making them easily obscured by numerous complex background features in conventional detection networks. This leads to missed detections or inaccurate localization, especially in applications with drastic scale changes and complex backgrounds, where detection robustness significantly decreases. Existing problems in UAV small target detection using infrared images can be summarized into the following two main aspects:
[0003] Existing methods do not pay enough attention to low-level feature details: The low-level features extracted by the detection model mainly include structural information such as the target's edges and shape. This information is crucial for the detection of small targets by infrared UAVs and is easily affected by background interference and unstable structural representation. However, existing detection methods do not pay enough attention to the low-level features of the model. During training, low-level features are prone to overfitting to the background texture, resulting in unstable structural representation in different scenarios. Small targets on UAVs are easily overlooked or misidentified.
[0004] Degradation of semantic information in high-level features: Existing detection methods lack consistent modeling in the extraction of mid-to-high-level features. In the absence of rich texture details in infrared images, the semantic information of small target regions in high-level features is easily lost or damaged. The consistency of contextual expression ability among high-level features is insufficient, and problems such as target structure breakage and blurred boundaries often occur, making it difficult to accurately distinguish between the target and the background.
[0005] The Chinese invention patent authorization announcement, CN120032271B, dated December 26, 2025, discloses a method for detecting small targets on unmanned aerial vehicles (UAVs) inspired by the eagle-eye vision mechanism. This method utilizes the structural differences in the fovea region of an eagle's eye, designing two functional modules to guide spatial attention in the neural network structure, thereby improving small target detection performance. However, this method is primarily based on visible light image detection and is not applicable to infrared images. Furthermore, it lacks the ability to enhance low-level feature details and preserve high-level feature semantics, leading to the loss of boundary details in low-level UAV small target features and the loss of semantic information in high-level features.
[0006] Chinese invention patent authorization announcement number CN120953858B, with an authorization announcement date of December 26, 2025, discloses a method for small target recognition of unmanned aerial vehicles (UAVs) based on wavelet decomposition and motion vectors. This method enhances multiple high-frequency detail information regions in the small target image, then reconstructs the image based on the enhanced image, and finally detects the UAV based on the reconstructed image, thus improving the accuracy of the target recognition algorithm. However, this method is mainly designed for visible light images and relies on significant features such as texture details, motion boundaries, or grayscale changes of the target in the image. It is not suitable for infrared images. Directly applying a high-frequency enhancement and low-frequency reconstruction strategy to infrared images may lead to false enhancement or blurred edges of infrared small targets, thus affecting detection performance. Furthermore, it lacks enhancement of shallow features of infrared UAV small targets and reinforcement of edge and semantic structure information of small targets, making it sensitive to infrared background noise and posing a risk of feature overload.
[0007] The Chinese invention patent authorization announcement, CN119888542B, dated November 28, 2025, discloses a method for small target detection on UAVs based on an improved YOLOv8s model. This method embeds a BiFormer module into the C2f module of Backbone, introducing the concept of capturing long-range dependencies and preserving fine-grained contextual feature information to improve the model's detection accuracy for small targets. While this approach performs well in natural light environments, it suffers from limitations in infrared images, including a small grayscale dynamic range, weak contrast, and sparse texture. Furthermore, small targets rely on low-level spatial features and are prone to failure during training due to complex backgrounds, thus compromising the stability of small target detection and making robust detection of small targets in infrared backgrounds difficult. Moreover, its network model lacks the ability to handle background interference and weak target structure in infrared images. Infrared small targets are often close in intensity to the background, with blurred edges that are easily misdetected or missed, limiting its detection capability.
[0008] Chinese invention patent authorization announcement number CN119478739B, with an authorization announcement date of October 28, 2025, discloses a method for detecting small targets on unmanned aerial vehicles (UAVs). This method utilizes channel attention and spatial attention to filter important feature data and uses self-attention calculation to obtain multi-granularity feature information, thereby improving the localization and detection capabilities of small targets. However, this scheme also relies on visible light image processing methods and is difficult to adapt to the feature representation of infrared imaging data. It also lacks sufficient preservation of detailed structural information of small targets, and during downsampling and multi-layer feature extraction, it lacks a special protection mechanism for small target features, which may lead to weakened features and blurred boundaries, affecting positioning accuracy.
[0009] In summary, existing UAV small target detection methods have the following significant drawbacks in infrared image scenarios, especially when dealing with issues such as "weak texture, low contrast, and difficulty in separating small targets from complex backgrounds":
[0010] ① Lack of modeling mechanisms for infrared image characteristics: Existing UAV small target detection methods are mostly designed based on visible light image data, relying on image texture, color, grayscale gradient, and other information as references for target saliency. However, in infrared images, due to differences in imaging mechanisms, targets typically exhibit the following characteristics: no obvious color differences; blurred edges and degraded shapes; weak grayscale contrast and strong background perturbation. Existing technologies do not structurally model these problems in infrared images, leading to potential performance degradation of detection models in infrared scenes, making them unable to adapt to real-world environments with complex lighting, distance variations, and uneven thermal signals.
[0011] ② Lack of consideration for the stability and robustness of low-level network features: Small targets in infrared images are typically very small, heavily relying on low-level structural information such as details and edges in shallow feature maps. However, existing detection schemes generally neglect the robust training of low-level features, resulting in the following problems: During training, shallow features are easily affected by background patterns, overfitting to the fixed environment in the training set; lack of constraints on structural expressive power leads to the failure of small target features during inference.
[0012] ③ Insufficient semantic preservation of mid-to-high-level semantic features: Small targets in infrared images are often "submerged" or "fragmented" in high-level semantic feature maps, especially in complex backgrounds or after downsampling, resulting in severe loss of original structural information. Existing detection methods mainly integrate information through general multi-scale feature fusion, which has the following problems: lack of perception of the continuity of feature structure; lack of adaptive modeling mechanism for long-range context; and easy to cause the disappearance of small target boundaries or positional shift, reducing positioning accuracy.
[0013] ④ The structural enhancement methods are not systematic enough and lack low-to-high level coordination mechanisms: Existing technologies mostly optimize the structure on the detection backbone or feature fusion path, lacking a systematic design approach from low-level detail protection to high-level semantic enhancement, and have not formed a unified coordination mechanism. Summary of the Invention
[0014] The purpose of this invention is to provide a method, device, and medium for infrared image-based unmanned aerial vehicle target detection based on complementary design, so as to solve the technical problem that existing target detection methods lack modeling of infrared image characteristics.
[0015] To address the aforementioned technical problems, the first aspect of this invention provides an infrared image-based UAV target detection method based on complementary design, the method comprising:
[0016] S1. Acquire the infrared image to be detected;
[0017] S2. Input the infrared image into a pre-trained infrared image drone detection model to obtain the drone category and location information in the infrared image;
[0018] The infrared image drone detection model is an improvement on the YOLOv8 network. The improvement includes adding low-level feature detail enhancement modules between the first C2f layer and the third Conv layer and between the second C2f layer and the fourth Conv layer of the backbone network, respectively.
[0019] The low-level feature detail enhancement module includes: enhancing the feature map output by the C2f layer. Apply random perturbation to obtain perturbation feature map In the frequency domain, based on the feature map Corresponding phase and perturbation feature maps The corresponding amplitude is used for spectrum reconstruction; the reconstructed frequency domain features are then restored to the feature map. The spatial domain is used to obtain the feature map output by the low-level feature detail enhancement module.
[0020] In one possible implementation, the improvement further includes adding high-level feature semantic preservation modules between the first C2f layer and the second upsampling layer of the neck network, and between the second C2f layer and the first Conv layer, respectively.
[0021] The high-level feature semantic preservation module includes: processing the input feature map Pooling is performed to obtain low-pass downsampled features; depthwise separable convolutions are used to extract the input feature maps. Gradient response characteristics in each direction, and corresponding gradient magnitude characteristics calculated based on gradient response characteristics;
[0022] The low-pass downsampling features are concatenated with the gradient magnitude features along the channel dimension, and the number of channels in the concatenated feature map is compressed to the input feature map. The number of channels is used to obtain directional enhancement features;
[0023] Input feature map After aligning with the size of the directional enhancement feature, feature fusion is performed with the directional enhancement feature to obtain the fused feature;
[0024] After the fused features are sequentially processed by column-direction attention and contextual semantic self-calibration, the feature size is restored to the input feature map by upsampling. The size is determined to obtain the feature map output by the high-level feature semantic preservation module.
[0025] In one possible implementation, the spectrum reconstruction is a weighted spectrum reconstruction as shown below:
[0026]
[0027] in, For reconstructed frequency domain features; These are learnable weight parameters; Perturbation feature map The corresponding amplitude; For feature map Corresponding phase; For complex number operators.
[0028] In one possible implementation, the perturbation feature map The feature map is obtained as follows: Each channel is independently convolved to obtain a perturbation feature map. The convolution kernel for the two-dimensional convolution operation is a randomly generated depth-separable convolution kernel that is randomly sampled using a Laplace distribution.
[0029] In one possible implementation, the gradient magnitude feature is obtained as follows:
[0030]
[0031]
[0032]
[0033] in, Features of the gradient magnitude in the horizontal direction; The feature is the gradient magnitude in the vertical direction; The gradient magnitude is a characteristic of the diagonal direction; The characteristics are those of the gradient response in the horizontal direction; This refers to the gradient response characteristics in the vertical direction; The gradient response characteristics are diagonal. This represents element-wise multiplication; It is a positive number whose value is unstable when taking the square root.
[0034] In one possible implementation, the orientation enhancement feature is based on the stitched feature map via... The result is obtained after convolution.
[0035] In one possible implementation, the column-directed attention involves: simultaneously calculating the mean of the fused features in both the height and channel dimensions to generate a column description vector; and sequentially passing the column description vector through... After convolution and activation function processing, column-direction attention weights are obtained; the attention weights are broadcast in the height and channel dimensions and then multiplied element-wise with the fused features to obtain the output features of column-direction attention.
[0036] In one possible implementation, the context semantic self-calibration is as follows: the output features of column direction attention are sequentially processed by downsampling compression, convolution transformation and upsampling feature recovery to obtain context compensation features; then, the context compensation features are superimposed with the output features of column direction attention to obtain the output features of context semantic self-calibration.
[0037] To address the aforementioned technical problems, a second aspect of the present invention provides an infrared image-based unmanned aerial vehicle target detection device based on complementary design, comprising a processor for executing a computer program to implement the steps of the method in any possible implementation of the first aspect of the present invention.
[0038] To address the aforementioned technical problems, a third aspect of the present invention provides a computer-readable storage medium storing a computer program internally, the computer program being executed by a processor to implement the steps of the method in any possible implementation of the first aspect of the present invention.
[0039] The beneficial effects of this invention are as follows: This invention addresses the characteristic of infrared image target detection that heavily relies on low-level structural information in shallow feature maps. For the backbone network used for feature extraction, a low-level feature detail enhancement module is added after the first two C2f modules. Since the "phase" part of an image determines its structure (edges and shape) in the frequency domain, this invention preserves the original phase after perturbation to maintain the target structure. The low-level feature detail enhancement module enhances the expressive power of shallow feature maps for small infrared targets through random perturbation convolution and frequency domain reconstruction mechanisms. This module can effectively preserve the edge structure information of small targets even in complex background noise and low target contrast conditions, while improving robustness to complex backgrounds, providing a more robust feature representation for subsequent detection tasks. This invention solves the technical problem of existing target detection methods lacking modeling of infrared image characteristics.
[0040] This invention also addresses the issue that small target features in infrared image target detection are often "submerged" or "fragmented" in high-level semantic feature maps, resulting in severe loss of original structural information. For the neck network used for feature fusion, a high-level feature semantic preservation module is added after the first two C2f modules. This module explicitly extracts and enhances the edge and structural information of small targets in high-level features through a "low-pass branch + multi-directional gradient filter group" approach. Based on residual semantic compensation, column-direction attention is introduced to weight features column by column, enhancing the structural continuity of small target features in infrared scenes and suppressing background interference. Simultaneously, a semantic compensation with a larger receptive field is established by combining a context self-calibration branch, making small target features under weak texture and low contrast conditions less prone to being submerged or fragmented during multi-scale fusion, thereby improving the localization accuracy and robustness of small target detection in infrared UAVs.
[0041] By complementing the design of the low-level feature detail enhancement module and the high-level feature semantic preservation module, the accuracy and robustness of small target detection in infrared images are significantly improved. The detail preservation capability of low-level features and the semantic consistency of high-level features are enhanced, and robust detection under complex background and weak texture conditions is achieved. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the improved YOLOv8 network in the implementation of the infrared image UAV target detection method based on complementary design of the present invention;
[0043] Figure 2 This is the low-level feature detail enhancement module network framework in the implementation of the infrared image UAV target detection method based on complementary design of the present invention;
[0044] Figure 3 This is the high-level feature semantic preservation module network framework in the implementation of the infrared image UAV target detection method based on complementary design of the present invention;
[0045] Figure 4 This is a flowchart illustrating the method implementation of the infrared image UAV target detection method based on complementary design according to the present invention.
[0046] Figure 5 This is a schematic diagram of the device architecture in an embodiment of the infrared image UAV target detection device based on complementary design of the present invention. Detailed Implementation
[0047] This invention addresses the issue of infrared image target detection heavily relying on low-level structural information in shallow feature maps. For the backbone network used for feature extraction, a low-level feature detail enhancement module is added after the first two C2f modules. Since the "phase" component of an image determines its structure (edges and shape) in the frequency domain, this invention preserves the original phase after perturbation to maintain the target structure. The low-level feature detail enhancement module enhances the expressive power of shallow feature maps for small infrared targets through random perturbation convolution and frequency domain reconstruction mechanisms. This module effectively preserves the edge structure information of small targets even in complex background noise and low target contrast conditions, while improving robustness to complex backgrounds and providing a more robust feature representation for subsequent detection tasks. This invention solves the technical problem of existing target detection methods lacking modeling of infrared image characteristics.
[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings.
[0049] Explanation of technical terms:
[0050] Infrared images: typically single-channel (grayscale) images, lacking color and clear texture, with low contrast, making it easy to confuse the target with the background.
[0051] Small objective: The pixel size in the input image is very small (usually less than 32×32), and its key information is easily lost after downsampling, so shallow features are needed to participate in the detection.
[0052] Phase: In the frequency domain, the "phase" part of an image determines the structure (edges and shape). This invention preserves the original phase after perturbation to keep the target structure unchanged.
[0053] Column-oriented attention: an attention mechanism that establishes global perception along the vertical direction (column).
[0054] Context self-calibration: global background information is obtained through downsampling → convolution → upsampling, and then fed back to local regions to achieve consistent adjustment of spatial semantics.
[0055] Implementation of an infrared image-based UAV target detection method based on complementary design:
[0056] This invention presents an infrared image UAV target detection method based on complementary design, utilizing the YOLOv8 detection framework and introducing two innovative modules: a low-level feature detail enhancement module and a high-level feature semantic preservation module. These two modules operate on the shallow and deep feature paths of the model, respectively, thereby constructing feature representations with robust detail and semantic preservation capabilities, improving the accuracy and robustness of small UAV target detection in infrared images.
[0057] like Figure 4 As shown, the infrared image UAV target detection method based on complementary design in this embodiment includes the following steps:
[0058] Step 1: Input image data acquisition and preprocessing.
[0059] 1.1 Dataset Preparation
[0060] This dataset is a supervised learning dataset for UAV target detection tasks using infrared images. It mainly contains infrared images and their annotation information. Each image in the dataset is accompanied by annotation information for target localization and classification.
[0061] Infrared UAV Detection Supervised Learning Dataset Represented as: ;in, This represents the total number of samples in the dataset; For sample index, ; For the first One infrared image sample (original input image); For the first Annotation information for each infrared image sample:
[0062]
[0063] in, For the first The number of labeled targets in the image. For the target index, , For the first The first image Category labels for each target. For the first The first image The bounding box parameters for each target are represented as follows:
[0064]
[0065] in, The x-coordinate of the center point of the target box; The ordinate of the center point of the target bounding box; The width of the target bounding box; The height of the target bounding box.
[0066] 1.2 Image Preprocessing and Enhancement
[0067] The input image data is a preprocessed single-channel infrared image. The preprocessing module enhances image features through a series of image enhancement operations (such as normalization, flipping, affine transformation, etc.) to improve the accuracy and robustness of small target detection. For details on how to perform image enhancement, please refer to existing technologies.
[0068] Preprocessed image The input tensor is formed as , Number of channels; Let the height be the tensor height. Where is the tensor width. During subsequent training, multiple image tensors will be input into the model in batches for training. A batch of input tensors is represented as: ,in: The batch size represents the number of samples processed in each training batch. The number of tensor channels; Let the height be the tensor height. The width of the tensor;
[0069] Step 2: The backbone network (i.e., the main network) performs feature extraction.
[0070] The preprocessed image data enters the Backbone network, such as... Figure 1 As shown, multiple Conv layers, C2f layers, and SPPF layers are used to extract low-level features (including edges and textures) and high-level features (including abstract semantic and positional information).
[0071] Step 3: Low-level feature detail enhancement module.
[0072] This step, as a processing module in the YOLOv8 Backbone, enhances the shallow feature map's ability to represent the structure of small infrared targets by adding a low-level feature detail enhancement module after the first and second C2f layers in the Backbone. This provides robust enhancement, especially when background noise varies greatly and target grayscale differences are small. The low-level feature detail enhancement module is as follows: Figure 2 As shown.
[0073] 3.1 Data Input
[0074] Receive shallow feature map of the input image after passing through the C2f layer :
[0075]
[0076] in: This is the shallow feature map after the C2f layer, with dimension 1. .
[0077] 3.2 Generating the perturbation kernel
[0078] For each channel, a depthwise separable convolution kernel with randomly sampled parameters using a Laplace distribution is generated:
[0079]
[0080] in, : Number of channels in the feature map. : Size of the convolution kernel. Location parameters of the Laplace distribution. The value is 0, which is the scale parameter of the Laplace distribution. The value is set to 0.1 to control the noise intensity. The Laplace distribution has a sharper peak, so its randomly sampled parameter values are more concentrated, which can avoid the instability of training caused by excessive differences in parameters between groups.
[0081] 3.3 Perform random perturbation convolution
[0082] right Apply channel-independent group convolutional perturbations to obtain perturbation feature maps. :
[0083]
[0084] in, The perturbated feature map, with dimensions and The same indicates that the new feature map is generated after convolution. : Two-dimensional convolution operation, using group convolution, with each channel convolved independently.
[0085] 3.4 Frequency Domain Amplitude and Phase Extraction
[0086] right and Perform a two-dimensional Fourier transform:
[0087]
[0088] in, and : respectively and The Fourier transform result represents the representation of the feature map in the frequency domain. This represents the Fast Fourier Transform.
[0089] After extracting the perturbation amplitude and phase :
[0090]
[0091] in, : After Fourier transform The amplitude represents the spectral amplitude after the disturbance. : Fourier transform Phase refers to the phase information of the original feature.
[0092] 3.5 Weighted Restructuring
[0093] In the frequency domain, the "phase" part of an image determines its structure (edges and shape). This invention preserves the original phase after perturbation to keep the target structure unchanged.
[0094] Define learnable parameters The amplitude after Fourier transform perturbation and the phase information of the original features after Fourier transform are weighted and reconstructed, and the spectrum is reconstructed using the following formula:
[0095] Symbol explanation: : The reconstructed frequency domain characteristics are obtained by reconstructing the spectrum through amplitude and original phase. : A complex exponential function representing phase information, used to recover phase information. Learnable parameters. Through... Weighting controls the influence of the perturbed amplitude on the reconstruction process, allowing the model to flexibly balance the diversity introduced by the perturbation with the importance of the original features. The reconstructed perturbation feature map is then obtained through inverse Fourier transform.
[0096]
[0097] in The reconstructed feature map has a dimension of 1. The spatial domain is restored through inverse Fourier transform. This represents the inverse Fourier transform.
[0098] Enhanced feature map output This method preserves details and edge information of small drone targets and improves their robustness to complex backgrounds, serving as input features for subsequent networks.
[0099] Step 4: Neck section processing.
[0100] The Neck part of YOLOv8 employs an optimized feature pyramid structure. It uses a top-down feature pyramid path and a bottom-up path aggregation network to upsample and downsample the feature maps of different scales output by the Backbone and fuse them with the C2f layer to output a further fused high-level feature map.
[0101] However, since the features of small target regions in infrared UAV images are easily diluted and weakened during multi-scale fusion and downsampling, it is necessary to specifically enhance the semantic structure consistency of the high-level fusion feature map in this area.
[0102] Step 5: High-level feature semantic preservation module.
[0103] The high-level feature semantic preservation module is embedded as a plug-in enhancement module after the first and second C2f layers of the YOLOv8 Neck network. It introduces attention mechanisms and contextual self-correction on high-level features. Under weak infrared texture conditions, this module enhances the semantic continuity and contextual expressiveness of features, and suppresses background noise and long-range inference interference on the detection of infrared small target UAV features. The high-level feature semantic preservation module is as follows: Figure 3 As shown
[0104] 5.1 Low-pass branch downsampling
[0105] For the input feature map Perform average pooling with a step size of 2 to obtain low-pass downsampling features:
[0106]
[0107] in, : 2×2 average pooling with a step size of 2; The low-pass branch output characteristics and spatial size become... .
[0108] 5.2 Multi-directional gradient branching
[0109] Let the three sets of directional gradient convolution kernels be... (Corresponding to gradient filter kernels in the horizontal / vertical / diagonal directions respectively), perform depthwise convolution on each channel, and use stride. Perform downsampling:
[0110]
[0111]
[0112]
[0113] in:
[0114]
[0115] To obtain a stable directional structural strength response, the gradient magnitude is calculated, ensuring it is non-negative and numerically stable.
[0116]
[0117]
[0118]
[0119] in, Features of the gradient magnitude in the horizontal direction; The feature is the gradient magnitude in the vertical direction; The gradient magnitude is a characteristic of the diagonal direction; The characteristics are those of the gradient response in the horizontal direction; This refers to the gradient response characteristics in the vertical direction; The gradient response characteristics are diagonal. The horizontal gradient filter convolution kernel; The vertical gradient filtering convolution kernel; Diagonal gradient filtering convolution kernel; This represents depthwise separable convolution; This serves as the input feature map for the high-level feature semantic preservation module. Step size; This represents element-wise multiplication; It is a very small positive number (in this embodiment, the value is taken as ). This avoids instability in the value when taking the square root.
[0120] 5.3 Feature splicing and channel compression
[0121] The low-pass filter and the three-directional amplitude are spliced together in the channel dimension:
[0122]
[0123] Then compressed back by 1×1 convolution Number of channels:
[0124]
[0125] in, : Splicing along the channel; : 1×1 convolution, used for channel compression and fusion; : The directional structural features after splicing; : Directional enhancement features after compression.
[0126] 5.4 Residual Downsampling Fusion
[0127] To supplement the semantic representation of the original features and achieve directional enhancement features with the compressed version. Size alignment, relative to the original input Perform convolutional downsampling to form residual branches:
[0128]
[0129] Then with Addition and fusion:
[0130]
[0131] in, : 3×3 convolution with a stride of 2, used for downsampling; LeakyReLU activation function; Residual branch downsampling characteristics; The fused feature map serves as the input for column-direction attention.
[0132] 5.5 Column pooling generates column description vectors
[0133] Fusion features In the channel dimension With high dimension Calculate the mean, retaining only the width dimension. , thus obtaining the column description vector :
[0134]
[0135] in, : For input tensors, simultaneously in the channel dimension With high dimension Find the mean; : Column description vector, reflecting the overall response strength of each column; The shape is .
[0136] Column description vector Perform convolution, then pass it through a sigmoid activation function to obtain the column direction attention weights. :
[0137]
[0138] in, : 1×1 convolution; The sigmoid function maps the output to... ; Column attention: Each column corresponds to a weight coefficient, giving higher weights to columns that are more likely to contain information related to structural semantics.
[0139] 5.6 Column Weight Broadcasting and Weighting
[0140] Column weights In high dimensions With channel dimension Broadcasting on, with Performing element-wise multiplication yields the column-direction consistency enhancement feature. :
[0141]
[0142] in, Element-wise multiplication; Broadcast rules: The size is During multiplication, the copy is expanded to... ; : Weighted feature map.
[0143] 5.7 Contextual Semantic Self-calibration
[0144] This step establishes a longer-range context compensation through a process of "pooling compression → convolution transformation → upsampling backfill".
[0145] Context compression:
[0146]
[0147] Contextual convolution transformation:
[0148]
[0149] Upsampling and stacking residuals:
[0150]
[0151]
[0152] in, Features after context compression; Features after contextual convolution transformation; Upsampling operators (e.g., bilinear interpolation); Backfill to Contextual compensation features; : Feature map after semantic self-calibration.
[0153] 5.8 Upsampling recovery output
[0154] The self-calibrated features are restored to the input resolution, and the output module results are then processed using a 1×1 convolution:
[0155] (1) Upsampling recovery size:
[0156]
[0157] (2) Output mapping:
[0158]
[0159] in, Restore to Feature map; The module outputs a feature map with the same size as the input map.
[0160] Output The input model is then further processed in subsequent modules, where it is fused and stitched with features from other scales, and finally fed into the YOLOv8 Head to complete bounding box regression and classification prediction.
[0161] Step 6: Head Part Detection Output and Loss Calculation
[0162] The multi-scale features fused from the Neck portion are fed into the Head portion, outputting a prediction set. It includes the class probability and bounding box regression value for each candidate box, and also calculates the training loss. .
[0163] Step 7: Parameter Update and Model Saving
[0164] Regarding the loss Perform backpropagation to update the parameters of the Backbone, Neck, and Head components, and iteratively train the model until it converges. Save the final model weights after training. .
[0165] Step 8: Model Deployment and Inference
[0166] Configure the inference environment on the target hardware platform and load the trained weights. The acquired infrared images The input is processed using a preprocessing procedure consistent with training to obtain the inference input tensor. The data is then fed into the model and passed through the Backbone, Neck, and Head parts for forward inference, outputting candidate boxes and class probabilities. Finally, the output candidate set is subjected to confidence thresholding and non-maximum suppression post-processing to output the final detection results. It includes the target bounding box coordinates, category, and score.
[0167] Implementation of an infrared image-based UAV target detection device based on complementary design:
[0168] A schematic diagram of the architecture of the infrared image UAV target detection device based on complementary design in this embodiment is shown below. Figure 5As shown, the system includes a memory, a processor, a system bus, and a computer program stored in the memory. The processor and memory communicate and exchange data via the system bus. The processor executes the computer program to implement the steps of the complementary design-based infrared image UAV target detection method of the present invention. The specific complementary design-based infrared image UAV target detection method has been described in sufficient detail in the above-described embodiments and will not be repeated here. The processor can be a microprocessor (MCU) or other processing device; the memory can be any type of memory that stores information using electrical energy, such as non-volatile storage media (including computer programs, databases), or other types of memory.
[0169] The infrared image drone target detection device based on complementary design in this embodiment can be deployed on a drone platform or carrier to achieve visual detection of drones during takeoff or recovery; it can also be deployed in places where drone monitoring is required, such as in urban security, where drone detection helps to promptly detect and stop drones from illegally filming or delivering contraband; in security work for important events or venues, drone detection can enhance the monitoring and defense capabilities of drones and ensure the safety of the event.
[0170] Implementation of computer-readable storage media:
[0171] A computer-readable storage medium storing a computer program internally, the computer program being executed by a processor to implement the steps of the infrared image-based UAV target detection method based on complementary design as described above. The specific infrared image-based UAV target detection method based on complementary design has been described in sufficient detail in the above-described embodiments and will not be repeated here.
[0172] The computer-readable storage medium can be the memory in the infrared image UAV target detection device based on complementary design, which can be used directly as a step in the infrared image UAV target detection method based on complementary design; or it can be used only as a storage medium for the computer program, such as a hard disk, which can transfer the computer program to another storage medium to implement the steps of the infrared image UAV target detection method based on complementary design.
[0173] Specifically, computer-readable storage media can be volatile memory or non-volatile memory, or both, depending on the application scenario. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. For example, Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), or Direct Rambus RAM (DRRAM).
[0174] This invention has the following characteristics:
[0175] This invention addresses the core challenges of small target detection in infrared images by proposing two complementary technical modules, applied to different layers of the neural network, to achieve stronger preservation of small target feature structures and robust detection performance: 1. The low-level module enhances the representation of small targets on infrared UAVs by adding random perturbation convolution and frequency domain reconstruction mechanisms. Under conditions of complex background noise and low contrast of small targets, this module effectively preserves the edge structure information of small targets by randomly generating depthwise separable convolution kernels and performing group convolution perturbations, combined with weighted reconstruction of amplitude and phase information from Fourier transform, while improving robustness to complex backgrounds. This design allows the model to balance the feature diversity brought by perturbation with the importance of the original features, providing a more robust feature representation for subsequent detection tasks. 2. The high-level feature semantic preservation module explicitly extracts and enhances the edge and structural information of small targets on high-level features through a "low-pass branch + multi-directional gradient filter group" approach. Based on residual semantic compensation, column-direction attention is introduced to weight the features column by column, enhancing the structural continuity of small target features in infrared scenes and suppressing background interference. At the same time, combined with the context self-calibration branch, a semantic compensation with a larger receptive field is established, making it less likely for small target features under weak texture and low contrast conditions to be submerged or broken during multi-scale fusion, thereby improving the positioning accuracy and robustness of infrared UAV small target detection.
[0176] This invention is specifically designed for infrared images and has been structurally adapted to address the following characteristics: target textures are blurred, making direct image detail enhancement difficult; small targets have small areas and low contrast, requiring preservation of their structural integrity in the feature space; and the target's low distinguishability from the background easily leads to model overfitting. Therefore, this invention combines a low-level feature detail enhancement module with a high-level feature semantic preservation module, leveraging the complementary functions of the modules to improve overall detection robustness and accuracy. By separately enhancing low-level detail features and high-level semantic features, this invention achieves a modeling mechanism from shallow to deep layers, enabling clearer and more stable detection of small targets in infrared images. This invention is specifically adapted to the characteristics of infrared images, with optimizations in module structure design, training perturbation strategies, and feature extraction paths for small target detection in grayscale infrared images, making it suitable for scenarios such as nighttime surveillance, infrared inspection, and long-range remote sensing drone monitoring.
[0177] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still make modifications to the technical solutions described in the foregoing embodiments without creative effort, or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for detecting unmanned aerial vehicle targets using infrared images based on complementary design, characterized in that, The method includes: S1. Acquire the infrared image to be detected; S2. Input the infrared image into a pre-trained infrared image drone detection model to obtain the drone category and location information in the infrared image; The infrared image drone detection model is an improvement on the YOLOv8 network. The improvement includes adding low-level feature detail enhancement modules between the first C2f layer and the third Conv layer and between the second C2f layer and the fourth Conv layer of the backbone network, respectively. The low-level feature detail enhancement module includes: enhancing the feature map output by the C2f layer. Apply random perturbation to obtain perturbation feature map In the frequency domain, based on the feature map Corresponding phase and perturbation feature maps The corresponding amplitude is used for spectrum reconstruction; the reconstructed frequency domain features are then restored to the feature map. The spatial domain is used to obtain the feature map output by the low-level feature detail enhancement module.
2. The infrared image UAV target detection method based on complementary design according to claim 1, characterized in that, The improvement also includes adding high-level feature semantic preservation modules between the first C2f layer and the second upsampling layer of the neck network and between the second C2f layer and the first Conv layer, respectively. The high-level feature semantic preservation module includes: processing the input feature map Pooling is performed to obtain low-pass downsampled features; depthwise separable convolutions are used to extract the input feature maps. Gradient response characteristics in each direction, and corresponding gradient magnitude characteristics calculated based on gradient response characteristics; The low-pass downsampling features are concatenated with the gradient magnitude features along the channel dimension, and the number of channels in the concatenated feature map is compressed to the input feature map. The number of channels is used to obtain directional enhancement features; Input feature map After aligning with the size of the directional enhancement feature, feature fusion is performed with the directional enhancement feature to obtain the fused feature; After the fused features are sequentially processed by column-direction attention and contextual semantic self-calibration, the feature size is restored to the input feature map by upsampling. The size is determined to obtain the feature map output by the high-level feature semantic preservation module.
3. The infrared image UAV target detection method based on complementary design according to claim 1, characterized in that, The spectrum reconstruction is a weighted spectrum reconstruction as shown below: in, For reconstructed frequency domain features; These are learnable weight parameters; Perturbation feature map The corresponding amplitude; For feature map Corresponding phase; For complex number operators.
4. The infrared image UAV target detection method based on complementary design according to claim 1, characterized in that, Perturbation feature map The feature map is obtained as follows: Each channel is independently convolved to obtain a perturbation feature map. The convolution kernel for the two-dimensional convolution operation is a randomly generated depth-separable convolution kernel that is randomly sampled using a Laplace distribution.
5. The infrared image UAV target detection method based on complementary design according to claim 2, characterized in that, The gradient magnitude feature is obtained in the following way: in, Features of the gradient magnitude in the horizontal direction; The feature is the gradient magnitude in the vertical direction; The gradient magnitude is a characteristic of the diagonal direction; The characteristics are those of the gradient response in the horizontal direction; This refers to the gradient response characteristics in the vertical direction; The gradient response characteristics are diagonal. This represents element-wise multiplication; It is a positive number whose value is unstable when taking the square root.
6. The infrared image UAV target detection method based on complementary design according to claim 2, characterized in that, The directional enhancement feature is based on the stitched feature map. The result is obtained after convolution.
7. The infrared image UAV target detection method based on complementary design according to claim 2, characterized in that, The column direction attention is as follows: the mean of the fused feature is calculated simultaneously in both the height and channel dimensions to generate a column description vector; the column description vector is then sequentially processed... After convolution and activation function processing, column-direction attention weights are obtained; the attention weights are broadcast in the height and channel dimensions and then multiplied element-wise with the fused features to obtain the output features of column-direction attention.
8. The infrared image UAV target detection method based on complementary design according to claim 2 or 7, characterized in that, The context semantic self-calibration is as follows: the output features of column direction attention are sequentially processed by downsampling compression, convolution transformation and upsampling feature recovery to obtain context compensation features; then, the context compensation features are superimposed with the output features of column direction attention to obtain the output features of context semantic self-calibration.
9. A target detection device for unmanned aerial vehicles based on infrared imaging with complementary design, comprising a processor, characterized in that, The processor is used to execute a computer program to implement the steps of the infrared image UAV target detection method based on complementary design as described in any one of claims 1 to 8.
10. A computer-readable storage medium, wherein a computer program is stored internally, characterized in that, The computer program is executed by a processor to implement the steps of the infrared image UAV target detection method based on complementary design as described in any one of claims 1 to 8.