Unmanned aerial vehicle infrared remote sensing image weak and small target detection method and device

By introducing the WIoU loss function and non-maximum suppression of the in-scale feature efficient interaction and fusion module, the WIoU loss function and non-maximum suppression of the distance attention mechanism, the speed and accuracy of weak target detection in the infrared remote sensing images of the drone is solved, and a fast and lightweight detection effect is achieved, suitable for environmental monitoring and military reconnaissance.

CN120580591AActive Publication Date: 2025-09-02GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202510755027.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-02
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing drone infrared remote sensing image object detectors have shortcomings in speed, lightweight and accuracy, and cannot effectively detect small and low contrast weak targets, and have a high missed detection rate.

Method used

Introduce the in-scale feature efficient interaction and fusion module, WIoU loss function based on distance attention mechanism and non-maximum suppression (soft-NMS), and design a lightweight network to improve detection accuracy and robustness.

Benefits of technology

It realizes fast and lightweight infrared remote sensing image weak target detection, reducing missed detection, and is suitable for real-time applications, multi-scene environmental monitoring and military reconnaissance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580591A_ABST
    Figure CN120580591A_ABST
Patent Text Reader

Abstract

The invention relates to an unmanned aerial vehicle infrared remote sensing image weak and small target detection method and device. The method comprises the steps of obtaining a to-be-detected infrared remote sensing image; the to-be-detected infrared remote sensing image is input into a preset target detection model, a weak and small target detection result is output, the target detection model is obtained by inputting a training set and adopting a WIoU loss function based on a distance attention mechanism for training, the training set comprises the infrared remote sensing image and a corresponding target label, and the target label is a target label corresponding to the infrared remote sensing image. The target detection model is constructed by introducing an intra-scale feature efficient interaction mechanism and a non-maximum suppression mechanism into a yolk-based target detector. According to the method, the feature information in the infrared remote sensing image can be effectively utilized, and the detection accuracy and robustness of the weak and small target are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of infrared small target detection, and in particular to a method and device for detecting small targets in infrared remote sensing images of unmanned aerial vehicles. Background Art

[0002] As an emerging, safe, and efficient detection method, drone infrared remote sensing imagery target detection technology has the potential for widespread application in the field of infrared small target detection. Drone infrared remote sensing systems can operate with high precision around the clock, making it possible to detect small infrared targets even in complex backgrounds. Drone infrared remote sensing imagery target detection technology offers wide coverage and high mobility, making it suitable for disaster monitoring, military reconnaissance, and other applications.

[0003] In recent years, the rapid development of deep learning technology has provided powerful tools for image recognition and detection. However, traditional object detectors have limitations for detecting objects in drone infrared remote sensing images. Existing object detectors are not ideal in terms of speed and lightweightness, failing to meet the requirements of real-time applications or applications on edge devices such as mobile devices. Furthermore, the small and faint objects in drone infrared remote sensing images are often tiny and low-contrast, making traditional detection methods challenging in terms of accuracy and robustness. Therefore, developing a faster, lighter, and more accurate object detection algorithm is of great significance for object detection in drone infrared remote sensing images. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and device for detecting small targets in unmanned aerial vehicle infrared remote sensing images to address the problems of low detection accuracy, high small target missed detection rate, and high computational complexity in existing algorithms. The method introduces an efficient intra-scale feature interaction and fusion module, a WIoU loss function based on a distance attention mechanism, and non-maximum suppression (soft-NMS). At the same time, a lightweight network is designed to reduce the complexity of the network, which can effectively utilize the feature information in infrared remote sensing images and improve the detection accuracy and robustness of small targets.

[0005] To achieve the above objectives, the present invention provides a method for detecting small targets in infrared remote sensing images of unmanned aerial vehicles, comprising:

[0006] Acquire the infrared remote sensing image to be measured;

[0007] The infrared remote sensing image to be tested is input into a preset target detection model, and a weak target detection result is output, wherein the target detection model is obtained by inputting a training set and training with a WIoU loss function based on a distance attention mechanism, the training set includes infrared remote sensing images and corresponding target annotations, and the target detection model is constructed by introducing an efficient intra-scale feature interaction mechanism and a non-maximum suppression mechanism in a Yolo-based target detector.

[0008] Optionally, obtaining the training set includes:

[0009] Acquire a number of original infrared remote sensing images containing faint targets, and convert the original infrared remote sensing images from a single-channel to a pseudo three-channel RGB format;

[0010] The categories of small targets and the coordinates of the upper left corner and lower right corner of the small targets are marked in the infrared remote sensing image after format conversion to obtain the final infrared remote sensing image.

[0011] Optionally, the target detection model includes:

[0012] The backbone network is used to extract, interact and fuse features of the input image and output several first feature maps of different scales;

[0013] The neck network is used to perform feature splicing and extraction on the first feature maps of several different scales and output the second feature maps of several different scales;

[0014] The head network is used to detect small targets on several second feature maps of different scales and output detection boxes.

[0015] Optionally, the input image in the backbone network passes through a convolutional layer, a feature extraction module, a fast spatial pyramid pooling layer, and a cross-stage efficient attention mechanism to output three feature maps of different scales I b1 , I b2 and I b3 , Feature Map I b1 , I b2 After the feature map is spliced ​​to connect the feature extraction module and the feature transfer of the neck network, the feature map I b3 After feature map splicing, the cross-stage efficient attention mechanism and the feature transfer of the neck network are connected, and the feature map I b3 After the self-attention operation and the intra-scale feature interaction are performed by the efficient intra-scale feature interaction and fusion module, the feature maps in the neck network are added and spliced. The neck network outputs three feature maps of different scales I n1 , I n2 and I n3 , respectively input into the detection head in the head network, and output the weak target detection box.

[0016] Optionally, the intra-scale feature efficient interaction and fusion module adds two-dimensional sine-cosine position encoding to the input features, and then outputs the interactively fused features through a multi-head attention mechanism, residual connection and layer normalization, and a feed-forward fully connected network.

[0017] Optionally, after obtaining several small and weak target detection frames output by the head network, splicing and dimensionality transformation operations are performed in combination with the non-maximum suppression mechanism to output a final small and weak target detection frame, wherein the non-maximum suppression mechanism is:

[0018]

[0019] Among them, s i is the bounding box score, b i is the initial bounding box, is the bounding box with the largest score, D is the final result box set, and σ is the standard deviation of the Gaussian function.

[0020] Optionally, the WIoU loss function based on the distance attention mechanism is:

[0021] L WIoUv1 =R WIoU ×L IoU ;

[0022]

[0023] Among them, L WIoUv1 is the WIoU v1 function, L IoU is the IOU loss function, R WIoU is the penalty term of WIOU, x and y are the coordinates of the upper left corner of the anchor box, x gt and y gt is the coordinate of the upper left corner of the target box, W g and H g Indicates the width and height of the minimum bounding box. The superscript * indicates that W g and H g Detach from the computation graph.

[0024] On the other hand, the present invention also provides a device for detecting small targets in infrared remote sensing images of unmanned aerial vehicles, comprising:

[0025] Image acquisition equipment, used to obtain infrared remote sensing images to be measured;

[0026] The image processing device is used to input the infrared remote sensing image to be tested into a preset target detection model and output a weak target detection result, wherein the target detection model is obtained by inputting a training set and training with a WIoU loss function based on a distance attention mechanism. The training set includes infrared remote sensing images and corresponding target annotations. The target detection model is constructed by introducing an efficient intra-scale feature interaction mechanism and a non-maximum suppression mechanism in a Yolo-based target detector.

[0027] The beneficial effects of the present invention are:

[0028] This invention effectively captures the detailed features of small and weak targets in infrared remote sensing images by introducing an efficient intra-scale feature interaction and fusion module, enabling cross-channel and cross-spatial information fusion between feature maps. Furthermore, the use of a distance-attention-based WIoU loss function and non-maximum suppression (Soft-NMS) improves detection accuracy and reduces missed detections. This invention boasts the advantages of speed, lightweight design, and accuracy, making it suitable for real-time applications and various scenarios in infrared image small and weak target detection. It has broad application prospects in environmental monitoring, military reconnaissance, and other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0030] Figure 1 This is a schematic diagram of the structure of a device for detecting small targets in infrared remote sensing images of a drone according to an embodiment of the present invention, wherein 101 is a computer, 102 is an infrared remote sensing imaging device of a drone, 103 is a transceiver, and 104 is an inspected area.

[0031] Figure 2 Schematic diagram of the target detection model structure according to an embodiment of the present invention;

[0032] Figure 3 Schematic diagram of the structure of the efficient intra-scale feature interaction and fusion module in an embodiment of the present invention. DETAILED DESCRIPTION

[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0034] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] This embodiment provides a device for detecting small targets in infrared remote sensing images of a drone, comprising:

[0036] Image acquisition equipment, used to obtain infrared remote sensing images to be measured;

[0037] The image processing device is used to input the infrared remote sensing image to be tested into a preset target detection model and output a weak target detection result, wherein the target detection model is obtained by inputting a training set and training with a WIoU loss function based on a distance attention mechanism. The training set includes infrared remote sensing images and corresponding target annotations. The target detection model is constructed by introducing an efficient intra-scale feature interaction mechanism and a non-maximum suppression mechanism in a Yolo-based target detector.

[0038] Specifically, such as Figure 1 As shown, in this embodiment, the image acquisition device uses a drone infrared remote sensing imaging device 102 to collect infrared remote sensing images of the inspected area 104, and sends them to the image processing device, namely the computer 101, through the transceiver device 103 to realize weak target detection in the infrared remote sensing image.

[0039] Based on the above device, this embodiment also provides a method for detecting small targets in infrared remote sensing images of a drone, including:

[0040] Acquire the infrared remote sensing image to be measured;

[0041] The infrared remote sensing image to be tested is input into a preset target detection model, and a weak target detection result is output, wherein the target detection model is obtained by inputting a training set and training with a WIoU loss function based on a distance attention mechanism, the training set includes infrared remote sensing images and corresponding target annotations, and the target detection model is constructed by introducing an efficient intra-scale feature interaction mechanism and a non-maximum suppression mechanism in a Yolo-based target detector.

[0042] Specifically, this embodiment introduces an efficient intra-scale feature interaction and fusion module to effectively capture the detailed features of small targets in infrared remote sensing images and realize cross-channel and cross-space information fusion between feature maps; at the same time, it uses the WIoU loss function based on the distance attention mechanism and non-maximum suppression soft-NMS to improve detection accuracy and reduce missed detections.

[0043] Furthermore, obtaining the training set includes:

[0044] Acquire a number of original infrared remote sensing images containing faint targets, and convert the original infrared remote sensing images from a single-channel to a pseudo three-channel RGB format;

[0045] The categories of small targets and the coordinates of the upper left corner and lower right corner of the small targets are marked in the infrared remote sensing image after format conversion to obtain the final infrared remote sensing image.

[0046] Specifically, in this embodiment, a plurality of infrared remote sensing images containing faint targets are captured by the infrared remote sensing imaging device 102 of the drone, and the images are input into the computer 101 through the transceiver device 103 to convert the single-channel images into a pseudo three-channel RGB format, thereby constructing an infrared remote sensing image dataset I. h =[I h1 , I h2 ,...I hK ], where dataset I h The total number of elements in is K, and the image size is n ch ×h×w,n ch represents the image channel, h represents the image height, and w represents the image width. Image annotation uses the open-source tool labelImg. The annotations include the infrared target category and the coordinates of the target's upper left and lower right corners. The annotated information is stored in a txt file.

[0047] Furthermore, the target detection model includes:

[0048] The backbone network is used to extract, interact and fuse features of the input image and output several first feature maps of different scales;

[0049] The neck network is used to perform feature splicing and extraction on the first feature maps of several different scales and output the second feature maps of several different scales;

[0050] The head network is used to detect small targets on several second feature maps of different scales and output detection boxes.

[0051] The specific structure of the target detection model is as follows Figure 2 As shown, using size n ch The RGB image of size h×w is used as input to the backbone network of the target detection model, which passes through the convolution layer, feature extraction module, fast spatial pyramid pooling layer, cross-stage efficient attention mechanism, and outputs three feature maps of different scales I b1 , I b2 and I b3 , the scales are n c1 ×h / 8×w / 8,n c2 ×h / 16×w / 16 and n c3 ×h / 32×w / 32. Among them, the feature map I b1 , I b2 The feature extraction module in the backbone network is connected with the feature transfer of the neck network through feature map splicing. b3 The cross-stage efficient attention mechanism in the backbone network is connected with the feature transfer of the neck network through feature map splicing. Then, the feature map I b3The backbone network's efficient intra-scale feature interaction and fusion module performs self-attention and intra-scale feature interaction to achieve information fusion. The output and input feature scales of the efficient intra-scale feature interaction and fusion module are kept consistent.

[0052] Output three feature maps of different scales I in the neck network n1 , I n2 and I n3 , the scales are n c1 ×h / 8×w / 8,n c2 ×h / 16×w / 16 and n c3 ×h / 32×w / 32, to achieve multi-scale feature fusion, thereby improving the performance of weak target detection. Then the feature map I n1 , I n2 and I n3 The detection head in the head network is input separately and outputs three images containing weak target detection boxes.

[0053] Among them, the efficient intra-scale feature interaction and fusion module adds two-dimensional sine-cosine position encoding to the input features, and then outputs the interactively fused features through a multi-head attention mechanism, residual connection and layer normalization, and a feedforward fully connected network.

[0054] Specifically, such as Figure 3 As shown in the figure, during the forward propagation process, the shape of the input feature is [B, C, H, W] (batch size B, number of channels C, height H, width W), which is flattened and transposed to the format of [B, H×W, C] before processing. Then, 2D Sin-Cos Positional Embedding is added and embedded into the feature to provide position information. Finally, the output feature is transposed back to the original format [B, C, H, W].

[0055] The multi-head attention mechanism is an improvement on the self-attention mechanism. When generating q, k, v (query, index, content), q, k, v are split into num_heads (attention heads) parts, and self-attention operation is performed on each part. Finally, the results are spliced ​​together.

[0056] Residual connection and layer normalization consists of two parts: Add and Norm. The calculation process of Add&Norm layer can be expressed as follows:

[0057] Add&Norm(X)=LayerNorm(X+MultiHeadAttention(X));

[0058] Here, Add refers to X + MultiHeadAttention(X), which performs a residual connection; Norm refers to LayerNorm(), also known as LayerNormalization, which performs layer normalization. X refers to the feature map input vector.

[0059] The feedforward fully connected network, abbreviated as FFN, is essentially a two-layer fully connected layer. The activation function of the first layer is Relu, and the second layer does not use an activation function. The calculation process can be expressed in mathematical formulas as follows:

[0060] FFN(X)=max(0,XW1+b1)W2+b2;

[0061] W1 and W2 are weight matrices, mapping the input dimension to the hidden dimension and back to the output dimension, respectively; b1 and b2 are bias terms. max(0,x) converts negative numbers to zero, enhancing nonlinearity and reflecting the ReLU activation function; X refers to the feature map input vector.

[0062] Furthermore, the target detection model is trained:

[0063] First, freeze the weights of the feature extraction module in the backbone network and train for several epochs. Then, unfreeze all weights and train for several epochs. Set the network training parameters: learning rate lr, batch size, training set / validation set split, optimizer, and training period.

[0064] The loss function used is the WIoU loss function based on the distance attention mechanism, which effectively eliminates the influence of hindering convergence and does not introduce new indicators (such as aspect ratio). The calculation process can be expressed as follows:

[0065] L WIoUv1 =R WIoU ×L IoU ;

[0066] Among them, L WIoUv1 is the WIoU v1 function, L IoU is the IOU loss function, R WIoU is the penalty term of WIOU, calculated as follows:

[0067]

[0068] Among them, x and y are the coordinates of the upper left corner of the anchor box, x gt and y gt is the coordinate of the upper left corner of the target box, W g and H g Indicates the width and height of the minimum bounding box; at the same time, in order to eliminate R WIoU The influence of hindering convergence, W g and Hg Detach from the computation graph (superscript * denotes this operation).

[0069] Further, target detection model testing and application:

[0070] Use the trained target detection model to make predictions, input the test image, and output the target box predicted by the infrared remote sensing image. t Input to the network, the image size is n ch ×h×w, after network inference, the output of the detection head is obtained. The output feature map scales are three feature maps of 80×80, 40×40 and 20×20. The classification and regression prediction results are extracted from the feature maps of different scales, and splicing and dimension transformation operations are performed. For ease of processing, the original channel dimension is permuted to the end, and the shapes of the category prediction branch and the bbox prediction branch are obtained as (b,8400,80) and (b,8400,4) respectively. Then soft-NMS is used to replace the traditional NMS for non-maximum suppression, and the confidence of the overlapping boxes is reduced by Gaussian weighting. The formula is:

[0071]

[0072] Among them, s i is the bounding box score, b i is the initial bounding box, is the bounding box with the largest score, D is used to store the final result box set, and σ is the standard deviation of the Gaussian function.

[0073] By converting the pixel coordinates of the target detection box to a normalized scale and visualizing the detection results in the image, the presence of the target is determined based on whether the detection box exists. If the input contains the detection box parameters (x, y, w, h), the normalization operation is performed to (X, Y, W, H), and a rectangular box is drawn, and the target is determined to be found; otherwise, the target is not found.

[0074] The small target detection device and apparatus for infrared remote sensing images of unmanned aerial vehicles, as well as the specific target detection model proposed in this embodiment, have the advantages of being fast, lightweight, and accurate. They are suitable for real-time applications and various scenarios of small target detection in infrared images, and have broad application prospects in environmental monitoring, military reconnaissance and other fields.

[0075] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A method for detecting small targets in infrared remote sensing images of unmanned aerial vehicles, characterized in that: include: Acquire the infrared remote sensing image to be measured; The infrared remote sensing image to be tested is input into a preset target detection model, and a weak target detection result is output, wherein the target detection model is obtained by inputting a training set and training with a WIoU loss function based on a distance attention mechanism, the training set includes infrared remote sensing images and corresponding target annotations, and the target detection model is constructed by introducing an efficient intra-scale feature interaction mechanism and a non-maximum suppression mechanism in a Yolo-based target detector.

2. The method for detecting small targets in infrared remote sensing images of unmanned aerial vehicles according to claim 1, characterized in that: Obtaining the training set includes: Acquire a number of original infrared remote sensing images containing faint targets, and convert the original infrared remote sensing images from a single-channel to a pseudo three-channel RGB format; The categories of small targets and the coordinates of the upper left corner and lower right corner of the small targets are marked in the infrared remote sensing image after format conversion to obtain the final infrared remote sensing image.

3. The method for detecting small targets in infrared remote sensing images of unmanned aerial vehicles according to claim 1, characterized in that: The target detection model includes: The backbone network is used to extract, interact and fuse features of the input image and output several first feature maps of different scales; The neck network is used to perform feature splicing and extraction on the first feature maps of several different scales and output the second feature maps of several different scales; The head network is used to detect small targets on several second feature maps of different scales and output detection boxes.

4. The method for detecting small targets in infrared remote sensing images of unmanned aerial vehicles according to claim 3, characterized in that: The input image in the backbone network passes through the convolution layer, feature extraction module, fast spatial pyramid pooling layer, and cross-stage efficient attention mechanism, and outputs three feature maps of different scales I b1 , I b2 and I b3 , Feature Map I b1 , I b2 After the feature map is spliced ​​to connect the feature extraction module and the feature transfer of the neck network, the feature map I b3 After feature map splicing, the cross-stage efficient attention mechanism and the feature transfer of the neck network are connected, and the feature map I b3 After the self-attention operation and the intra-scale feature interaction are performed by the efficient intra-scale feature interaction and fusion module, the feature maps in the neck network are added and spliced. The neck network outputs three feature maps of different scales I n1 , I n2 and I n3 , respectively input into the detection head in the head network, and output the weak target detection box.

5. The method for detecting small targets in infrared remote sensing images of unmanned aerial vehicles according to claim 4, characterized in that: The efficient intra-scale feature interaction and fusion module adds two-dimensional sine-cosine position encoding to the input features, and then outputs the interactively fused features through a multi-head attention mechanism, residual connection and layer normalization, and a feed-forward fully connected network.

6. The method for detecting small targets in infrared remote sensing images of unmanned aerial vehicles according to claim 4, characterized in that: After obtaining several small target detection frames output by the head network, the non-maximum suppression mechanism is combined with the splicing and dimension transformation operations to output the final small target detection frame, wherein the non-maximum suppression mechanism is: Among them, s i is the bounding box score, b i is the initial bounding box, is the bounding box with the largest score, D is the final result box set, and σ is the standard deviation of the Gaussian function.

7. The method for detecting small targets in infrared remote sensing images of unmanned aerial vehicles according to claim 1, characterized in that: The WIoU loss function based on the distance attention mechanism is: L WIoUv1 =R WIoU ×L IoU ; Among them, L WIoUv1 is the WIoU v1 function, L IoU is the IOU loss function, R WIoU is the penalty term of WIOU, x and y are the coordinates of the upper left corner of the anchor box, x gt and y gt is the coordinate of the upper left corner of the target box, W g and H g Indicates the width and height of the minimum bounding box. The superscript * indicates that W g and H g Detach from the computation graph.

8. A device for detecting small targets in infrared remote sensing images of unmanned aerial vehicles, characterized in that: include: Image acquisition equipment, used to obtain infrared remote sensing images to be measured; The image processing device is used to input the infrared remote sensing image to be tested into a preset target detection model and output a weak target detection result, wherein the target detection model is obtained by inputting a training set and training with a WIoU loss function based on a distance attention mechanism. The training set includes infrared remote sensing images and corresponding target annotations. The target detection model is constructed by introducing an efficient intra-scale feature interaction mechanism and a non-maximum suppression mechanism in a Yolo-based target detector.

Citation Information

Patent Citations

  • Infrared weak and small target detection method based on asymmetric attention feature fusion

    CN113591968A

  • Small target detection method for remote sensing image

    CN118115893A

  • Infrared weak and small target detection method based on double-branch attention mechanism

    CN119600291A

  • Crystalline silicon photovoltaic cell defect real-time detection method based on multi-scale feature fusion

    CN119741470A

  • Object detection in an image

    US20210303862A1

Cited By

  • Light-weight unmanned aerial vehicle infrared remote sensing target detection method and device based on knowledge distillation

    CN121280703A