A small target detection method for infrared images based on U-shaped network and patch attention for intelligent reconnaissance equipment

By combining the U-shaped network and the patch attention module, the infrared image small target detection method solves the problem that intelligent reconnaissance equipment has difficulty detecting small infrared targets in complex backgrounds, and achieves higher detection accuracy and robustness.

CN120279248BActive Publication Date: 2025-10-03GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510341242.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-10-03
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

Existing intelligent reconnaissance equipment has difficulty distinguishing targets in complex backgrounds when processing weak targets in infrared images. In addition, due to limited computing resources, it is difficult to deploy large-scale deep learning models, and the detection accuracy and robustness are insufficient.

Method used

A small target detection method for infrared images is adopted that combines the U-net and the patch attention module. By improving the UIU-Net model and adding the patch attention module to learn local knowledge and global information, the detection performance is optimized.

Benefits of technology

It improves the detection capability of intelligent reconnaissance equipment in complex environments, enhances the ability to identify small infrared targets, and improves detection accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279248B_ABST
    Figure CN120279248B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of target detection technology and proposes a method for detecting small targets in infrared images using intelligent reconnaissance equipment by combining a U-shaped network and patch attention. The method comprises the following steps: using an infrared camera within the intelligent reconnaissance equipment to capture and preprocess an infrared image of a target; inputting the preprocessed infrared image into an infrared image small target detection model based on UIU-Net for feature extraction, and decoding the extracted features to obtain a target detection result; wherein the attention module in the infrared image small target detection model includes a patch attention module. The present invention uses the patch attention module to improve the infrared image small target detection model based on UIU-Net, so that the model can not only collect global information but also learn knowledge about local patches, thereby rejecting interference from far-away targets and enhancing weak targets, further enhancing the network's ability to detect small infrared targets, and enabling the network to have better performance when processing complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology, and more specifically, to a method for detecting small targets in infrared images using a combination of a U-shaped network and patch attention for intelligent reconnaissance equipment. Background Art

[0002] Intelligent reconnaissance equipment is widely used in both civilian and military applications. It can be equipped with infrared sensors to identify and track targets. However, small, faint targets in infrared images are often lost in complex backgrounds, making them difficult to distinguish. Furthermore, small targets lack distinct texture and color features, making them challenging to detect using traditional image processing methods. Furthermore, the limited computing resources and storage space in intelligent reconnaissance equipment make it difficult to deploy large deep learning models. Furthermore, the complex and ever-changing operating environment of these equipment requires robust image processing methods. To address these issues, specialized infrared small target detection algorithms are needed that balance detection accuracy, speed, and resource requirements, while also improving algorithm robustness to adapt to complex environments.

[0003] Deep learning-based methods have become the most widely used technology for detecting small infrared objects. Currently, some propose modeling the problem as semantic segmentation rather than detection, which can help improve small object detection. Among them, UIU-Net (U-Net in U-Net) uses a resolution-preserving network to learn multi-scale features and an interactive cross-attention module to fuse features. However, its attention module suffers from shortcomings, such as a poor understanding of local semantics. Summary of the Invention

[0004] In order to overcome the defects of the above-mentioned prior art, the present invention provides a method for detecting small targets in infrared images by combining a U-shaped network and patch attention for intelligent reconnaissance equipment.

[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows:

[0006] A method for detecting small targets in infrared images using a U-shaped network and patch attention for intelligent reconnaissance equipment includes the following steps:

[0007] Use the infrared camera in the intelligent reconnaissance equipment to collect target infrared images and perform pre-processing;

[0008] The preprocessed infrared image is input into the infrared image small target detection model based on UIU-Net for feature extraction, and the extracted features are decoded to obtain the target detection result; wherein, the attention module in the infrared image small target detection model includes a patch attention module.

[0009] Furthermore, the present invention also proposes a small target detection system for infrared images using a U-shaped network and patch attention for intelligent reconnaissance equipment, which applies the small target detection method for infrared images proposed in the present invention. The system includes:

[0010] An image acquisition module, including an infrared camera within the intelligent reconnaissance equipment, is used to acquire infrared images of the target;

[0011] A preprocessing module, used for performing image preprocessing on the collected infrared images;

[0012] The target detection module is equipped with a UIU-Net-based infrared image small target detection model, which is used to extract features from the input preprocessed infrared image and decode the extracted features to obtain target detection results; wherein, the attention module in the infrared image small target detection model includes a patch attention module.

[0013] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0014] In this paper, the patch attention module is used to improve the infrared image small target detection model based on UIU-Net. The model can not only collect global information, but also learn the knowledge of local patches, thereby rejecting interference from far-away targets and enhancing weak targets, further strengthening the network's detection ability for infrared small targets, and making the network have better performance when processing complex scenes.

[0015] The present invention is applied to an intelligent pod to improve its adaptability to changing environments and to more easily detect weak targets in a blurred background in an infrared image. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 The figure is a flow chart of a method for detecting small targets in infrared images according to one embodiment of the present invention.

[0017] Figure 2 The figure is an architecture diagram of a small target detection model in infrared images according to one embodiment of the present invention.

[0018] Figure 3 FIG. 1 is an architecture diagram of an RSU unit according to an embodiment of the present invention.

[0019] Figure 4 FIG. 1 is a schematic diagram showing a skip connection according to an embodiment of the present invention.

[0020] Figure 5 This is an architectural diagram of a patch attention module according to one embodiment of the present invention.

[0021] Figure 6FIG1 is an architecture diagram of a system for detecting small targets in infrared images according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0023] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0024] It should be understood that although the terms "first," "second," "third," etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information, without departing from the scope of the present invention. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."

[0025] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. Example 1

[0026] This embodiment proposes a method for detecting small targets in infrared images. Figure 1 FIG. 1 is a flow chart of the infrared image small target detection method according to the present embodiment.

[0027] The infrared image small target detection method proposed in this embodiment includes the following steps:

[0028] S1. Use the infrared camera in the intelligent reconnaissance equipment to collect the target infrared image and perform pre-processing;

[0029] S2. Input the preprocessed infrared image into the infrared image weak target detection model based on UIU-Net for feature extraction, and decode the extracted features to obtain the target detection result;

[0030] Among them, the attention module in the infrared image small target detection model includes a patch attention module.

[0031] The infrared image small target detection model in this embodiment uses the UIU-Net network as its foundational architecture. UIU-Net is a deep learning framework that improves the detection accuracy of small objects in infrared images through a nested U-Net structure. Conventional UIU-Net networks typically use a resolution-maintaining network to learn multi-scale features and employ an interactive cross-attention module to fuse features. However, this cross-attention module lacks a good understanding of local semantics, resulting in low detection accuracy when applied to infrared image small target detection.

[0032] This embodiment uses a patch attention module to replace the interactive cross attention module in the original network. This module includes low-level feature branches to learn details, and high-level feature branches to learn semantic long-range dependencies to extract the advantages of multi-scale features. During use, the two features are calibrated and fused to learn dynamic weights. In addition, this module can be inserted into various networks and combined with loss functions, model compression and other technologies to optimize the detection performance and speed of small infrared targets, thereby enhancing the detection capabilities of intelligent reconnaissance equipment for weak infrared targets.

[0033] In this embodiment, the patch attention module is used to improve the infrared image small target detection model based on UIU-Net, so that the model can not only collect global information, but also learn the knowledge of local patches, thereby rejecting interference from far-away targets and enhancing weak targets, further strengthening the network's detection ability for small infrared targets, and making the network have better performance when processing complex scenes.

[0034] Applying the present invention to an intelligent pod can improve its adaptability to changing environments and make it easier to detect weak targets in a blurred background in an infrared image.

[0035] In an optional embodiment, in step S1, the step of preprocessing the acquired target infrared image includes:

[0036] Convert the acquired target infrared image into A tensor of size; where H and W are the height and width of the image respectively;

[0037] Normalize the pixel value range of the image tensor to between.

[0038] For example, the infrared camera used in this embodiment is a Tigris 640 medium-wave infrared camera with a resolution of 320×320, capable of providing thermal imaging information of the target. When preprocessing the captured infrared image, the infrared image is first converted into a tensor of size (320, 320). Then, the pixel values ​​are divided by 255, subtracted by 0.5, and finally multiplied by 2 to normalize the pixel values ​​to the range [-1, 1].

[0039] Further optionally, the data set is divided into a training set, a validation set and a test set at approximately 50%, 20% and 30% respectively, for pre-training, validation and testing of the infrared image small target detection model.

[0040] In an optional embodiment, the infrared image small target detection model includes an encoding network and a decoding network; wherein, the encoding network includes at least 6 first convolutional layers connected in sequence, and each first convolutional layer includes an RSU unit composed of several layers of U-Net; the decoding network includes at least 5 second convolutional layers, each second convolutional layer has the same composition as the first convolutional layer of its corresponding level, and the second convolutional layer and the first convolutional layer of the corresponding level are jump connected.

[0041] For example, Figure 2 FIG. 1 is a diagram showing the architecture of the infrared image small target detection model of this embodiment.

[0042] The RSU unit in this embodiment is a network module that combines a U-shaped structure and a residual connection, and is intended to extract and fuse multi-scale features from an input feature map.

[0043] The RSU unit includes several CBR modules. Specifically, the CBR module includes a 3×3 convolution layer Conv, a batch normalization layer BN and a ReLU activation function layer connected in sequence, which are used to extract features at different scales of the input image. Figure 3 , which is a diagram of the architecture of the RSU unit of this embodiment.

[0044] Furthermore, in an optional embodiment, the encoding network includes six first convolutional layers stage1 to stage6 connected in sequence, wherein:

[0045] The first convolutional layer stage1 consists of 7 layers of U-Net, and the output of the 7th layer of U-Net is connected to an extended convolution block with an expansion rate of 2;

[0046] The structures of the first convolutional layers stage2, stage3, and stage4 are the same as that of the first convolutional layer stage1, and the number of U-Net layers inside them decreases layer by layer compared to the first convolutional layer stage1.

[0047] The first convolutional layers stage5 and stage6 include 4 layers of U-Net, and the output of each layer of U-Net is connected to an extended convolution block with an expansion rate greater than 1; the output of the first convolutional layer stage5 is connected to a maximum pooling layer; the output of the first convolutional layer stage6 is input to the decoding network after passing through the upsampling layer.

[0048] Furthermore, in an optional embodiment, the decoding network includes five second convolutional layers stage1d to stage5d connected in sequence, which are respectively at the same level and have the same structure as the first convolutional layers stage1 to stage5; wherein the input of the second convolutional layer is the upsampling result of the output of the previous layer. The feature map output by the first convolutional layer at the same level Input the output of the patch attention module .

[0049] In this embodiment, the output of the decoding network Expressed as:

[0050]

[0051] in, Indicates the first k The output features of the layer, K is the number of layers in the entire network; Represents the intermediate feature map of the input; Represents the RSU module of the kth layer.

[0052] For example, Figure 4 FIG. 1 is a schematic diagram of a jump connection according to the present embodiment.

[0053] As an example, in the first convolutional layer stage1, after one CBR operation, the feature map of the same size is output. After 5 CBR and downsampling operations, 1 CBR operation, and 1 extended CBR operation with an expansion rate of 2, the deep feature map output by stage1 is D 1 is input into the corresponding layer of the decoding network.

[0054] Furthermore, in the second convolutional layer stage1d of the corresponding level in the decoding network, its input includes the deep feature map output by stage1 D 1, and the previous layer output of stage1d, that is, the feature map output by stage6 D 6. The upsampling result after upsampling. Stage 1d concatenates the two feature maps to implement skip connection, then performs deconvolution upsampling operation, gradually obtains the original size feature map and outputs it to stage 2.

[0055] Furthermore, the patch attention module includes a first branch and a second branch of a jump connection; wherein the input of the first branch is the upsampling result , which includes two PW convolutional layers and one sigmoid layer connected in sequence; the input of the second branch is the feature map , which includes a global average pooling layer and two fully connected layers connected in sequence, and an adaptive average pooling layer and two PW convolutional layers connected in sequence. In the second branch, the outputs of the two are multiplied and then output through the sigmoid layer.

[0056] For example, Figure 5 , which is the architecture diagram of the patch attention module of this embodiment.

[0057] Among them, the output of the skip connection is:

[0058]

[0059] Where, Indicates that the multiplication operation is performed on each element one by one; and Represents the branches for processing low-level features and high-level features respectively:

[0060]

[0061] Where, represents the sigmoid function, represents the ReLu activation function, represents the batch normalization function; and is a point-by-point convolutional layer with BN, whose kernel sizes are (C / 4)×C×1×1 and C×(C / 4)×1×1, respectively, where C is the number of channels of the feature map.

[0062] This example uses a series of pixel-level convolutional layers to construct a simple encoder-decoder for each pixel in the low-level features. Therefore, the features of each pixel are aggregated separately along the channel dimension, providing key details of the low-level features. In the high-level feature branch:

[0063]

[0064] in, and Represents global average pooling and adaptive average pooling, corresponding to different pooling strategies of the two branches respectively; and Represent different linear weights respectively. The output size of adaptive average pooling is B×C×S×S, where S is adjustable.

[0065] The patch channel attention branch of this embodiment divides the features into S×S patches and performs channel attention calculation on each patch. Furthermore, the feature map output by the jump connection is passed through the RSU of this layer to obtain the output feature F of this layer, and F is output to the RSU of the next layer after upsampling. Thus, the output of the entire network model is obtained .

[0066] In an optional embodiment, the method further comprises the following steps:

[0067] Use the infrared camera in the intelligent reconnaissance equipment to collect raw infrared images and preprocess them to construct a training data set;

[0068] The training data set is input into the infrared image small target detection model to obtain D likelihood graph;

[0069] The objective function is constructed based on BCELoss, and the likelihood graph is input into the objective function to calculate the loss; its expression is:

[0070]

[0071]

[0072] in, is the index of the likelihood graph, is the predicted probability for each pixel, is the true label; N is the number of pixels, i is the pixel index; with the goal of minimizing the loss value, the infrared image weak target detection model is optimized according to the back propagation algorithm and the parameters are saved.

[0073] To obtain a better model, this embodiment establishes an objective function to minimize the loss value and modifies the model parameters according to the back-propagation algorithm to achieve the purpose of optimizing the model. Furthermore, new infrared images from intelligent reconnaissance equipment are used to evaluate the model.

[0074] This embodiment can also optionally further optimize the model based on indicators such as detection effect and speed to meet the needs of intelligent reconnaissance equipment for small target detection.

[0075] In an optional embodiment, the method further includes: inputting the test data set into the trained infrared image small target detection model, and evaluating the model effect according to the evaluation indicators IoU, nIoU, and ROC curve.

[0076] Among them, the closer the values ​​of IoU and nIoU are to 1, the closer the ROC curve is to the upper left corner, and the better the model effect.

[0077] Example 2

[0078] This embodiment proposes a small target detection system for infrared images that combines a U-shaped network and patch attention with intelligent reconnaissance equipment, and applies the small target detection method for infrared images proposed in Example 1. Figure 6 FIG. 1 is a diagram showing the architecture of the infrared image small target detection system of this embodiment.

[0079] The infrared image small target detection system proposed in this embodiment includes:

[0080] An image acquisition module, including an infrared camera within the intelligent reconnaissance equipment, is used to acquire infrared images of the target;

[0081] A preprocessing module, used for performing image preprocessing on the collected infrared images;

[0082] The target detection module is equipped with a UIU-Net-based infrared image small target detection model, which is used to extract features from the input preprocessed infrared image and decode the extracted features to obtain target detection results; wherein, the attention module in the infrared image small target detection model includes a patch attention module.

[0083] It can be understood that the system of this embodiment corresponds to the method of the above-mentioned embodiment 1, and the options in the above-mentioned embodiment 1 are also applicable to this embodiment, so they will not be described again here.

[0084] The terms in the drawings are for illustrative purposes only and are not to be construed as limiting the present invention;

[0085] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A method for detecting small targets in infrared images using a U-shaped network and patch attention for intelligent reconnaissance equipment, characterized in that: The following steps are involved: Use the infrared camera in the intelligent reconnaissance equipment to collect target infrared images and perform pre-processing; Inputting the preprocessed infrared image into an infrared image small target detection model based on UIU-Net for feature extraction, and decoding the extracted features to obtain a target detection result; wherein the attention module in the infrared image small target detection model includes a patch attention module; The infrared image small target detection model includes an encoding network and a decoding network; wherein the encoding network includes at least 6 first convolutional layers connected in sequence, each of which includes an RSU unit composed of several layers of U-Net; the decoding network includes at least 5 second convolutional layers, each second convolutional layer has the same composition as the first convolutional layer of its corresponding level, and the second convolutional layer and the first convolutional layer of the corresponding level are skipped; The patch attention module includes a first branch and a second branch of a jump connection; wherein the input of the first branch is the upsampling result X low , which includes two PW convolutional layers and one sigmoid layer connected in sequence; the input of the second branch is the feature map X high , which includes a global average pooling layer and two fully connected layers connected in sequence, and an adaptive average pooling layer and two PW convolutional layers connected in sequence. In the second branch, the outputs of the two are multiplied and then output through the sigmoid layer; then, the output of the patch attention module is expressed as: L(X low )=σ(PW2(ReLU(PW1(X low ))) P(X high )=σ(BN(W2ReLU(BN(W1GAP(X high ))))+PW2(ReLU(PW1(AAP(X high ))) in, Indicates the element-by-element multiplication operation; L(·) and P(·) represent the branches for processing low-level features and high-level features, respectively; σ(·) represents the sigmoid function, ReLU(·) represents the ReLu activation function, and BN(·) represents the batch normalization function; PW1 and PW2 are point-by-point convolution layers with BN, and their kernel sizes are (C / 4)×C×1×1 and C×(C / 4)×1×1, respectively, where C is the number of channels of the feature map; GAP(·) and AAP(·) represent global average pooling and adaptive average pooling, corresponding to different pooling strategies of the two branches, respectively; W1 and W2 represent different linear weights, respectively.

2. The infrared image small target detection method according to claim 1, characterized in that: The steps of preprocessing the acquired target infrared image include: Convert the acquired target infrared image into a tensor of size H×W; where H and W are the height and width of the image respectively; Normalize the pixel values ​​of the image tensor to the range [-1, 1].

3. The infrared image small target detection method according to claim 2, characterized in that: The encoding network includes six first convolutional layers stage1 to stage6 connected in sequence, where: The first convolutional layer stage1 consists of 7 layers of U-Net, and the output of the 7th layer of U-Net is connected to an extended convolution block with an expansion rate of 2; The structures of the first convolutional layers stage2, stage3, and stage4 are the same as that of the first convolutional layer stage1, and the number of U-Net layers inside them decreases layer by layer compared to the first convolutional layer stage1. The first convolutional layers stage5 and stage6 include 4 layers of U-Net, and the output of each layer of U-Net is connected to an extended convolution block with an expansion rate greater than 1; the output of the first convolutional layer stage5 is connected to a maximum pooling layer; the output of the first convolutional layer stage6 is input to the decoding network after passing through the upsampling layer.

4. The infrared image small target detection method according to claim 3, characterized in that: The decoding network includes five second convolutional layers stage1d to stage5d connected in sequence, which are located at the same level and have the same structure as the first convolutional layers stage1 to stage5; wherein the input of the second convolutional layer is the upsampling result X of the output of the previous layer low The feature map X output by the first convolutional layer at the same level high Input the output Z of the patch attention module; the output of the decoding network Expressed as: Among them, F k represents the output feature of the kth layer in the decoding network, K is the number of layers in the entire network; f(x) represents the intermediate feature map of the input; U k (f(x)) represents the RSU module of the kth layer.

5. The infrared image small target detection method according to any one of claims 1 to 4, characterized in that: The method further comprises the following steps: Use the infrared camera in the intelligent reconnaissance equipment to collect raw infrared images and preprocess them to construct a training data set; Inputting the training data set into the infrared image small target detection model to obtain D likelihood graphs; The objective function is constructed based on BCELoss, and the likelihood graph is input into the objective function to calculate the loss; its expression is: Wherein, d is the index of the likelihood graph, p is the predicted probability of each pixel, and y is the true label; N is the number of pixels, and i is the pixel index; with the goal of minimizing the loss value, the infrared image weak target detection model is optimized and the parameters are saved according to the back propagation algorithm.

6. The infrared image small target detection method according to claim 5, characterized in that: The method further comprises the following steps: The test dataset is input into the trained infrared image small target detection model, and the model effect is evaluated based on the evaluation indicators IoU, nIoU, and ROC curve.

7. An infrared image small target detection system combining a U-shaped network and patch attention for intelligent reconnaissance equipment, applying the infrared image small target detection method according to any one of claims 1 to 6, characterized in that: include: An image acquisition module, including an infrared camera within the intelligent reconnaissance equipment, is used to acquire infrared images of the target; A preprocessing module, used for performing image preprocessing on the collected infrared images; The target detection module is equipped with a UIU-Net-based infrared image small target detection model, which is used to extract features from the input preprocessed infrared image and decode the extracted features to obtain target detection results; wherein, the attention module in the infrared image small target detection model includes a patch attention module.

Citation Information

Patent Citations

  • Infrared weak and small target detection method based on asymmetric attention feature fusion

    CN113591968A

  • Multi-task joint perception network model and detection method for traffic road surface information

    US20240420487A1