Weak and small target detection method of intelligent reconnaissance equipment in combination with U-shaped network and patch attention infrared image

By combining the infrared image weak target detection method with U-shaped network and patch attention module, the problem that intelligent reconnaissance equipment is difficult to detect weak targets in complex backgrounds is solved, and the detection accuracy and robustness are improved.

CN120279248AActive Publication Date: 2025-07-08GUANGDONG UNIV OF TECH
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510341242.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

When existing intelligent reconnaissance equipment deals with weak infrared images, it is difficult to distinguish targets in complex backgrounds, and has limited computing resources, making it difficult to deploy large-scale deep learning models, and lacks detection accuracy and robustness.

Method used

The infrared image weak object detection method combining U-shaped network and patch attention module is adopted. By improving the UIU-Net model, the patch attention module is introduced to learn local and global information and enhance detection capabilities.

Benefits of technology

The target detection performance of intelligent reconnaissance equipment in complex environments is improved, the detection ability of weak targets in infrared images is enhanced, and the robustness and detection accuracy of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279248A_ABST
    Figure CN120279248A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection, and provides a weak and small target detection method for intelligent reconnaissance equipment in combination with a U-shaped network and a patch attention infrared image, which comprises the following steps of: acquiring a target infrared image by using an infrared camera in the intelligent reconnaissance equipment and preprocessing the target infrared image; inputting the preprocessed infrared image into an infrared image weak and small target detection model based on UIU-Net for feature extraction, and decoding the extracted features to obtain a target detection result; wherein the attention module in the infrared image weak and small target detection model comprises a patch attention module. A UIU-Net-based infrared image weak and small target detection model is improved by adopting a patch attention module, so that the model can collect global information and learn knowledge of local patches, interference far away from a target is rejected, a weak target is enhanced, the detection capability of a network on the infrared small target is further enhanced, and the detection efficiency of the network on the infrared small target is improved. And the network has better performance when processing a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and more specifically, to a method for detecting small and weak targets in infrared images of intelligent reconnaissance equipment by combining a U-shaped network and patch attention. Background Art

[0002] Intelligent reconnaissance equipment is widely used in civilian and military applications. It can carry an infrared sensor to identify and track targets. However, for small and weak targets in infrared images, they are often submerged in complex backgrounds and are difficult to distinguish from the images. Moreover, small targets do not have obvious texture and color characteristics, and it is difficult to detect them relying on traditional image processing methods. In addition, the computing resources and storage space in intelligent reconnaissance equipment are limited, and it is difficult to deploy large deep learning models. Moreover, its operating environment is complex and changeable, requiring the image processing method to have a certain degree of robustness. To solve these problems, it is necessary to design a special infrared small target detection algorithm, taking into account multiple factors such as detection accuracy, speed, and resources, and at the same time improving the robustness of the algorithm to adapt to complex environments.

[0003] Deep learning-based methods have become the most widely used technology for detecting small infrared objects. At present, there is a proposal to model the problem as semantic segmentation rather than detection, which helps to improve the detection effect of small targets. Among them, UIU-Net (U-Net in U-Net) uses a resolution-maintaining network to learn multi-scale features and an interactive cross-attention module to fuse features. However, its attention module has disadvantages such as being unable to well understand local semantics. Summary of the Invention

[0004] The present invention aims to overcome the above-mentioned defects of the prior art and provides a method for detecting small and weak targets in infrared images of intelligent reconnaissance equipment by combining a U-shaped network and patch attention.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows: A method for detecting small and weak targets in infrared images of intelligent reconnaissance equipment by combining a U-shaped network and patch attention, comprising the following steps: Collect a target infrared image using an infrared camera in the intelligent reconnaissance equipment and perform preprocessing; Input the preprocessed infrared image into an infrared image small and weak target detection model based on UIU-Net for feature extraction, and decode the extracted features to obtain a target detection result; wherein, the attention module in the infrared image small and weak target detection model includes a patch attention module.

[0006] Furthermore, the present invention also proposes a system for detecting small and weak targets in infrared images of intelligent reconnaissance equipment by combining a U-shaped network and patch attention, applying the infrared image small and weak target detection method proposed by the present invention. Among them, the system includes: The image acquisition module includes an infrared camera in the intelligent reconnaissance equipment, which is used to acquire the infrared image of the target; The preprocessing module is used to preprocess the acquired infrared image; The target detection module is equipped with an infrared image small and weak target detection model based on UIU-Net, which is used to extract features from the input preprocessed infrared image and decode the extracted features to obtain the target detection result; wherein, the attention module in the infrared image small and weak target detection model includes a patch attention module.

[0007] Compared with the prior art, the beneficial effects of the technical solution of the present invention are: In the present invention, a patch attention module is used to improve the infrared image small and weak target detection model based on UIU-Net, so that the model can not only collect global information, but also learn the knowledge of local patches, thereby rejecting the interference far from the target and enhancing the weak target, further strengthening the network's detection ability for infrared small targets, and enabling the network to have better performance when dealing with complex scenes; The present invention is applied to the intelligent pod, which can improve its adaptability to the changing environment and make it easier to detect small and weak targets in the blurred background of the infrared image. Description of the Drawings

[0008] Figure 1 It is a flowchart of the infrared image small and weak target detection method shown according to an embodiment of the present invention.

[0009] Figure 2 It is an architecture diagram of the infrared image small and weak target detection model shown according to an embodiment of the present invention.

[0010] Figure 3 It is an architecture diagram of the RSU unit shown according to an embodiment of the present invention.

[0011] Figure 4 It is a schematic diagram of the skip connection shown according to an embodiment of the present invention.

[0012] Figure 5 It is an architecture diagram of the patch attention module shown according to an embodiment of the present invention.

[0013] Figure 6 It is an architecture diagram of the infrared image small and weak target detection system shown according to an embodiment of the present invention. Detailed Embodiments

[0014] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0015] The terms used in the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0016] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0017] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. Embodiment 1

[0018] This embodiment proposes an infrared image small and weak target detection method. As Figure 1 shown, it is a flowchart of the infrared image small and weak target detection method of this embodiment.

[0019] The infrared image small and weak target detection method proposed in this embodiment includes the following steps: S1. Use the infrared camera in the intelligent reconnaissance equipment to collect the target infrared image and perform preprocessing; S2. Input the preprocessed infrared image into the infrared image small and weak target detection model based on UIU-Net for feature extraction, and decode the extracted features to obtain the target detection result; Among them, the attention module in the infrared image small and weak target detection model includes a patch attention module.

[0020] In the infrared image small and weak target detection model of this embodiment, the UIU-Net network is used as the basic architecture. Among them, UIU-Net is a deep learning framework that improves the detection accuracy of small objects in infrared images through a nested U-Net structure. The common UIU-Net network generally uses a resolution-maintaining network to learn multi-scale features and adopts an interactive cross-attention module to fuse features. However, the cross-attention module cannot well understand local semantics. Therefore, when applied to the detection of small and weak targets in infrared images, it has the disadvantage of low detection accuracy.

[0021] In this embodiment, a patch attention module is used to replace the interactive cross-attention module in the original network. This module includes a low-level feature branch to learn details and a high-level feature branch to learn semantic long-range dependencies to extract the advantages of multi-scale features. During use, after the two branches of features are calibrated and fused, dynamic weights can be learned. In addition, this module can be inserted into each network and cooperate with technologies such as loss functions and model compression to optimize the detection performance and speed of infrared small targets and enhance the detection ability of intelligent reconnaissance equipment for infrared small and weak targets.

[0022] In this embodiment, a patch attention module is used to improve the infrared image small and weak target detection model based on UIU-Net, enabling the model to not only collect global information but also learn the knowledge of local patches, thereby rejecting interference far from the target and enhancing weak targets, further strengthening the network's detection ability for infrared small targets and enabling the network to have better performance when processing complex scenes.

[0023] Applying the present invention to an intelligent pod can improve its adaptability to changing environments and make it easier to detect small and weak targets in a blurred background in an infrared image.

[0024] In an optional embodiment, in step S1, the steps of preprocessing the collected target infrared image include: Converting the collected target infrared image into a tensor of size ; where H and W are the height and width of the image respectively; Normalizing the pixel value range of the image tensor to between.

[0025] Exemplarily, the infrared camera used in this embodiment is a Tigris 640 mid-wave infrared camera with a resolution of 320×320, which can provide thermal image information of the target. When preprocessing the collected infrared image, first convert the infrared image into a tensor of size (320, 320), then divide the pixel value by 255, subtract 0.5, and finally multiply by 2, so that the pixel value range is normalized to between [-1, 1].

[0026] Further optionally, the data set is divided into a training set, a validation set, and a test set approximately according to 50%, 20%, and 30% respectively, for pre-training, validating, and testing the infrared image small and weak target detection model.

[0027] In an alternative embodiment, the infrared image small and weak target detection model includes an encoding network and a decoding network; wherein, the encoding network includes at least six sequentially connected first convolutional layers, and each of the first convolutional layers includes an RSU unit composed of several layers of U-Net; the decoding network includes at least five second convolutional layers, each second convolutional layer has the same composition as the first convolutional layer at its corresponding level, and the second convolutional layer and the first convolutional layer at the corresponding level are skip-connected.

[0028] Exemplarily, as Figure 2 shown, is the architecture diagram of the infrared image small and weak target detection model of this embodiment.

[0029] Among them, the RSU unit in this embodiment is a network module that combines a U-shaped structure and a residual connection, aiming to extract and fuse multi-scale features from the input feature map.

[0030] The RSU unit includes several CBR modules. Specifically, the CBR module includes a sequentially connected 3×3 convolutional layer Conv, a batch normalization layer BN, and a ReLU activation function layer, for extracting features of the input image at different scales. As Figure 3 shown, is the architecture diagram of the RSU unit of this embodiment.

[0031] Furthermore, in an alternative example, the encoding network includes six sequentially connected first convolutional layers stage1 to stage6, where: The first convolutional layer stage1 includes seven layers of U-Net, and the output of the seventh layer of U-Net is connected to an expansion convolutional block with an expansion rate of 2; The structures of the first convolutional layers stage2, stage3, and stage4 are the same as the structure of the first convolutional layer stage1, and the number of U-Net layers inside them decreases layer by layer compared with the first convolutional layer stage1; The first convolutional layers stage5 and stage6 include four layers of U-Net, and the output of each layer of U-Net is connected to an expansion convolutional block with an expansion rate greater than 1; the output of the first convolutional layer stage5 is connected to a max pooling layer; the output of the first convolutional layer stage6 is input to the decoding network after passing through an upsampling layer.

[0032] Further, in an optional embodiment, the decoding network includes five successively connected second convolutional layers stage1d to stage5d, which are respectively at the same level as and have the same structure as the first convolutional layers stage1 to stage5; wherein, the input of the second convolutional layer is the upsampling result of the output of the previous layer and the feature map output by the first convolutional layer at the same level input into the patch attention module to obtain the output .

[0033] In this embodiment, the output of the decoding network is expressed as:

[0034] wherein, represents the output feature of the k th layer in the decoding network, K is the number of layers of the entire network; represents the input intermediate feature map; represents the RSU module of the

[0035] Exemplarily, as Figure 4 shown, it is a schematic diagram of the skip connection in this embodiment.

[0036] As an exemplary illustration, in the first convolutional layer stage1, after one CBR operation, a feature map of the same size is output. Then, after 5 CBR and downsampling operations, 1 CBR operation, and 1 expanded CBR operation with an expansion rate of 2, the deep feature map D 1 output by stage1 is input into the corresponding level of the decoding network.

[0037] Further, in the second convolutional layer stage1d of the corresponding level in the decoding network, its input includes the deep feature map D 1 output by stage1, and the output of the previous layer of stage1d, that is, the upsampling result of the feature map D 6 output by stage6 after upsampling operation. stage1d splices the two feature maps to achieve a skip connection, and then performs a transposed convolution upsampling operation to gradually obtain the original-size feature map and output it to stage2.

[0038] Further, the patch attention module includes a first branch and a second branch of the skip connection; wherein, the input of the first branch is the upsampling result , which includes 2 successively connected PW convolutional layers and 1 sigmoid layer; the input of the second branch is the feature map , which includes a globally average pooling layer and two fully connected layers connected in sequence, and an adaptive average pooling layer and two PW convolutional layers connected in sequence. After multiplying the outputs of the two in the second branch, the result is output through a sigmoid layer.

[0039] Exemplarily, as Figure 5 shown, it is the architecture diagram of the patch attention module of this embodiment.

[0040] Among them, the output of the skip connection is:

[0041] In the formula, represents the element-wise multiplication operation; and respectively represent the branches for processing low-level features and high-level features:

[0042] In the formula, represents the sigmoid function, represents the ReLu activation function, represents the batch normalization function; and are pointwise convolutional layers with BN, and their kernel sizes are (C / 4)×C×1×1 and C×(C / 4)×1×1 respectively, where C is the number of channels of the feature map.

[0043] In this embodiment, a continuous pixel-level convolutional layer is used to construct a simple encoder-decoder for each pixel in the low-level features. Therefore, the features of each pixel will be separately aggregated in the channel dimension, thereby providing the key details of the low-level features. In the branch of high-level features:

[0044] Among them, and represent global average pooling and adaptive average pooling, corresponding to different pooling strategies for the two branches respectively; and represent different linear weights respectively. The output size of the adaptive average pooling is B×C×S×S, where S is adjustable.

[0045] And the patch channel attention branch of this embodiment divides the features into S×S patches and calculates channel attention for each patch. Further, the feature map output by the skip connection is passed through the RSU of this layer to obtain the output feature F of this layer, and F is upsampled and output to the next layer of RSU. Thus, the output of the entire network model is obtained .

[0046] In an alternative embodiment, the method further comprises the following steps: Collecting and preprocessing the original infrared images by using the infrared camera in the intelligent reconnaissance equipment to construct a training data set; Inputting the training data set into the infrared image small and weak target detection model to obtain D a likelihood map; Constructing an objective function based on BCELoss, and inputting the likelihood map into the objective function to calculate the loss; its expression is:

[0047]

[0048] where is the index of the likelihood map, is the predicted probability of each pixel, is the true label; N is the number of pixels, i is the pixel index; aiming to minimize the loss value, optimizing the infrared image small and weak target detection model according to the backpropagation algorithm and saving the parameters.

[0049] In order to obtain a better model, in this embodiment, by establishing an objective function, aiming to minimize the loss value, and according to the backpropagation algorithm, modifying the model parameters to achieve the purpose of optimizing the model. Further, new infrared images from the intelligent reconnaissance equipment are used for model evaluation.

[0050] This embodiment also optionally further optimizes the model according to indicators such as detection effect and speed to meet the requirements of the intelligent reconnaissance equipment for small target detection.

[0051] In an alternative embodiment, the method further comprises: inputting the test data set into the trained infrared image small and weak target detection model, and evaluating the model effect according to the evaluation indicators IoU, nIoU, and ROC curve.

[0052] Among them, the closer the values of IoU and nIoU are to 1, and the closer the ROC curve is to the upper left corner, the better the model effect.

[0053] Embodiment 2 This embodiment proposes a small and weak target detection system for infrared images of intelligent reconnaissance equipment that combines a U-shaped network and patch attention, and applies the infrared image small and weak target detection method proposed in Embodiment 1. As Figure 6 shown, it is the architecture diagram of the infrared image small and weak target detection system of this embodiment.

[0054] In the infrared image small and weak target detection system proposed in this embodiment, it includes: The image acquisition module includes an infrared camera in the intelligent reconnaissance equipment for acquiring target infrared images; The preprocessing module is used for preprocessing the acquired infrared images; The target detection module is equipped with an infrared image small target detection model based on UIU-Net, which is used for extracting features from the input preprocessed infrared images and decoding the extracted features to obtain target detection results; among them, the attention module in the infrared image small target detection model includes a patch attention module.

[0055] It can be understood that the system in this embodiment corresponds to the method in Embodiment 1 above, and the optional items in Embodiment 1 above also apply to this embodiment, so they will not be described repeatedly here.

[0056] The terms in the accompanying drawings are only for illustrative purposes and should not be construed as limitations on the present invention; Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for detecting small and weak targets in infrared images by combining a U-shaped network and patch attention in an intelligent reconnaissance equipment, characterized in that, It includes the following steps: Collect the infrared image of the target by using the infrared camera in the intelligent reconnaissance equipment and perform preprocessing; Input the preprocessed infrared image into the infrared small target detection model based on UIU-Net for feature extraction, and decode the extracted features to obtain the target detection result; wherein, the attention module in the infrared small target detection model includes a patch attention module.

2. The infrared image small and weak target detection method according to claim 1, characterized in that The steps of preprocessing the collected infrared image of the target include: Convert the collected target infrared image into a tensor of size; where H and W are the height and width of the image respectively; Normalize the pixel value range of the image tensor to between.

3. The infrared image small and weak target detection method according to claim 1, characterized in that The infrared small target detection model includes an encoding network and a decoding network; wherein, the encoding network includes at least 6 first convolutional layers connected in sequence, and each of the first convolutional layers includes an RSU unit composed of several layers of U-Net; the decoding network includes at least 5 second convolutional layers, and the structure of each second convolutional layer is the same as that of the corresponding first convolutional layer at its level, and the second convolutional layer and the corresponding first convolutional layer at its level are skip-connected.

4. The infrared image small and weak target detection method according to claim 3, characterized in that The encoding network includes 6 first convolutional layers stage1 to stage6 connected in sequence, wherein: The first convolutional layer stage1 includes 7 layers of U-Net, and the output of the 7th layer of U-Net is connected with an expansion convolutional block with an expansion rate of 2; The structures of the first convolutional layers stage2, stage3 and stage4 are the same as that of the first convolutional layer stage1, and the number of layers of the U-Net inside them decreases layer by layer compared with the first convolutional layer stage1; The first convolutional layers stage5 and stage6 include 4 layers of U-Net, and an expansion convolutional block with an expansion rate greater than 1 is connected to the output of each layer of U-Net; the output of the first convolutional layer stage5 is connected with a max pooling layer; the output of the first convolutional layer stage6 is input into the decoding network after passing through an upsampling layer.

5. The infrared image small and weak target detection method according to claim 4, characterized in that The decoding network includes five second convolutional layers stage1d to stage5d connected in sequence, which are respectively at the same level as and have the same structure as the first convolutional layers stage1 to stage5; wherein, the input of the second convolutional layer is the upsampling result output by the previous layer and the feature map output by the first convolutional layer at the same level The output obtained by inputting into the patch attention module ; The output of the decoding network is expressed as: , in, Indicates the decoding network k The output features of the layer, K is the number of layers of the entire network; Represents the intermediate feature map of the input; Represents the RSU module of the kth layer.

6. The method for detecting a small and weak target in an infrared image according to claim 5, wherein The patch attention module includes a first branch and a second branch with skip connections; wherein, the input of the first branch is the upsampling result , which includes 2 PW convolutional layers and 1 sigmoid layer connected in sequence; the input of the second branch is the feature map , which includes a global average pooling layer and 2 fully connected layers connected in sequence, and an adaptive average pooling layer and 2 PW convolutional layers connected in sequence. In the second branch, the outputs of the two are multiplied and then output through a sigmoid layer; then, the output of the patch attention module is expressed as: ; ; ; Among them, represents the operation of multiplying elements one by one; and represent the branches for processing low-level features and high-level features respectively; represents the sigmoid function, represents the ReLu activation function, represents the batch normalization function; and are pointwise convolutional layers with BN, and their kernel sizes are (C / 4)×C×1×1 and C×(C / 4)×1×1 respectively, where C is the number of channels of the feature map; and represent global average pooling and adaptive average pooling, corresponding to different pooling strategies of the two branches respectively; and represent different linear weights respectively.

7. The method for detecting a small and weak target in an infrared image according to any one of claims 1 to 6, characterized in that, The method further includes the following steps: Collect the original infrared image by using the infrared camera in the intelligent reconnaissance equipment and perform preprocessing to construct a training data set; Input the training data set into the infrared image small and weak target detection model to obtain D likelihood maps; Construct an objective function based on BCELoss, and input the likelihood map into the objective function to calculate the loss; its expression is: ; ; Among them, is the index of the likelihood map, is the predicted probability of each pixel, is the ground truth label; N is the number of pixels, i is the pixel index; aiming at minimizing the loss value, the infrared image small target detection model is optimized according to the backpropagation algorithm and the parameters are saved.

8. The infrared image small and weak target detection method according to claim 7, wherein The method further includes the following steps: Input the test data set into the trained infrared small target detection model, and evaluate the model effect according to the evaluation metrics IoU, nIoU, and ROC curve.

9. A small and weak target detection system for infrared images that combines a U-shaped network and patch attention in an intelligent reconnaissance equipment, applying the infrared image small and weak target detection method according to any one of claims 1 to 8, characterized in that, It includes: An image acquisition module, including an infrared camera in the intelligent reconnaissance equipment, for collecting the infrared image of the target; A preprocessing module, for performing image preprocessing on the collected infrared image; A target detection module, on which an infrared small target detection model based on UIU-Net is mounted, for extracting features from the input preprocessed infrared image and decoding the extracted features to obtain the target detection result; wherein, the attention module in the infrared small target detection model includes a patch attention module.

Citation Information

Patent Citations

  • Infrared weak and small target detection method based on asymmetric attention feature fusion

    CN113591968A

  • Single-frame image infrared weak and small target detection method based on deep U-shaped network

    CN115311508A

  • Remote sensing image change detection method combining U-shaped network and self-attention mechanism

    CN116740527A

  • Segmentation method for identifying infrared small target

    CN117115443A

  • Infrared small target identification method based on distraction mining network

    CN117934814A