Infrared Single-frame Small Target Detection Method Based on Attention Mechanism
By constructing a multi-dimensional attention perception network MDA-Net, combining the coded-end interactive guidance and false alarm attention module, the infrared small object detection algorithm has solved the problem of low detection rate and high false alarm in complex backgrounds, and efficient object detection is achieved.
Patent Information
- Application Number
- CN202211086622.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-09-07
AI Technical Summary
The existing infrared small object detection algorithm has a low detection rate and a high false alarm rate under complex backgrounds. The traditional manual design feature operator model has poor generalization capabilities. The feature information extraction of convolutional neural networks is limited by local features and cannot effectively fuse global information.
A multi-dimensional attention perception network MDA-Net is constructed, combining the code-side decoding-side interactive guidance module EDIG and the false alarm attention module AFF, through the channel attention and point-by-point attention of shallow and deep features, the characteristics with high contributions are screened, and non-local attention modules and feature fusion modules are introduced to obtain global context information and reduce the false alarm rate.
It improves the infrared small object detection rate, reduces the false alarm rate, and realizes accurate detection in complex backgrounds, which is highly robust.
Smart Images

Figure CN115375668B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing and computer vision, and particularly relates to an infrared single-frame small target detection method, which can be used for accurate detection of infrared small targets in complex backgrounds. Background Art
[0002] In recent years, computer vision technology has developed rapidly and has been widely applied in various fields. As an important branch of computer vision technology, infrared small target detection has high application value in precise guidance, weapon manufacturing, monitoring and early warning, etc. due to the unique advantages of infrared sensors such as all-weather operation, strong anti-interference performance, and convenient on-board installation. Therefore, infrared small target detection technology has attracted the attention of experts and scholars around the world and has become one of the research hotspots in recent years.
[0003] Currently, the key problem of infrared small target detection algorithms is how to accurately locate and segment targets in infrared images with complex backgrounds, improve the detection rate and reduce the false alarm rate. The main detection algorithms are divided into traditional infrared small target detection algorithms and deep learning-based infrared small target detection algorithms.
[0004] Traditional infrared small target detection mainly relies on traditionally hand-designed features for detection, that is, modeling infrared small targets as outliers popping out from the background, mainly divided into filter-based methods, local contrast-based methods, and low-rank-based methods, among which:
[0005] Filter-based methods directly perform target segmentation from the filtered image by setting a threshold. Typical methods include those proposed by Deshpande et al. in "Max-mean and max-median filters for detection of small targets" [C] / / Signal and Data Processing of Small Targets 1999. International Society for Optics and Photonics, 1999, 3809: 74-83, which apply Max-mean and Max-median filters to the detection of infrared small targets; Zeng et al. in "The design of top-hat morphological filter and application to infrared target detection" [J]. Infrared physics & technology, 2006, 48(1): 67-76, who proposed the Top-hat filter. This method uses mathematical morphology for small target detection. By performing an opening operation to remove small targets in the image to obtain the background image, and then subtracting the original image from the background image, small target detection can be achieved. Such filter-based methods are easily affected by clutter and noise in the background, resulting in poor detection robustness.
[0006] Local-contrast-based methods mainly enhance the target signal through local contrast processing to improve detection accuracy. Typical methods include those by Chen et al. in "A local contrast method for small infrared target detection" [J]. IEEE transactions on geoscience and remote sensing, 2013, 52(1): 574-581, which enhance the target or suppress the background by obtaining the local contrast map of the infrared image and calculating the difference between the current position and its surrounding area; Han et al. in "A robust infrared small target detection algorithm based on human visual system" [J]. IEEE Geoscience and Remote Sensing Letters, 2014, 11(12): 2168-2172, who proposed an improved local contrast metric based on the HVS contrast mechanism, which has good robustness for infrared small target detection. Such local-contrast-based methods have a high false detection rate due to being easily affected by factors such as edges and noise.
[0007] The low-rank based method assumes that the target image patch is a sparse matrix and the background is a low-rank matrix, transforms the small target detection problem into an optimization problem to recover the low-rank matrix and the sparse matrix, so as to achieve small target detection. Representative ones include the IPI algorithm proposed by Gao et al. in Infrared patch-image model for small target detection in asingle image[J].IEEE transactions on image processing,2013,22(12):4996-5009; the NRAM algorithm proposed by Zhang et al. in Infrared small target detection via non-convex rankapproximation minimization joint l2,1norm[J].Remote Sensing,2018,10(11):1821. The deficiencies of this kind of method are: being sensitive to pixel values and taking a long time.
[0008] With the rise of deep learning technology, infrared small target detection methods based on deep learning have gradually become a research hotspot in the current field. They have powerful model fitting capabilities. By inputting a large amount of data into a convolutional neural network for training to actively learn features, the purpose of detecting small targets is achieved, and better detection performance is realized. Liu et al. first used a convolutional neural network for infrared small target detection in "Imagesmall target detection based on deep learning with SNR controlled samplegeneration[J].Current Trends in Computer Science and Mechanical Automation,2017,1:211-220". This method first randomly extracts the background part from some cloudy sky images, then adds randomly generated target points to the background with a controlled signal-to-noise ratio, and then conducts training and testing to achieve the detection of small targets. Dai et al. proposed a public dataset SIRST for small target detection in single-frame infrared images and designed an asymmetric context ACM module in "Asymmetric contextual modulation forinfrared small target detection[C] / / Proceedings of the IEEE / CVF WinterConference on Applications of Computer Vision.2021:950-959". This module has good detection performance on the SIRST dataset. In the same year, Dai et al. also proposed an attentional local contrast network ALCNet in "Attentional local contrastnetworks for infrared small target detection[J].IEEE Transactions onGeoscience and Remote Sensing,2021,59(11):9813-9824". This method combines a convolutional neural network with a conventional model-driven method to highlight and retain the features of small targets, further improving the detection performance.Tong et al. proposed an enhanced asymmetric attention EAA in "EAAU-Net: Enhanced Asymmetric Attention U-Net for Infrared Small Target Detection" [J]. Remote Sensing, 2021, 13(16): 3200. This method effectively realizes the exchange of spatial and channel feature information within the network layer, achieves the effective fusion of context information, and further improves the detection performance on the SIRST dataset. However, the attention mechanisms used in these deep learning-based detection methods are relatively single and limited, and less global information is considered. More effective attention is not fused and applied to infrared small target detection. In the case of a more complex background, a large number of false alarms usually appear in the detection results.
[0009] In summary, the existing single-frame infrared small target detection algorithms mainly have the following deficiencies: First, using traditional manually designed feature operators for detection results in poor generalization ability of the algorithm model, resulting in serious misdetection phenomena in complex scenarios; second, the extraction of feature information by the convolutional neural network model is limited within the field of view of the convolutional kernel, so it is impossible to extract more effective information by combining global information from limited local features, resulting in a lower detection rate and a higher false alarm rate for infrared small target detection. Summary of the Invention
[0010] The purpose of the present invention is to propose an infrared small target detection method based on an attention mechanism in view of the above deficiencies of the existing technology, so as to improve the detection rate of infrared small targets, reduce the false alarm rate, and enhance the detection performance in complex backgrounds.
[0011] To achieve the above purpose, the technical solution of the present invention includes the following:
[0012] (1) Select a set of labeled datasets from the publicly available infrared small target datasets, and perform random scaling, random cropping, or zero-padding operations in the range of 0.7 to 1.7 in sequence to obtain training sets and test sets with a unified size of 480×480;
[0013] (2) Construct a multi-dimensional attention perception network MDA-Net under the Pytorch framework:
[0014] (2a) Establish an encoding-decoding end interaction guidance module EDIG composed of a shallow channel attention sub-module, a deep channel attention sub-module, and a pointwise attention sub-module;
[0015] (2b) Establish a false alarm attention module AFF composed of a connection of a non-local attention module and a non-local feature fusion module;
[0016] (2c) Select three existing convolutional operation units, one max-pooling operation unit, two upsampling modules, and eighteen residual blocks to form the backbone network of an eight-layer encoder-decoder structure;
[0017] (2d) Embed two EDIG modules constructed in (2a) and one AFF module constructed in (2b) into the backbone network of the eight-layer encoder-decoder structure to form a multi-dimensional attention perception network under the Pytorch framework, and use the IoU Loss function as the loss function of this network;
[0018] (3) Use the training set and its annotation information to train the multi-dimensional attention perception network by the gradient descent method to obtain a trained multi-dimensional attention perception network;
[0019] (4) Input the test set into the trained multi-dimensional attention perception network and output the infrared small target detection results.
[0020] Compared with the prior art, the present invention has the following beneficial effects:
[0021] First, the present invention establishes a multi-dimensional attention perception network MDA-Net model based on the attention mechanism, and introduces an encoder-decoder interaction guidance EDIG module at the encoding end and decoding end of the network. By applying channel attention blocks to shallow and deep features to screen out channel features with a greater contribution to the target, the effective learning of small target features by the network is improved. At the same time, pointwise attention is used to aggregate the spatial position context information of shallow features, and it is embedded into deep features in a bottom-up modulation manner, which can realize the guidance of low-level detail information to high-level semantic information, effectively restore the full resolution space of the target, and improve the infrared small target detection rate.
[0022] Second, aiming at the problems that infrared small targets occupy few pixels, lack clear texture shapes, are vulnerable to clutter and noise interference in complex environments, and there are serious missed detections and false detections in the target detection task, the present invention designs a false alarm attention module AFF. This module includes a non-local attention module and a non-local feature fusion module. Since the non-local attention module introduces non-local operations in convolutional and pooling operations, it effectively breaks out of the limitation of the local receptive field, can realize the exploration of global features, and obtain rich context information. At the same time, since the non-local feature fusion module can obtain the dependency relationship between deep features and shallow features, it can assist high-level features to better learn false alarm information, and significantly reduce the detection false alarm rate.
[0023] Experimental results show that the present invention can accurately locate and segment small targets in different scenarios, performs outstandingly in quantitative and qualitative results, has high robustness, effectively improves the detection rate of infrared small targets, reduces the detection false alarm rate, and has good detection performance. Description of the Drawings
[0024] Figure 1 This is the overall flowchart for the implementation of the present invention;
[0025] Figure 2 This is the schematic structural diagram of the Encoding-Decoding Interaction Guidance (EDIG) module constructed in the present invention;
[0026] Figure 3 This is the schematic structural diagram of the False Alarm Attention (AFF) module constructed in the present invention;
[0027] Figure 4 This is the structural diagram of the Multi-Dimensional Attention Perception Network (MDA-Net) constructed in the present invention;
[0028] Figure 5 This is the comparison chart of the detection effects of the infrared small target data using the present invention and the existing infrared small target detection algorithms. Detailed implementation manners
[0029] The following further illustrates the embodiments and effects of the present invention in conjunction with the accompanying drawings:
[0030] In this embodiment, the single-frame infrared small and weak target image dataset SIRST established by Dai et al. is used for infrared small target detection.
[0031] Refer to Figure 1 , and the specific implementation of this example is as follows:
[0032] Step 1: Dataset preprocessing.
[0033] To enhance the network's processing ability for input targets of different scales, the dataset needs to be preprocessed, and the specific implementation is as follows:
[0034] 1.1) Statistically analyze the sizes of the infrared images in the dataset. According to the statistical result that 99.9% of the image widths and heights are within 500, the basic size of the input image is selected as 512×512;
[0035] 1.2) Randomly scale the input image within the range of 512×0.7 to 512×1.7;
[0036] 1.3) Perform random cropping or zero-padding operations on the randomly scaled images to obtain training and test sets with a unified size of 480×480.
[0037] Step 2: Construct the Encoding-Decoding Interaction Guidance (EDIG) module.
[0038] Refer to Figure 2 , and the specific implementation of this step is as follows:
[0039] 2.1) Establish a shallow channel attention sub-module and a deep channel attention sub-module, as shown in Figure 2 (a), where:
[0040] The shallow channel attention sub-module and the deep channel attention sub-module have the same structure, and both include a global average pooling layer, two fully connected layers, a ReLU activation function layer, and a sigmoid function layer;
[0041] The structure of the shallow channel attention sub-module is: shallow input port → global average pooling layer → first fully connected layer → ReLU activation function layer → second fully connected layer → sigmoid function layer. After the output of the sigmoid function layer is multiplied by the original input features of this sub-module, the output result of this sub-module is obtained;
[0042] The structure of the deep channel attention sub-module is: deep input port → global average pooling layer → first fully connected layer → ReLU activation function layer → second fully connected layer → sigmoid function layer. After the output of the sigmoid function layer is multiplied by the original input features of this sub-module, the output result of this sub-module is obtained;
[0043] 2.2) Establish a pointwise attention sub-module, as shown in Figure 2 (b), where:
[0044] The pointwise attention sub-module includes two pointwise convolutional layers, a ReLU activation function layer, and a sigmoid function layer;
[0045] This module has two input ports (shallow and deep) and one output port. Its structure is: input port from the shallow layer → first pointwise convolutional layer → ReLU activation function layer → second pointwise convolutional layer → sigmoid function layer. The output of this sigmoid function layer is multiplied by the input features from the deep layer of this sub-module to obtain the output result of this sub-module;
[0046] 2.3) Establish an Encoder-Decoder Interaction Guidance Module (EDIG) composed of a shallow channel attention sub-module, a deep channel attention sub-module, and a pointwise attention sub-module, as shown in Figure 2 (c). The structural relationship of this EDIG module is:
[0047] The shallow channel attention sub-module and the deep channel attention sub-module are respectively connected to the shallow input port and the deep input port of the pointwise attention sub-module. And the output results of the deep channel attention sub-module and the shallow channel attention sub-module are multiplied pixel by pixel and then added to the result output by the pointwise attention sub-module. The added result is the output result of the Encoder-Decoder Interaction Guidance (EDIG) module.
[0048] Step 3: Construct the false alarm attention module AFF.
[0049] Refer to Figure 3 , the specific implementation of this step is as follows:
[0050] 3.1) Establish a non-local attention module composed of three parallel branches. This module has one input port and one output port, as shown in Figure 3 (a): The structure of each branch is: input port → convolutional layer → Reshape layer, the convolutional kernel size of each is 1*1, and the convolutional stride of each is 1;
[0051] The input of this non-local attention module is X. The output R1(f(X)) of the first branch is multiplied by the output R2(f(X)) of the second branch to obtain the first matrix: E(X) = R1(f(X)) · R2(f(X)), where f(·) represents the convolutional operation and R(·) represents the Reshape operation;
[0052] The output R3(f(X)) of the third branch is multiplied by the first matrix E(X) to obtain the second matrix: D(X) = R3(f(X)) · E(X);
[0053] The second matrix D(X) passes through a convolutional layer with a convolutional kernel size of 1*1 and a convolutional stride of 1 to obtain the output feature f(D(X)). This output feature f(D(X)) is added to the input X pixel by pixel to obtain the output result of this module: Y(X) = f(D(X)) + X;
[0054] 3.2) Establish a non-local feature fusion module composed of 3 parallel branches, as shown in Figure 3 (b), where:
[0055] The structure of the first branch is: deep input port → convolutional layer → Reshape layer, the convolutional kernel size of this convolutional layer is 1*1, and the convolutional stride of each is 1;
[0056] The structures of the second and third branches are the same. They are in turn: shallow input port → max pooling layer → convolutional layer → Reshape layer, the convolutional kernel size of this convolutional layer is 1*1, the convolutional stride of each is 1, and the convolutional kernel size of the max pooling layer is 1*1;
[0057] The input of the deep input port is X h , the input of the shallow input port is X l , the output R1(f(X h )) of the first branch is multiplied by the output R2(f(MaxPool(X l ))) of the second branch to obtain the matrix: E(X hl ) = R1(f(X h))·R2(f(MaxPool(X l ))), where MaxPool(·) represents the max pooling operation;
[0058] The output of the third branch R3(f(MaxPool(X l ))) is multiplied pixel by pixel with the matrix E(X hl ) to obtain the final output result of the module: Y(X hl ) = R3(f(MaxPool(X l )))·E(X hl );
[0059] 3.3) Connect the output port of the non-local attention module to the deep input port of the first branch of the non-local feature fusion module to form the false alarm attention module AFF.
[0060] The relevant non-local operations mainly capture long-range dependencies by calculating the similarity between two positions. Commonly used similarity functions mainly include Gaussian function, embedded Gaussian function, dot product similarity, and cascade function, etc. This example uses but is not limited to using the dot product similarity function for calculation.
[0061] Step 4: Construct a multi-dimensional attention perception network MDA-Net under the Pytorch framework.
[0062] Refer to Figure 4 , the specific implementation of this step is as follows:
[0063] 4.1) Select three existing convolutional operation units, one max pooling operation unit, two upsampling modules, and eighteen residual blocks to form a backbone network with an eight-layer encoder-decoder structure. Its first four layers are the encoding layers, and the last four layers are the decoding layers. The structures of each layer are as follows:
[0064] The first layer: The first convolutional operation unit → the max pooling operation unit, and its output feature is 8-dimensional, with a size of 240×240;
[0065] The second layer: The first residual block → the second residual block → the third residual block, and its output feature is 16-dimensional, with a size of 240×240;
[0066] The third layer: The fourth residual block → the fifth residual block → the sixth residual block → the seventh residual block, and its output feature is 32-dimensional, with a size of 120×120;
[0067] The fourth layer: The eighth residual block → the ninth residual block → the tenth residual block → the eleventh residual block, and its output feature is 64-dimensional, with a size of 60×60;
[0068] The fifth layer: the first upsampling module → the twelfth residual block → the thirteenth residual block → the fourteenth residual block → the fifteenth residual block, whose output features are 32 - dimensional and the size is 120×120;
[0069] The sixth layer: the second upsampling module → the sixteenth residual block → the seventeenth residual block → the eighteenth residual block, whose output features are 16 - dimensional and the size is 240×240;
[0070] The seventh layer: the second convolutional operation unit, whose output features are 4 - dimensional and the size is 240×240;
[0071] The eighth layer: the third convolutional operation unit, whose output features are 1 - dimensional and the size is 240×240;
[0072] The convolutional kernel sizes of the first convolutional operation unit and the second convolutional operation unit are both 3*3, and the strides are both 1; the convolutional kernel size of the third convolutional operation unit is 1*1, and the stride is 1;
[0073] The convolutional kernel size of the max - pooling operation unit is 3*3, and the stride is 2;
[0074] For the two upsampling modules, the convolutional kernel sizes are both 4*4, and the strides are both 2;
[0075] The convolutional kernel sizes of the eighteen residual blocks are all 3*3, and the strides are all 1;
[0076] The fourth residual block and the eighth residual block both contain two convolutional layers and an average - pooling layer; the remaining residual blocks are all composed of two convolutional layers;
[0077] 4.2) Embed the Encoding - Decoding Interaction Guidance (EDIG) module and the False Alarm Attention (AFF) module in the backbone network of the eight - layer encoding - decoding structure, and the specific implementation is as follows:
[0078] The shallow - channel attention sub - module of the EDIG1 module is connected to the third residual block of the second layer in the backbone network, the deep - channel attention sub - module of the EDIG1 module is connected to the second upsampling module of the sixth layer in the backbone network, and the output port of the EDIG1 module is connected to the sixteenth residual block of the sixth layer in the backbone network;
[0079] The non - local attention module in the AFF module is connected to the seventh residual block of the third layer in the backbone network, the non - local feature fusion module in the AFF module is connected to the third residual block of the second layer in the backbone network, and the output port of the non - local feature fusion module is respectively connected to the eighth residual block of the fourth layer in the backbone network and the shallow - channel attention sub - module of the EDIG2 module;
[0080] The deep channel attention sub-module of the EDIG2 module is connected to the first upsampling module of the fifth layer in the backbone network, and the output port of the EDIG2 module is connected to the twelfth residual block of the fifth layer in the backbone network.
[0081] Step 5: Use the training set and its annotation information to train the multi-dimensional attention perception network by the gradient descent method.
[0082] The specific implementation of this step is as follows:
[0083] 5.1) Divide the infrared small target training set and its corresponding annotation data into multiple paired image groups according to the batch size, and input the first image group into the multi-dimensional attention perception network MDA-Net to obtain the weights and biases of each convolutional operation of the network and the results predicted by the network;
[0084] 5.2) According to the results predicted by the network and the annotation data of the infrared image, calculate its loss value through the loss function of the network:
[0085]
[0086] where p i,j represents the prediction result of the network at the i-th row and j-th column, t i,j represents the annotation data of the infrared image at the i-th row and j-th column, p represents the prediction result of the network for a group of images, t represents the annotation data for a group of images, and L IoU (p, t) represents the network loss value under this set of data of the network prediction result p and the annotation data t;
[0087] 5.3) Use the Nesterov Accelerated Gradient algorithm to update the gradient direction, aim to minimize the loss function value, update the parameters in the network, and obtain the multi-dimensional attention perception network after one parameter update;
[0088] 5.4) Input the second image group into the multi-dimensional attention perception network after one parameter update, repeat steps 5.1) to 5.3), and obtain the multi-dimensional attention perception network after two parameter updates; and so on, until the last image group is input into the multi-dimensional attention perception network after the previous update, and obtain the multi-dimensional attention perception network after one training;
[0089] 5.5) Input all image groups into the multi-dimensional attention perception network after one training in turn, repeat steps 5.1) to 5.4), and obtain the multi-dimensional attention perception network after two trainings; and so on, until all image groups are input 1200 times to obtain the trained multi-dimensional attention perception network.
[0090] Step 6: Input the test set into the trained multi-dimensional attention-aware network and output the infrared small target detection results.
[0091] The effects of the present invention are further illustrated by the following simulations:
[0092] I. Test conditions
[0093] Data: The SIRST dataset proposed by Dai et al. is adopted;
[0094] Experimental platform: The CPU is Intel(R) Core(TM) i7-7500U @ 2.70GHz, 8GB RAM, the operating system is Ubuntu18.04, the graphics card uses RTX 2080, the cuda version is 10.2, and the Pytorch version is 1.10;
[0095] Parameter settings: The training batch size of the present invention is set to 32, the learning rate is set to 0.05, and a total of 1200 rounds of training are performed; the parameter settings of the deep learning methods U-Net algorithm, ACM-U-Net algorithm, and ALCNet algorithm are the same as those in the original papers; the parameter settings of the traditional methods Top-hat algorithm, Max-median algorithm, FKRW algorithm, and IPI algorithm are as shown in Table 1 below.
[0096] Table 1 Hyperparameter settings in traditional methods
[0097]
[0098] II. Simulation test content
[0099] Test 1: Use 8 methods including the Top-hat algorithm, Max-median algorithm, FKRW algorithm, IPI algorithm, U-Net algorithm, ACM-U-Net algorithm, ALCNet algorithm, and the present invention to perform infrared small target detection on the SIRST dataset respectively, and calculate 5 objective evaluation indicators including the intersection over union IoU, normalized intersection over union nIoU, detection rate P d , false alarm rate F a and the number of parameters Params. The results are shown in Table 2:
[0100] Table 2 Comparison of detection indicators of different algorithms
[0101]
[0102]
[0103] The calculation formulas of each index in Table 2 are as follows:
[0104]
[0105]
[0106]
[0107]
[0108] Params = (2 × C i × w 2 ) × C o 。
[0109] In the formula, TP is the number of positive samples predicted as positive samples, FP is the number of negative samples predicted as positive samples, TN is the number of negative samples predicted as negative samples, FN is the number of positive samples predicted as negative samples, T is the number of correctly predicted samples, P is the total number of positive samples, N is the total number of samples, w is the size of the convolution kernel, C i is the number of input channels, and C o is the number of output channels.
[0110] As can be seen from Table 2, the present invention has achieved the best performance in the IoU, nIoU, P d , F a indicators, improving the detection rate P d , significantly reducing the false alarm rate F a , and the number of parameters Params is also lower than that of the U-Net algorithm and the ACM-U-Net algorithm, only slightly higher than that of the ALCNet algorithm.
[0111] Test 2: Use 4 methods including the Top-hat algorithm, the Max-median algorithm, the ACM-U-Net algorithm and the present invention to perform infrared small target detection on the SIRST dataset respectively, and the results are as Figure 5 . Among them:
[0112] Figure 5 (a) is the 3D diagram of the detection results of the 4 methods;
[0113] Figure 5 (b) is the plan view of the detection results of the 4 methods. The frame is marked with the detected targets. The enlarged view is at the lower right corner of the image to more intuitively present the fine segmentation results. The dotted circle indicates the misdetection area.
[0114] From Figure 5It can be seen that traditional methods such as the Top-hat algorithm and the Max-median algorithm are prone to multiple false detections and missed detections in complex scenarios. This is because the performance of traditional methods highly depends on manually extracted features and cannot adapt to changes in target scale and scenarios. Although the performance of the ACM-U-Net algorithm has been greatly improved compared with traditional methods, there are still false detection phenomena in the first three figures of this method. The method of the present invention can not only achieve accurate positioning of the target, but also almost achieve the detection performance of zero false detection.
[0115] Thus, it can be seen that the MDA-Net network proposed by the present invention has very high robustness to target size and scene changes, obtains the best detection performance, and can effectively reduce the false alarm rate while increasing the detection rate.
Claims
1. An infrared single-frame small target detection method based on an attention mechanism, characterized in that, The steps are as follows: (1) Select a set of labeled datasets from the publicly available infrared small target dataset, and perform random scaling, random cropping, or zero-padding operations within the range of 0.7 to 1.7 in sequence to obtain training set and test set datasets with a unified size of 480×480; (2) Construct a multi-dimensional attention perception network MDA-Net under the Pytorch framework: (2a) Establish an Encoder-Decoder Interaction Guidance module EDIG composed of a shallow channel attention sub-module, a deep channel attention sub-module, and a pointwise attention sub-module; (2b) Establish a False Alarm Attention module AFF composed of a connection of a non-local attention module and a non-local feature fusion module; (2c) Select three existing convolutional operation units, one max-pooling operation unit, two upsampling modules, and eighteen residual blocks to form a backbone network with an eight-layer encoder-decoder structure; (2d) Embed two EDIG modules constructed in (2a) and one AFF module constructed in (2b) into the backbone network with an eight-layer encoder-decoder structure to form a multi-dimensional attention perception network under the Pytorch framework, and use the IoU Loss function as the loss function of this network; (3) Train the multi-dimensional attention perception network by the gradient descent method using the training set and its annotation information to obtain a trained multi-dimensional attention perception network; (4) Input the test set into the trained multi-dimensional attention perception network to output the infrared small target detection result.
2. The method according to claim 1, wherein: For the Encoder-Decoder Interaction Guidance module EDIG established in step (2a), its structural relationship is as follows: Connect the shallow channel attention sub-module and the deep channel attention sub-module to the shallow input port and the deep input port of the pointwise attention sub-module respectively, and multiply the output results of the deep channel attention sub-module and the shallow channel attention sub-module pixel by pixel and then add them to the result output by the pointwise attention sub-module. The added result is the output result of the Encoder-Decoder Interaction Guidance EDIG module; The shallow channel attention sub-module and the deep channel attention sub-module have the same structure, both of which include a global average pooling layer, two fully connected layers, a ReLU activation function layer, and a sigmoid function layer; the structure of each sub-module is: input port → global average pooling layer → first fully connected layer → ReLU activation function layer → second fully connected layer → sigmoid function layer. After multiplying the output of the sigmoid function layer by the original input features of this sub-module, the output result of this sub-module is obtained; The pointwise attention sub-module includes two pointwise convolutional layers, a ReLU activation function layer, and a sigmoid function layer; this module has two input ports, a shallow one and a deep one, and one output port. Its structure is: input port from the shallow layer → first pointwise convolutional layer → ReLU activation function layer → second pointwise convolutional layer → sigmoid function layer. The output of the sigmoid function layer is multiplied by the input features from the deep layer of this sub-module to obtain the output result of this sub-module.
3. The method according to claim 1, wherein: The non-local attention module in step (2b) consists of three branches in parallel. The structure of each branch is: input port → convolutional layer → Reshape layer. The kernel size of the convolutional layer is 1*1, and the stride is 1. Among them: The output R1(f(X)) of the first branch is multiplied by the output R2(f(X)) of the second branch to obtain the first matrix: E(X) = R1(f(X)) · R2(f(X)), where f(·) represents the convolution operation and R(·) represents the Reshape operation; The output R3(f(X)) of the third branch is multiplied by the first matrix E(X) to obtain the second matrix: D(X) = R3(f(X)) · E(X); The second matrix D(X) passes through a convolutional layer with a kernel size of 1*1 and a stride of 1 to obtain the output feature f(D(X)). This output feature f(D(X)) is added to the input X of the non-local attention module pixel by pixel to obtain the output result of this module: Y(X) = f(D(X)) + X.
4. The method according to claim 1, characterized in that: The non-local feature fusion module in step (2b) consists of three branches in parallel. Among them: The structure of the first branch is: deep input port → convolutional layer → Reshape layer. The kernel size of this convolutional layer is 1*1, and the stride is 1; The structures of the second and third branches are the same. They are in turn: shallow input port → max pooling layer → convolutional layer → Reshape layer. The kernel size of this convolutional layer is 1*1, and the stride is 1. The kernel size of the max pooling layer is 1*1; The output R1(f(X h )) of the first branch is multiplied by the output R2(f(MaxPool(X l ))) of the second branch to obtain the matrix: E(X hl ) = R1(f(X h )) · R2(f(MaxPool(X l ))), where MaxPool(·) represents the max pooling operation; The output R3(f(MaxPool(X l ))) of the third branch is multiplied pixel - by - pixel with the matrix E(X hl ) to obtain the final output result of the module: Y(X hl ) = R3(f(MaxPool(X l )))·E(X hl ).
5. The method according to claim 1, characterized in that: The eight-layer encoder-decoder backbone network constructed in step (2c) has four encoding layers in the front and four decoding layers in the back. The structures of each layer are as follows: The first layer: the first convolutional operation unit → the max pooling operation unit, and its output feature is 8-dimensional, with a size of 240×240; The second layer: the first residual block → the second residual block → the third residual block, and its output feature is 16-dimensional, with a size of 240×240; The third layer: the fourth residual block → the fifth residual block → the sixth residual block → the seventh residual block, and its output feature is 32-dimensional, with a size of 120×120; The fourth layer: the eighth residual block → the ninth residual block → the tenth residual block → the eleventh residual block, and its output feature is 64-dimensional, with a size of 60×60; The fifth layer: the first upsampling module → the twelfth residual block → the thirteenth residual block → the fourteenth residual block → the fifteenth residual block, and its output feature is 32-dimensional, with a size of 120×120; The sixth layer: the second upsampling module → the sixteenth residual block → the seventeenth residual block → the eighteenth residual block, and its output feature is 16-dimensional, with a size of 240×240; The seventh layer: the second convolutional operation unit, and its output feature is 4-dimensional, with a size of 240×240; The eighth layer: the third convolutional operation unit, and its output feature is 1-dimensional, with a size of 240×240; The kernel sizes of the first convolutional operation unit and the second convolutional operation unit are both 3*3, and the strides are both 1; the kernel size of the third convolutional operation unit is 1*1, and the stride is 1; The kernel size of the max pooling operation unit is 3*3, and the stride is 2; For the two upsampling modules, the size of their convolution kernels is 4*4 and the stride is 2 for both; For the eighteen residual blocks, the size of their convolution kernels is 3*3 and the stride is 1 for all; The fourth residual block and the eighth residual block both contain two convolutional layers and one average pooling layer; the remaining residual blocks are each composed of two convolutional layers.
6. The method according to claim 1, characterized in that: In step (2d), the EDIG module and the AFF module are embedded into the backbone network of the eight-layer encoder-decoder structure, and the implementation is as follows: The shallow channel attention sub-module of the EDIG1 module is connected to the third residual block of the second layer in the backbone network, the deep channel attention sub-module of the EDIG1 module is connected to the second upsampling module of the sixth layer in the backbone network, and the output port of the EDIG1 module is connected to the sixteenth residual block of the sixth layer in the backbone network; The non-local attention module in the AFF module is connected to the seventh residual block of the third layer in the backbone network, the non-local feature fusion module in the AFF module is connected to the third residual block of the second layer in the backbone network, and the output port of the non-local feature fusion module is respectively connected to the eighth residual block of the fourth layer in the backbone network and the shallow channel attention sub-module of the EDIG2 module; The deep channel attention sub-module of the EDIG2 module is connected to the first upsampling module of the fifth layer in the backbone network, and the output port of the EDIG2 module is connected to the twelfth residual block of the fifth layer in the backbone network.
7. The method according to claim 1, characterized in that: In step (3), the multi-dimensional attention perception network is trained by the gradient descent method using the training set and its annotation information, and the implementation is as follows: (3a) The infrared small target training set and its corresponding annotation data are evenly divided into multiple paired image groups according to the batch size, and the first image group is input into the multi-dimensional attention perception network MDA-Net to obtain the weights and biases of each convolutional operation of the network and the result predicted by the network; (3b) According to the result predicted by the network and the annotation data of the infrared image, calculate its loss value through the loss function of the network: Among them, p i,j represents the prediction result of the network at the i-th row and j-th column, and t i,j represents the annotation data of the infrared image at the i-th row and j-th column. p represents the network prediction results of a set of images, and t represents the annotation data of a set of images. L IoU (p, t) represents the network loss value under this set of data of the network prediction result p and the annotation data t; (3c) Use the NesterovAcceleratedGradient algorithm to update the gradient direction, with the goal of minimizing the loss function value, update the parameters in the network to obtain the multi-dimensional attention perception network after one parameter update; (3d) Input the second image group into the multi-dimensional attention perception network after one parameter update, and repeat steps (3a) to (3c) to obtain the multi-dimensional attention perception network after two parameter updates; and so on, until the last image group is input into the multi-dimensional attention perception network updated in the previous time to obtain the multi-dimensional attention perception network after one training; (3e) Then input all the image groups into the multi-dimensional attention perception network after one training in turn, and repeat steps (3a) to (3d) to obtain the multi-dimensional attention perception network after two trainings; and so on, until all the image groups have been input 1200 times to obtain the trained multi-dimensional attention perception network.
Citation Information
Patent Citations
Lightweight small target detection method in combination with attention mechanism
CN113065558A
Infrared weak and small target detection method based on attention mechanism convolutional neural network
CN114863097A