Infrared small target detection method based on diversified feature learning and coordination
By proposing a diverse feature capture and coordination network in infrared small object detection, the problems of small target size, complex background and lack of texture features are solved in infrared small object detection, and higher detection accuracy and effect are achieved.
Patent Information
- Application Number
- CN202411912211.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-24
AI Technical Summary
The infrared small object detection task faces the problems of small target size, complex background and lack of texture features, which makes it difficult for existing detection networks to retain local details and restore subtle target features when dealing with irregular shapes and lack of textures.
A diversified feature capture and coordination network is proposed. Local and global features are extracted through FFT hybrid coding blocks, and the network is reconstructed to enhance image clarity. The cross-layer feature adaptive selection method is used to refine high-level features, and a loss function based on coordinate calibration is designed to optimize the position of the prediction target.
It improves the accuracy and effectiveness of infrared small object detection, especially when dealing with irregular shapes and lack of texture, it can better retain local details and restore subtle object characteristics, improving detection accuracy.
Smart Images

Figure CN119942297A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of infrared target detection, and relates to research on a method for detecting small infrared targets by coordinating multiple features. Background Art
[0002] Infrared images have strong detection capabilities and high sensitivity to targets under low light conditions, so they are widely used in military, agricultural and civilian fields. The goal of the infrared small target detection (IRSTD) task is to extract small targets from infrared images in complex backgrounds. However, due to the long imaging distance and poor imaging conditions, IRSTD faces many challenges:
[0003] (1) Small target size: Small targets occupy very few pixels in the image, resulting in weak signal strength and easy to be drowned out by noise. Typically, the target only occupies about 0.05% of the total pixels in an infrared image.
[0004] (2) Complex background: A complex background environment contains many interference factors, such as temperature changes, environmental noise, and other heat sources that produce infrared radiation similar to the target signal. The complex background makes it difficult to separate the target signal from the noise background.
[0005] (3) Lack of texture features: Infrared imaging mainly relies on thermal radiation, and infrared targets are usually small and lack detailed features such as regular shapes and textures that can be seen under visible light.
[0006] Existing IRSTD methods can be roughly divided into traditional methods and deep learning-based methods. Traditional methods include filtering-based methods, methods based on human visual system models, and low-rank-based methods. These traditional methods rely heavily on prior knowledge, empirical assumptions, and manually adjusted hyperparameters, which limits their effectiveness in dealing with complex real-world scenarios.
[0007] In recent years, deep learning methods have shown good performance in IRSTD tasks. For example, Liu et al. proposed the first IRSTD method based on convolutional neural network (CNN) using multi-layer perceptron (MLP) for detection. In addition, generative adversarial network (GAN) has also been applied to this task, such as MDvsFA-cGAN, using adversarial training to balance false negatives and false positives, thereby improving detection performance. Recent studies have shown that treating the IRSTD task as a semantic segmentation problem is usually more effective than traditional detection methods, and can more accurately identify and locate small targets. For example, ACMNet is the first IRSTD method based on semantic segmentation, which introduces an asymmetric context module to replace the skip connection, realizes cross-layer feature fusion, and improves detection performance. DNANet adopts a densely nested interaction method to promote feature interaction between high-level and low-level features. ISNet supplements the precise shape information of infrared targets by integrating Taylor finite difference heuristic blocks and bidirectional attention aggregation modules. FC3-Net explores cross-level correlation and feature compensation to cope with the loss of details caused by downsampling layers. DMFNet designs a dual encoder to obtain more effective information, thereby enhancing the detection ability of small targets.
[0008] Despite remarkable progress in the field of IRSTD, the inherent characteristics of infrared images and targets still pose significant challenges. Specifically, targets in infrared images often lack well-defined shapes and textures, making it difficult to effectively extract multi-category target information. This limitation directly affects the detection and localization accuracy of small irregularly shaped targets. Currently, most detection networks, such as U-Net-based models, usually reduce the spatial resolution by downsampling. While this helps extract deeper features, it also leads to the loss of local contextual information, preventing the network from recovering and exploiting subtle target details. While existing network architectures are capable of handling regular targets, they have difficulty coping with the problem of detail loss in small target detection, especially when dealing with irregularly shaped and texture-poor targets. Therefore, advancing network architectures and feature learning strategies, especially in preserving local details and recovering subtle target features, has become the focus of current research. Inspired by the core idea of preserving and identifying fine target feature reconstruction, we propose an innovative infrared image reconstruction mechanism to capture and preserve important information from the original image. As a result, we apply the infrared image reconstruction mechanism to the IRSTD task for the first time. Specifically, we propose a new network, Diversified Feature Capture and Coordination Network, to recover and optimize target information to improve the effect of infrared small target detection. Summary of the invention
[0009] The object of the present invention is to provide a detection method for the field of infrared small target detection. With the application of deep learning in the field of infrared small target detection, significant progress has been made in detection accuracy. Despite the remarkable progress in the field of IRSTD, the inherent characteristics of infrared images and targets still pose significant challenges. Specifically, targets in infrared images often lack clear shapes and textures, making it difficult to effectively extract multi-category target information. This limitation directly affects the detection and positioning accuracy of small irregular shaped targets. At present, most detection networks, such as U-Net-based models, usually reduce spatial resolution by downsampling. Although this helps to extract deeper features, it also leads to the loss of local contextual information, which prevents the network from recovering and utilizing subtle target details. Therefore, the present invention proposes a diversified feature capture and coordination network for recovering and optimizing target information to improve the effect of infrared small target detection. Its core lies in the following four parts:
[0010] 1. FFT hybrid coding block
[0011] Infrared images usually have problems with low contrast and complex background, which may make it difficult for traditional convolution methods to extract subtle target features. Convolutional neural networks (CNNs) may also be affected by background noise, which affects the accuracy and effectiveness of infrared target detection (IRSTD). Spectral convolution can significantly improve feature extraction by converting the convolution operation into a frequency domain operation using Fourier transform. Fourier transform provides additional frequency information, which helps to identify targets in low-contrast areas, and can also process background information in the frequency domain, thereby separating the target from the background and improving detection accuracy. To improve computational efficiency, we use fast Fourier transform (FFT).
[0012] We designed a FFT-based hybrid coding block to effectively extract local and global features. The specific architecture of the coding block is as follows: Figure 1 As shown in (a) and (b). We transform the input feature X into the frequency domain through FFT and obtain the real part Re(X) and the imaginary part Im(X). By concatenating these two parts in the channel dimension, we obtain a more comprehensive frequency domain feature representation. Then, we use a convolution block to learn from these features, and finally convert the data back to the spatial domain through the inverse Fourier transform (IFFT) to obtain the feature X′. In addition, the residual connection is applied to the input feature X to obtain the feature X f , further enhancing the feature representation ability of the model.
[0013] Re(X),Im(X)=FFT(X)
[0014] X′=IFFT(Conv(Cat[Re(X),Im(X)]))
[0015]
[0016] where Gat[·,·] represents the concatenation operation along the channel dimension. represents element-wise addition. σ represents the Sigmoid function. Conv(·) represents a convolutional block.
[0017] The obtained features are further processed by convolution blocks to extract local information and provide a more complete feature representation F. Specifically, by capturing local features using convolution operations in the spatial domain and capturing global features through Fourier transforms in the frequency domain, this fusion method is able to capture both local details and global properties of the image. This complementary approach enhances the detection of small objects and improves the overall parsing of complex scenes.
[0018] 2. Feature enhancement brought by reconstruction network
[0019] Since infrared images have the characteristics of low contrast, noise and blur, it becomes difficult to identify and detect small targets in the image. Therefore, infrared images need to be enhanced to improve the visibility of targets and the clarity of images. Therefore, we introduce a reconstruction network branch so that the model can handle image reconstruction and segmentation tasks at the same time. By establishing a joint learning framework for reconstruction and segmentation, it helps the model to communicate features and learn more useful representations for target detection.
[0020] The purpose of image reconstruction is to obtain higher quality, clearer, more compact and useful images by processing, repairing and improving existing image data. After introducing the reconstruction network branch, the infrared image can be enhanced, thereby improving the visibility of the target and the clarity of the image, making it easier to identify and detect small targets. In addition, the introduction of the reconstruction network branch can provide a richer and more accurate image representation for the target detection branch. By learning and restoring the global structure and context information of the image, the reconstruction network can help the detection network better understand the location, shape and characteristics of the target in the image, thereby improving the detection accuracy.
[0021] Specifically, if Figure 2 As shown in Figure 1, we use UNet based on ResNet18 as the backbone of the reconstruction network and use CSAM (Spatial Channel Attention Module) between each convolutional block to integrate and enhance the obtained features. The original infrared image is preprocessed and used as the input of the reconstruction network. The feature L in the encoding process i_0 It can be expressed as:
[0022] L i_0 =ρ max (CSAM(Conv(L i-1_0 )))
[0023] Among them, ρ max (·) denotes a max pooling operation with a stride of 2. CSAM denotes a spatial channel attention module. Conv(·) denotes a convolutional block in the backbone network.
[0024] In the decoding stage, the deep features are upsampled and the features of the encoding stage are connected through ordinary skip connections as the input of the decoding block. The features of the decoding stage can be expressed as:
[0025] L i_1 =CSAM(Conv(Cat[L i_0 ,μ(L i+1_1 )]))
[0026] Among them, Cat[·,·] represents the concatenation operation along the channel dimension, μ(·) represents the upsampling operation with a sampling ratio of 2.
[0027] For the detection network, the encoding part includes a conventional convolution block and an FFT mixing block. The output of the FFT mixing block is represented as F i Since the output L from the corresponding stage of the reconstruction network i_0 It provides additional information for the detection network, so the final output feature D of each stage i_0 It can be expressed as:
[0028] D i_0 =ρ max (Cono(Cat[D i-1_0 , L i_0 , F i ]))
[0029] This approach combines convolutional features, frequency domain characteristics, and reconstruction information to enhance the encoder's ability to extract features from the input image.
[0030] 3. Cross-layer feature adaptive selection method
[0031] In the encoding stage, low-level features usually contain rich object details that are not obvious in high-level features. To address this problem, we propose a Cross-layer Feature Adaptive Selection (CFAS) method to refine high-level features by adaptively selecting the fusion weights of deep and shallow features to facilitate shape and edge detection.
[0032] like Figure 3 As shown, for the module input, the low-level features First, a 1×1 convolution operation is performed to adjust the channel dimension so that it is consistent with the high-level features. Then, the adjusted low-level features X l With high-level features X hAdd them element by element to get the fused feature, which is used as the input of the channel attention (ca) module. The generated feature X c Then, the channel attention weights f of high-level and low-level features are obtained through the sigmoid function. ca and 1-f ca , corresponding to high-level and low-level features respectively.
[0033] f ca =Sigmoid(ca(X))
[0034] The output of the channel attention module X c is input into the spatial attention (sa) module, and the result X s After sigmoid mapping, the weights f of high-level and low-level features in the spatial dimension are obtained. sa and 1-f sa :
[0035] f sa =Sigmoid(sa(X c ))
[0036] Finally, use the obtained weights to adjust the low-level features X l and high-level features X h The weighted results are weighted, and the residual operation is performed on the weighted results. Finally, the results are added in the channel dimension to obtain the final enhanced feature X′:
[0037]
[0038] in, and Represent element-wise addition and element-wise multiplication, respectively.
[0039] Most existing feature fusion methods cannot flexibly adjust the fusion weights of different features, and usually only consider a single channel and spatial dimension. Compared with these fusion methods, our method has several significant advantages. First, the weight of each feature is flexibly selected according to the characteristics of the feature itself. Second, the spatial and channel information of the feature is fully considered. Finally, by adding the residual part, the robustness and generalization ability of the fusion are further enhanced.
[0040] 4. Loss Function
[0041] The loss function is an important component in model training, which is used to compare the difference between the real target and the predicted result. In the infrared target detection (IRSTD) task, the IoU loss is usually used. However, the IoU loss focuses more on the difference in shape and area of the predicted target, and is insensitive to the accuracy of the target position. Therefore, we introduce the concept of positioning calibration detection and design a loss function based on coordinate calibration. In order to address the insensitivity of the detection network to the position of small targets, we design a target-level evaluation metric for the IRSTD task: the coordinate calibration (CC) loss, and propose a cooperative two-stage training strategy aimed at calibrating and optimizing the position of the predicted target. Specifically, as Figure 4 As shown in Figure 1, we first obtain the centroid coordinates of the real target and the predicted target. Since the infrared image may contain multiple small infrared targets, when there are multiple targets, we calculate the average of the centroid coordinates of these targets. Then we calculate the position information difference between the real target and the predicted target and use it to further optimize the model.
[0042] First, the IoU loss is used for the IRSTD task. It is a pixel-level evaluation metric that evaluates the detector's ability to describe the object contour. IoU is calculated by the area ratio of the intersection and union between the prediction and the label.
[0043]
[0044] in, and A represent the predicted and true target maps, respectively. and are the areas of intersection and union respectively.
[0045] The MSE loss is used for infrared image reconstruction tasks to measure the pixel-level difference between the model-reconstructed image and the original infrared image, specifically the average of the squared differences between the predicted value and the true value.
[0046]
[0047] in, and I i are the values of the i-th pixel in the reconstructed image and the original image, respectively, and M is the total number of pixels in the image. By minimizing the MSE loss, the reconstruction network can learn to generate results that are highly similar to the real image. Since the decoder is encouraged to reconstruct the entire original image, the encoder must learn to retain rich target-related information, which is easily lost in the traditional encoder-decoder framework. Through this mechanism, the encoder can make up for the lost target-related information.
[0048]
[0049] Among them, λ is used to adjust CCloss The proportion of the total loss, m is the number of targets (when there are multiple targets). is the center of mass coordinate of the real target.
[0050] The specific implementation of the two-stage training strategy is as follows:
[0051]
[0052] Among them, n represents the current training round, and N is the round threshold for switching the loss function. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 Schematic diagram of the infrared small target detection model of the invented diversified feature capture and coordination network;
[0054] Figure 2 A schematic diagram of the infrared image reconstruction branch structure in the invented model;
[0055] Figure 3 Schematic diagram of the cross-layer feature adaptive selection method;
[0056] Figure 4 This is a schematic diagram of coordinate calibration loss;
[0057] Figure 5 Visualization results of the method of the present invention and the comparative method in various typical scenes selected in the data set;
[0058] Figure 6 The three-dimensional visualization results of various typical scenes selected in the data set by the method of the present invention and the comparative method; DETAILED DESCRIPTION
[0059] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0060] Step 1: Preprocess the data used;
[0061] Step 2: Input the preprocessed data into the target detection branch and reconstruction branch for feature extraction;
[0062] Step 3: Input the features of the target detection branch into the cross-layer feature adaptive selection method module;
[0063] Step 4: Obtain the result of the reconstruction branch and fuse the outputs of different scales to obtain the final prediction map;
[0064] Step 5: Use the loss function to calculate the loss;
[0065] Further, the step 1 specifically comprises the following steps:
[0066] Step 1-1: The image sizes in the infrared small target dataset used are different. To ensure the consistency of model input, improve computing efficiency and memory utilization, simplify data expansion and other operations, and also facilitate the training of the model, we uniformly crop and scale the infrared images into images of size 256×256, and set the value of the expanded area to 0.
[0067] Step 1-2: Then normalize the image and perform data enhancement operations, including random image flipping, blurring, and other operations.
[0068] Further, the step 2 specifically includes the following steps:
[0069] Step 2-1: If Figure 1 The figure shows a schematic diagram of the infrared small target detection model of the invention's diversified feature capture and coordination network, wherein the implementation process of the infrared image reconstruction branch structure is described in detail as follows Figure 2 Before the infrared image is input into the image reconstruction branch, the image is first processed by adding Gaussian noise of random size, and the processed image is used as the input of the reconstruction branch, while the original image is used as the input of the target detection branch.
[0070] Step 2-2: We designed a FFT-based hybrid coding block to effectively extract local and global features. The specific architecture of the coding block is as follows: Figure 1 As shown in (a) and (b). We transform the input feature X into the frequency domain through FFT and obtain the real part Re(X) and the imaginary part Im(X). By concatenating these two parts in the channel dimension, we obtain a more comprehensive frequency domain feature representation. Then, we use a convolution block to learn from these features, and finally convert the data back to the spatial domain through the inverse Fourier transform (IFFT) to obtain the feature X′. In addition, the residual connection is applied to the input feature X to obtain the feature X f , further enhancing the feature representation ability of the model.
[0071] Re(X),Im(X)=FFT(X)
[0072] X′=IFFT(Cone(Cat[Re(X),Im(X)]))
[0073]
[0074] where Cat[·,·] represents the concatenation operation along the channel dimension. represents element-wise addition. σ represents the Sigmoid function. Conv(·) represents a convolutional block.
[0075] Step 2-3: By learning and restoring the global structure and contextual information of the image, the reconstruction network can help the detection network better understand the location, shape, and features of the target in the image, thereby improving the detection accuracy.
[0076] Specifically, if Figure 2 As shown in Figure 1, we use UNet based on ResNet18 as the backbone of the reconstruction network and use CSAM (Spatial Channel Attention Module) between each convolutional block to integrate and enhance the obtained features. The original infrared image is preprocessed and used as the input of the reconstruction network. The feature L in the encoding process i_0 It can be expressed as:
[0077] L i_0 =ρ max (CSAM(Conv(L i-1_0 )))
[0078] Among them, ρ max (·) denotes a max pooling operation with a stride of 2. CSAM denotes a spatial channel attention module. Conv(·) denotes a convolutional block in the backbone network.
[0079] In the decoding stage, the deep features are upsampled and the features of the encoding stage are connected through ordinary skip connections as the input of the decoding block. The features of the decoding stage can be expressed as:
[0080] L i_1 =CSAM(Conv(Cat[L i_0 ,μ(L i+1_1 )]))
[0081] Among them, Cat[·,·] represents the concatenation operation along the channel dimension, μ(·) represents the upsampling operation with a sampling ratio of 2.
[0082] For the detection network, the encoding part includes a conventional convolution block and an FFT mixing block. The output of the FFT mixing block is represented as F i Since the output L from the corresponding stage of the reconstruction network i_0 It provides additional information for the detection network, so the final output feature D of each stage i_0 It can be expressed as:
[0083] D i_0 =ρ max (Conv(Cat[D i-1_0 , L i_0 , F i ]))
[0084] Further, the step 3 specifically includes the following steps:
[0085] Step 3-1: If Figure 3 As shown, for the module input, the low-level features First, a 1×1 convolution operation is performed to adjust the channel dimension so that it is consistent with the high-level features. Then, the adjusted low-level features X l With high-level features X h Add them element by element to get the fused feature, which is used as the input of the channel attention (ca) module. The generated feature X c Then, the channel attention weights f of high-level and low-level features are obtained through the sigmoid function. ca and 1-f ca , corresponding to high-level and low-level features respectively.
[0086] f ca =Sigmoid(ca(X))
[0087] Step 3-2: Output X of the channel attention module c is input into the spatial attention (sa) module, and the result X s After sigmoid mapping, the weights f of high-level and low-level features in the spatial dimension are obtained. sa and 1-f sa :
[0088] f sa =Sigmoid(sa(X c ))
[0089] Step 3-3: Finally, use the obtained weights to analyze the low-level features X l and high-level features X h The weighted results are weighted, and the residual operation is performed on the weighted results. Finally, the results are added in the channel dimension to obtain the final enhanced feature X′:
[0090]
[0091] in, and Represent element-wise addition and element-wise multiplication, respectively.
[0092] Further, the step 4 specifically comprises the following steps:
[0093] The features obtained in the above process are upsampled to the same size as the original image, and the obtained feature map is transformed into a feature map with 3 channels using a 1×1 convolution layer. The global robust feature map G is calculated by the following formula:
[0094]
[0095] Finally, the obtained robust feature map is transformed into a final prediction map.
[0096] Further, the step 5 specifically includes the following steps:
[0097] like Figure 4 As shown in the figure, we first obtain the centroid coordinates of the real target and the predicted target. Since the infrared image may contain multiple small infrared targets, when there are multiple targets, we calculate the average of the centroid coordinates of these targets. Then the position information difference between the real target and the predicted target is calculated and used to further optimize the model. The intersection-over-union loss is used to predict the area difference between the real target and the predicted target. The specific implementation of the two-stage training strategy is as follows:
[0098]
[0099] Among them, n represents the current training round, and N is the round threshold for switching the loss function.
[0100] We use the publicly available NUAA, NUDT, and IRSTD-1k datasets from ACM Net, DNANet, and ISNet methods to verify the state-of-the-art and effectiveness of our method. The NUAA dataset contains 427 images from hundreds of real natural scenes, which is one of the most popular datasets for single-frame IRSTD. The NUDT dataset contains 1,327 images of more challenging scenes, such as multi-target, point target, and dark target scenes. The IRSTD-1k dataset contains 1,000 infrared images taken from the real world, with targets of different types and sizes, complex scenes, and severe clutter and noise. The selected datasets ensure that the datasets are diverse and cover different situations and challenges that the model may face, which can increase the evaluation of the model's generalization ability.
[0101] Top-hat and Max-median are selected in the model-driven method, and ACMNet, AGPCNet, DNANet, DCFRNet and the method of the present invention are selected in the data-driven method for comparison. The visualization result diagram and the three-dimensional visualization result diagram in the typical scene are shown in Figure 2. Figure 4 , Figure 5 As shown in Table 1. The present invention uses mIoU (Intersection over Union) to quantitatively evaluate and analyze the method of the present invention and several other infrared small target detection methods. The mIoU value range is 0 to 1. The closer the value is to 1, the better the performance of the detection method. The results are shown in Table 1. From the experimental results, the detection effect of the method of the present invention is obviously better than that of the comparison algorithm, thereby verifying the superiority of the method of the present invention.
[0102] Table 1 mIoU comparison table of infrared small target detection
[0103]
[0104] The results are shown in Table 1. From the experimental results, the detection effect of the method of the present invention is obviously better than that of the comparison algorithm, thus verifying the superiority of the method of the present invention.
Claims
1. A method for infrared small target detection based on diversified feature learning and coordination, characterized by: The following steps are involved: S1. Construct a diversified feature capture and coordination network, which learns and coordinates diversified features through multi-path coding. Construct a global feature extraction branch composed of FFT hybrid coding blocks to capture macro-level features of infrared images and provide extensive contextual information for target detection; S2, an infrared image reconstruction branch running in parallel with the detection branch, which preserves small target information and reduces feature loss through complementary context encoding; S3, cross-layer feature autonomous selection method: cross-layer feature fusion is performed in the decoding stage. This method effectively integrates and coordinates deep semantic information and shallow detail features; S4. In order to improve the accuracy of target positioning and detection, we introduced the concept of positioning calibration detection and proposed a new coordinate calibration (CC) loss function that focuses on optimizing the target position accuracy. At the same time, we adopted a matching two-stage training strategy to accurately capture and correct the target position.
2. The method for detecting small infrared targets by coordinating multiple features according to claim 1, characterized in that: The step S1 comprises the following steps: S1.1 We design a hybrid coding block based on FFT to effectively extract local and global features. We transform the input feature X into the frequency domain through FFT and obtain the real part Re(X) and the imaginary part Im(X). By concatenating these two parts in the channel dimension, we obtain a more comprehensive frequency domain feature representation. Re(X),Im(X)=FFT(X) S1.2 We then use a convolutional block to learn from these features and finally transform the data back to the spatial domain via an inverse Fourier transform (IFFT) to obtain the features X′. X′=IFFT(Conυ(Cat[Re(X),Im(X)])) S1.3 In addition, the residual is connected to the input feature X to obtain feature X f , further enhancing the feature representation ability of the model. where Cat[·,·] represents the concatenation operation along the channel dimension. represents element-wise addition. σ represents the Sigmoid function. Conυ(·) represents a convolutional block.
3. The infrared small target detection method based on a dual encoder multi-stage feature fusion network according to claim 1 is characterized in that: The step S2 comprises the following steps: S2.1 By learning and restoring the global structure and contextual information of the image, the reconstruction network can help the detection network better understand the position, shape, and features of the target in the image, thereby improving the detection accuracy. Specifically, we use UNet based on ResNet18 as the backbone of the reconstruction network and use CSAM (Spatial Channel Attention Module) between each convolutional block to integrate and enhance the obtained features. The original infrared image is preprocessed and used as the input of the reconstruction network. The feature L in the encoding process i_0 It can be expressed as: 50 i_0 =ρ max (CSAM(Conυ(L i-1_0 ))) Among them, ρ max (·) denotes a max pooling operation with a stride of 2. CSAM denotes a spatial channel attention module. Conυ(·) denotes a convolutional block in the backbone network. S2.2 In the decoding stage, the deep features are upsampled and the features of the encoding stage are connected through ordinary skip connections as the input of the decoding block. The features of the decoding stage can be expressed as: L i_1 =CSAM(Conυ(Cat[L i_0 ,μ(L i+1_1 )])) Among them, Cat[·,·] represents the concatenation operation along the channel dimension, μ(·) represents the upsampling operation with a sampling ratio of 2. S2.3 For the detection network, the encoding part includes a conventional convolution block and an FFT mixing block. The output of the FFT mixing block is represented as F i Since the output L from the corresponding stage of the reconstruction network i_0 It provides additional information for the detection network, so the final output feature D of each stage i_0 It can be expressed as: D i_0 =ρ max (Conυ(Cot[D i-1_0 ,L i_0 ,F i ]))。 4. The infrared small target detection method based on a dual encoder multi-stage feature fusion network according to claim 1 is characterized in that: The step S3 comprises the following steps: S3.1 For module input, low-level features First, a 1×1 convolution operation is performed to adjust the channel dimension so that it is consistent with the high-level features. match. S3.2 Then, the adjusted low-level features X l With high-level features X h Add them element by element to get the fused feature, which is used as the input of the channel attention (ca) module. The generated feature X c Then, the channel attention weights f of high-level and low-level features are obtained through the sigmoid function. ca and 1-f ca , corresponding to high-level and low-level features respectively. f ea =Sigmoid(ca(X)) S3.3 Output X of the channel attention module c is input into the spatial attention (sa) module, and the result X s After sigmoid mapping, the weights f of high-level and low-level features in the spatial dimension are obtained. sa and 1-f sa : f sa =Sigmoid(sa(X c )) S3.4Finally, use the obtained weights to adjust the low-level features X l and high-level features X h The weighted results are weighted, and the residual operation is performed on the weighted results. Finally, the results are added in the channel dimension to obtain the final enhanced feature X′: in, and Represent element-wise addition and element-wise multiplication, respectively.
5. The infrared small target detection method based on a dual encoder multi-stage feature fusion network according to claim 1 is characterized in that: The step S4 is as follows, We first obtain the centroid coordinates of the real target and the predicted target. Since the infrared image may contain multiple small infrared targets, when there are multiple targets, we calculate the average of the centroid coordinates of these targets. Then the position information difference between the real target and the predicted target is calculated and used to further optimize the model. The intersection-over-union loss is used to predict the area difference between the real target and the predicted target. The specific implementation of the two-stage training strategy is as follows: Among them, n represents the current training round, and N is the round threshold for switching the loss function. The loss function of the two-stage training strategy is used to calculate the detection branch loss and the image reconstruction branch loss for model training.
Citation Information
Patent Citations
Design method of interpretable multi-scale infrared weak and small target detection network
CN114998566A
Infrared small target detection method based on learnable guide filtering
CN118521767A
Infrared weak and small target detection method based on wavelet guidance
CN118587507A