Small target detection method capable of explaining infrared image

By constructing an infrared small object detection network based on low-rank sparse decomposition and robust principal component analysis, the problems of detection failure and missed detection and false detection under complex background noise are solved, and the real-time and lightweight deployment of the model is achieved.

CN120070867APending Publication Date: 2025-05-30JILIN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510224771.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing infrared small object detection models are prone to detection failure, missed detection and missed detection under complex background noise, and there are problems of insufficient real-time and lightweighting.

Method used

The low-rank sparse decomposition principle and robust principal component analysis optimization method are used to construct the object detection network. Through multiple stages of low-rank separation modules, sparse target extraction modules and joint reconstruction modules, the low-rank state space model and the sparse state space model are combined for global constraints, and a channel screening module is designed to improve the accuracy of sparse target extraction.

Benefits of technology

It improves the efficiency and stability of the model in complex visual dynamic systems, effectively overcomes detection errors, and significantly reduces model parameters, improving lightweight deployment and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070867A_ABST
    Figure CN120070867A_ABST
Patent Text Reader

Abstract

The invention discloses an explainable infrared image small target detection method, belongs to the technical field of computer vision, and aims to improve the real-time performance and light weight of an existing model while the problem that detection fails easily occurs under complex background noise in the prior art. Each of the training set and the test set comprises an original infrared image and an infrared small target mask corresponding to the original infrared image; decomposing into a low-rank background and a sparse target; updating the low-rank background; updating the sparse target; performing image reconstruction based on the image reconstruction network model to obtain a reconstructed original infrared image; repeating the operation for multiple times, and constructing an infrared small target detection network model; designing a loss function, and training the infrared small target detection network model by using the loss function; and for the original infrared image in the test set and the infrared small target mask corresponding to the original infrared image, based on the trained infrared small target detection network model, executing the above operation, and obtaining a finally detected small target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to an interpretable infrared image small target detection method. Background Art

[0002] Infrared small target detection (ISTD) is crucial in various fields, including modern military defense, long-range surveillance, and aerospace. In a thermally complex environment characterized by sea-sky background, ground clutter, or atmospheric disturbances, the accurate recognition and localization of remote targets with low signal-to-noise ratio or minimum contrast are of great importance. These capabilities are crucial for the effectiveness of early warning systems, precision guidance, and enhanced situation awareness. However, the inherent characteristics of infrared images, such as limited spatial resolution, thermal noise interference, and subtle thermal differences between targets and their backgrounds, pose significant technical challenges for ISTD operations.

[0003] Early ISTD algorithms were directly derived from traditional target detection techniques. However, due to the significant differences in shape and texture between visible light images and infrared images, the performance of these target detection-based ISTD algorithms is poor. Therefore, recent research has defined ISTD as a segmentation problem, which can better capture the salient features of small targets. However, using ImageNet-based segmentation networks for ISTD may limit the effective representation of infrared data. Subsequent research, especially those using CNN-based algorithms tailored for the ISTD dataset, has significantly improved the detection performance. However, these methods often emphasize local features at the expense of global context, resulting in missed detections. To address this issue, some have proposed using vision transformers (ViT) in ISTD, leveraging their ability to model long-range dependencies. However, ViT-based models greatly increase the computational requirements. Considering the need to process high-resolution images, state space models (SSMs), such as Mamba, have received attention for maintaining linear complexity while providing robust performance. However, when directly applying Mamba to the ISTD task, the specific image modeling methods of these SSMs still need to be further improved. Although these networks have improved in terms of detection accuracy, mainstream ISTD networks still lack interpretability regarding the specific principles of the ISTD domain.

[0004] With the development of deep unfolding networks (DUN), they have been introduced into the ISTD field as a method with strong interpretability and efficient network architecture. DUN unfolds the iterative solution algorithm of the existing model into the iterative layer of the network, thus promoting the optimization of network parameters. This realizes the systematic and precise integration between the iterative algorithm and the neural network. Specifically, the ISTD optimization scheme applies this technology to the image inpainting task, which involves complex model modeling and iterative solution steps, such as dealing with low-rank background and sparse targets, which is in line with the design concept of DUN. However, when applying deep unfolding networks (DUN) to ISTD-based research, the complex matrix operations involved pose a major challenge. These methods directly learn the parameters of the singular value threshold (SVT) and the soft threshold (ST). However, they often ignore the intrinsic correlations within the image, resulting in difficulty in selecting appropriate regularization parameters and further complicating model tuning.

[0005] In addition, existing ISTD algorithms mainly focus on constructing network models and often ignore the inherent defects of infrared images themselves. In the standard processing flow, these algorithms usually map single-channel grayscale images to a multi-dimensional feature space. However, the proportion of small pixels of targets in infrared images is extremely low, which means that only a few channels in the multi-dimensional feature space actually contain target information. Directly performing further feature extraction on these multi-dimensional feature maps will introduce a large amount of redundant calculations unrelated to the target. This not only increases the computational complexity but may also have a negative impact on the training efficiency and overall performance of the model.

[0006] To solve the above technical problems, Chinese Patent CN119399450A discloses "A method for constructing an infrared small target detection model and an infrared small target detection method", which designs a coordinated attention model and an end-to-end detection network model and uses collaborative attention. This integration enhances the ability of the model to effectively and stably simulate complex visual dynamic systems and solves the problem of easy detection failure under complex background noise. In addition, the network is also optimized for model lightweight, greatly reducing redundant model parameters and solving the problems of model lightweight deployment and real-time performance. Although this integration method can solve the problem of easy detection failure under complex background noise and achieve the purpose of model lightweight at the same time, there are still problems of easy missed detection and false detection under complex background noise, and the model obtained by this integration method still has problems of insufficient real-time performance and lightweight.

[0007] In summary, although the existing models can solve the problem of easy detection failure under complex background noise, they are still prone to missed detection and false detection problems, and the real-time performance and lightweight of the existing models are still insufficient. Summary of the Invention

[0008] The present invention solves the problem that the prior art is prone to detection failure under complex background noise, and at the same time improves the real-time performance and light weight of the existing model.

[0009] An interpretable infrared image small target detection method according to the present invention includes the following steps:

[0010] Step S1: Obtain a training set and a test set respectively. Both the training set and the test set include original infrared images and their corresponding infrared small target masks.

[0011] Step S2: Decompose the original infrared images in the training set into a low-rank background and a sparse target.

[0012] Step S3: Update the low-rank background based on the original infrared image, the low-rank background, and the sparse target using a low-rank separation network model.

[0013] Step S4: Extract the target based on the original infrared image, the low-rank background, and the sparse target using a sparse feature extraction network model, and update the sparse target.

[0014] Step S5: Reconstruct the image based on the updated low-rank background and the updated sparse target using an image reconstruction network model to obtain the reconstructed original infrared image.

[0015] Step S6: Repeat the operations of Step S3 to Step S5 multiple times, and construct an infrared small target detection network model based on the low-rank separation network model, the sparse feature extraction network model, and the image reconstruction network model.

[0016] Step S7: Design a loss function based on the updated low-rank background, the updated sparse target, the reconstructed original infrared image, and the infrared small target masks in the training set, and use the loss function to train the infrared small target detection network model.

[0017] Step S8: For the original infrared images and their corresponding infrared small target masks in the test set, perform the operations of Step S3 to Step S5 based on the trained infrared small target detection network model to obtain the finally detected small targets.

[0018] Further, in an embodiment of the present invention, in Step S3, the low-rank separation network model is:

[0019] It is divided into two paths. One path is sequentially connected in series with a Conv module, a Low-Rank SSM Blocks module, and a Conv module, and the other path is connected to the output of one path through adding residuals.

[0020] The Conv modules are all in the form of Conv+BN+ReLU.

[0021] Further, in an embodiment of the present invention, the Low-Rank SSM Blocks module is as follows:

[0022] First, extract features through multiple Conv modules, then use adaptive global average pooling to obtain the information of width and height, splice the dimensions of the width and height information, and split them into two one-dimensional quantities on average according to the channel dimension. Pass the two one-dimensional quantities through the local feature extraction module and the global feature extraction module respectively to obtain local information and global information, splice the local information and global information to obtain the fused feature, and finally input and connect it with its output by adding the residual to obtain the final output.

[0023] Further, in an embodiment of the present invention, the local feature extraction module is composed of Conv modules;

[0024] The global feature extraction module is composed of a low-rank state space model, and its low-rank system matrix is obtained through a low-rank system parameter identification module.

[0025] Further, in an embodiment of the present invention, the low-rank system parameter identification module is as follows:

[0026] First, obtain the system matrix for constructing the discrete state space model through linear transformation, and then use singular value decomposition to convert the system matrix of the discrete state space model into a low-rank system matrix;

[0027] The low-rank state space model is as follows:

[0028] Construct a discrete state space model with the low-rank system matrix as the basic parameter.

[0029] Further, in an embodiment of the present invention, in step S4, the sparse target extraction network model is as follows:

[0030] It is divided into two paths. One path is successively connected in series with Conv modules, a Channel Selection module, Sparse SSMBlocks modules, and Conv modules, and the other path is connected to the output of one path by adding the residual;

[0031] All Conv modules are in the form of Conv+BN+ReLU.

[0032] Further, in an embodiment of the present invention, the Channel Selection module is as follows:

[0033] First, use Maxpool to obtain one-dimensional channel intensity values. Obtain a difference matrix by constructing a directed difference between the one-dimensional channel intensity values and the transposed one-dimensional channel intensity values. Perform similarity measurement on the difference matrix to obtain the standard deviation and the mean respectively. Construct a Gaussian matrix using the standard deviation and the mean and normalize it to obtain the similarity weights. Extract the diagonal values of the similarity weights to generate the final weights. Finally, multiply with the input and perform an additive residual operation to obtain the final output.

[0034] Further, in an embodiment of the present invention, the Sparse SSM Blocks module is as follows:

[0035] First, extract features through multiple Conv modules, then use adaptive global average pooling to obtain the information of width and height, splice the dimensions of the width and height information, and split them into two one-dimensional quantities on average according to the channel dimension. Pass the two one-dimensional quantities through the local feature extraction module and the global feature extraction module respectively to obtain local and global information. Splice the local information and the global information to obtain the fused features. Finally, connect the input with its output through an additive residual operation to obtain the final output.

[0036] Further, in an embodiment of the present invention, the local feature extraction module is composed of Conv modules;

[0037] The global feature extraction module is composed of a sparse state space model, which obtains a sparse system matrix through a sparse system parameter identification module.

[0038] Further, in an embodiment of the present invention, the sparse system parameter identification module is as follows:

[0039] First, obtain the system matrix for constructing the discrete state space model through a linear transformation, and then use soft sparse representation to convert the system matrix of the discrete state space model into a sparse system matrix;

[0040] The sparse state space model is as follows:

[0041] Construct a discrete state space model with the sparse system matrix as the basic parameter.

[0042] The present invention solves the problem that the prior art is prone to detection failure under complex background noise, and at the same time improves the real-time performance and light weight of the existing model. The specific beneficial effects include:

[0043] 1. An interpretable small target detection method for infrared images according to the present invention constructs a target detection network based on the principle of low-rank sparse decomposition and a robust principal component analysis optimization method. The target detection network model consists of multiple stages, and each stage includes a low-rank separation module, a sparse target extraction module, and a joint reconstruction module. At the same time, in order to further meet the low-rank property of the background and the sparsity of the target, a low-rank state space model and a sparse state space model are respectively designed for global constraints. In addition, for sparse targets, since convolution has the problem of introducing too many invalid channels, the present invention designs a channel screening module to enhance the channel weights of the channels containing small targets. This integration improves the efficiency and stability of the model when simulating complex visual dynamic systems, effectively overcoming the problem of detection errors that are prone to occur in complex noise backgrounds. At the same time, the network lightweight processes the model, significantly reducing redundant model parameters and solving the challenges of lightweight deployment and real-time performance;

[0044] 2. An interpretable small target detection method for infrared images according to the present invention improves the real-time performance and lightweight of the existing model while the existing technology is prone to detection failure problems under complex background noise. To solve the above technical problems, the present invention decomposes the image into the background and the target, and trains a neural network using a deep learning optimizer to obtain a fast, real-time, and high-precision infrared image small target detection algorithm;

[0045] An interpretable small target detection method for infrared images according to the present invention is used to detect small targets in infrared images. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The above and / or additional aspects and advantages of the present invention will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0047] Figure 1 is a diagram of an interpretable small target detection model described in Embodiment 2;

[0048] Figure 2 is a comparison diagram of the detection effects of the images described in Embodiment 6. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The following will clearly and completely describe various embodiments of the present invention in conjunction with the drawings. The embodiments described by referring to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0050] Embodiment 1. An interpretable small target detection method for infrared images described in this embodiment includes the following steps:

[0051] Step S1, obtain a training set and a test set respectively. Both the training set and the test set include original infrared images and their corresponding infrared small target masks;

[0052] Step S2, decompose the original infrared images in the training set into a low-rank background and a sparse target;

[0053] Step S3, update the low-rank background based on the original infrared image, the low-rank background and the sparse target using a low-rank separation network model;

[0054] Step S4, perform target extraction on the original infrared image, the low-rank background and the sparse target using a sparse feature extraction network model, and update the sparse target;

[0055] Step S5, perform image reconstruction on the updated low-rank background and the updated sparse target using an image reconstruction network model to obtain the reconstructed original infrared image;

[0056] Step S6, repeat the operations of Step S3 to Step S5 multiple times, and construct an infrared small target detection network model based on the low-rank separation network model, the sparse feature extraction network model and the image reconstruction network model;

[0057] Step S7, design a loss function based on the updated low-rank background, the updated sparse target, the reconstructed original infrared image and the infrared small target masks in the training set, and use the loss function to train the infrared small target detection network model;

[0058] Step S8, for the original infrared images in the test set and their corresponding infrared small target masks, perform the operations of Step S3 to Step S5 based on the trained infrared small target detection network model to obtain the finally detected small targets.

[0059] Although the existing models can solve the problem of easy detection failure under complex background noise and achieve the purpose of model lightweighting, however, there are problems of easy missed detection and false detection under complex background noise, and such models still have problems of insufficient real-time performance and lightweighting.

[0060] To solve the above technical problems, this embodiment designs an interpretable infrared image small target detection method, including the following steps:

[0061] Step S1, obtain a training set and a test set respectively. Both the training set and the test set are original infrared images and their corresponding infrared small target masks;

[0062] Step S2, for the original infrared images in the training set, first use singular value decomposition to obtain a low-rank background and a sparse target;

[0063] Step S3: The original infrared image, the low-rank background, and the sparse target update the low-rank background based on the low-rank separation network model;

[0064] Step S4: The original infrared image, the low-rank background, and the sparse target perform target extraction based on the sparse feature extraction network model to update the sparse target;

[0065] Step S5: The low-rank background and the sparse target perform image reconstruction based on the joint reconstruction network model to obtain the reconstructed original infrared image;

[0066] The described image reconstruction network model includes 3 sequentially connected Conv modules, and its input is connected to its output by adding residuals.

[0067] The specific equation is defined as:

[0068] D k = D k-1 + h k (B k + T k )·δ k ;

[0069] Among them, B is the low-rank background, T is the sparse target, D is the reconstructed image, k represents the current stage, h k represents 3 sequentially connected Conv modules, and δ k is a trainable parameter with an initial value of 0.1;

[0070] Step S6: Repeat the operations of Step S3 to Step S5 3 times to construct the low-rank separation network model, the sparse feature extraction network model, and the image reconstruction network model into an infrared small target detection network model;

[0071] By repeating the operations of Step S3 to Step S5 3 times (i.e., 3 stages), an infrared small target detection network model is constructed, which can deepen the network;

[0072] The described infrared small target detection network model is an end-to-end detection network model, that is, the input and output are realized by one network. The described end-to-end detection network model consists of multiple stages, and each repetition of the operations of Step S3 to Step S5 is a stage;

[0073] Step S7: Design a loss function based on the low-rank background, the sparse target, the reconstructed original infrared image, and the infrared small target mask in the training set, and use the loss function to train the infrared small target detection network model;

[0074] The described loss function is:

[0075] Loss = λ 1 LossT + λ 2 LossB + λ3 LossD;

[0076] Among them, λ 1 , λ 2 , λ 3 are the weights corresponding to each loss. LossT directly constrains the target T, LossB indirectly constrains the target T of the current stage, and LossD indirectly affects the target T of the current stage. Therefore, λ 1 has the largest weight, followed by λ 2 , and finally λ 3 .

[0077] Step S8: For the original infrared image in the test set and its corresponding infrared small target mask, based on the trained infrared small target detection network model, perform the operations of steps S3 to S5 to obtain the finally detected small target.

[0078] Therefore, in this embodiment, by decomposing the image into the background and the target and training the neural network using a deep learning optimizer, a fast, real-time, and high-precision infrared image small target detection algorithm is obtained.

[0079] Embodiment 2: This embodiment further limits the interpretable infrared image small target detection method described in Embodiment 1. In step S3, the low-rank separation network model is as follows:

[0080] It is divided into two paths. One path is sequentially connected in series with a Conv module, a Low-Rank SSM Blocks module, and a Conv module, and the other path is connected to the output of the first path by adding the residual;

[0081] The Conv modules are all in the form of Conv+BN+ReLU.

[0082] In this embodiment, as Figure 1 shown, the low-rank separation network model includes a Conv module, a Low-Rank SSM Blocks module, and a Conv module connected in series in sequence, and its input is connected to its output by adding the residual; the Conv modules are all in the form of Conv+BN+ReLU.

[0083] The specific equation is defined as:

[0084] B k =(D k-1 -T k-1 +B k-1 )+f k (D k-1 -T k-1 +B k-1 )·ω k ;

[0085] Among them, B is the low-rank background, T is the sparse target, D is the reconstructed image, k is the current stage, k - 1 is the previous stage, and f k is a sequentially concatenated Conv module, Low-Rank SSM Blocks module, and Conv module, and ω k is a trainable parameter with an initial value of 0.1.

[0086] Embodiment 3. This embodiment further limits an interpretable infrared image small target detection method described in Embodiment 2. The Low-Rank SSM Blocks module is as follows:

[0087] First, extract features through multiple Conv modules, then use adaptive global average pooling to obtain the information of width and height, splice the dimensions of the width and height information, split it into two one-dimensional quantities on average according to the channel dimension, pass the two one-dimensional quantities through the local feature extraction module and the global feature extraction module respectively to obtain local information and global information, splice the local information and global information to obtain the fused feature, and finally input and connect it with its output by adding the residual to obtain the final output.

[0088] In this embodiment, the local feature extraction module is composed of a Conv module;

[0089] The global feature extraction module is composed of a low-rank state space model, and it obtains a low-rank system matrix through a low-rank system parameter identification module.

[0090] In this embodiment, the low-rank state space model is:

[0091] The low-rank system parameter identification module is:

[0092] First, obtain the system matrix for constructing the discrete state space model through linear transformation, and then use singular value decomposition to convert the system matrix of the discrete state space model into a low-rank system matrix;

[0093] The low-rank state space model is:

[0094] Construct a discrete state space model with the low-rank system matrix as the basic parameter.

[0095] In this embodiment, as Figure 1 shown, the specific composition of the Low-Rank SSM Blocks module is as follows;

[0096] First, the features are initially extracted by two layers of Conv modules. Then, adaptive global average pooling is used to obtain the information of width and height. The dimensions of the width and height information are concatenated and then split into two one-dimensional quantities along the channel dimension. The two one-dimensional quantities are respectively passed through the local feature extraction module and the global feature extraction module to obtain local and global information. The local feature extraction module is composed of Conv modules, and the global feature extraction module is composed of a Low-Rank State Space Model (Low-Rank SSM). The low-rank system matrix is obtained through the Low-Rank System Parameter Identification Module (LSPI). The low-rank state space model is constructed using the low-rank system matrix and the conventional discrete state space model construction method. The low-rank system parameter identification module first obtains the system matrices A, B, C, D for constructing the discrete state space model through linear transformation. Then, the singular value decomposition (SVD) is used to convert A, B, C, D into low-rank forms to obtain the low-rank system matrices LA, LB, LC, LD. The local information and the global information are concatenated to obtain the fused feature. Finally, the input is connected to its output by adding the residual to obtain the final output of the Low-Rank SSM Blocks module.

[0097] The specific equation is defined as:

[0098] x = σ(Conv2d(σ(BN(Conv2d(I))))) ;

[0099]

[0100] x c = Cat(x 1 , x 2 ) ;

[0101]

[0102] z 2 = σ(Conv1d(y 2 )) ;

[0103]

[0104] B, C = Split(Linear(σ(Conv1d(y 1 )))) ;

[0105] D = I + Δ ;

[0106] LA, LB, LC, LD = SVD(A, B, C, D) ;

[0107] x(k + 1) = LA d (k) · x(k) + LB d (k) · u(k) ;

[0108] y(k) = LC(k)·x(k) + LD(k)·u(k);

[0109]

[0110] z = Sigmoid(Linear(Cat(z 1 ,z 2 )));

[0111] O = x⊙z + I;

[0112] Where A, B, C, and D are initial system matrices, LA, LB, LC, and LD are low-rank system matrices, LA d and LB d are the low-rank state transition matrix and low-rank input matrix after discretization, x(k = 1) is the state update step, y(k) represents the final output, and ⊙ is the element-wise product.

[0113] Embodiment 4. This embodiment further limits an interpretable infrared image small target detection method described in Embodiment 1. In step S4, the sparse target extraction network model is:

[0114] Divided into two paths. One path is sequentially connected with a Conv module, a Channel Selection module, Sparse SSMBlocks module, and a Conv module, and the other path is connected to the output of one path by adding residuals;

[0115] The Conv modules are all in the form of Conv + BN + ReLU.

[0116] In this embodiment, as Figure 1 shown, the sparse target extraction network model includes a Conv module, a Channel Selection module, Sparse SSM Blocks module, and a Conv module connected in sequence, and its input is connected to its output by adding residuals; the Conv modules are all in the form of Conv + BN + ReLU.

[0117] The specific equation is defined as:

[0118] T k =(D k-1 -B k +T k-1 )-g k (D k-1 -T k-1 +B k-1 )·ε k ;

[0119] Among them, B is the low-rank background, T is the sparse target, D is the reconstructed image, k is the current stage, k - 1 is the previous stage, and g k is a sequentially concatenated Conv module, Channel Selection module, Sparse SSM Blocks module, and Conv module, and ε k is a trainable parameter with an initial value of 0.1.

[0120] Embodiment 5: This embodiment further limits the interpretable infrared image small target detection method described in Embodiment 4. The Channel Selection module is as follows:

[0121] First, use Maxpool to obtain one-dimensional channel intensity values. Obtain a difference matrix by constructing a directed difference between the one-dimensional channel intensity values and the transposed one-dimensional channel intensity values. Perform similarity measurement on the difference matrix to obtain the standard deviation and mean respectively. Construct a Gaussian matrix using the standard deviation and mean and normalize it to obtain the similarity weight. Extract the diagonal value of the similarity weight to generate the final weight. Finally, multiply it with the input and perform an additive residual operation to obtain the final output.

[0122] In this embodiment, the specific composition of the Channel Selection module is as follows;

[0123] First, use Maxpool to obtain one-dimensional channel intensity values α, and use β = α T to represent the transposed channel intensity values. Obtain a preliminary difference matrix d by constructing a directed difference between α and β. Perform similarity measurement on the matrix d to obtain the standard deviation σ and mean μ. Construct a Gaussian matrix using the standard deviation σ and mean μ and normalize it to obtain the similarity weight. Extract the diagonal value of the similarity weight to generate the final weight. Finally, multiply it with the input and perform an additive residual operation to obtain the final output.

[0124] The specific equation is defined as:

[0125] α = Sigmoid(Maxpool(I));

[0126] β = α T ;

[0127] d = αI C - I C β;

[0128]

[0129] W = Diag(Norm(N(x|μ, σ 2 )));

[0130] O = I ⊙ W + I;

[0131] Among them, I C is a second-order all-one square matrix with the number of channels as the dimension. R = 2 is to find the second minimum value, R = Xi is to find the Xi-th minimum value, Diag is to obtain the diagonal elements, and ⊙ is the element-wise product.

[0132] Embodiment 6: This embodiment further limits a method for detecting small targets in interpretable infrared images described in Embodiment 4. The Sparse SSM Blocks module is as follows:

[0133] First, extract features through multiple Conv modules, then use adaptive global average pooling to obtain the information of width and height, splice the dimensions of the width and height information, and split them into two one-dimensional quantities on average according to the channel dimension. Pass the two one-dimensional quantities through the local feature extraction module and the global feature extraction module respectively to obtain local and global information, splice the local information and the global information to obtain the fused feature, and finally input and connect the output by adding the residual to obtain the final output.

[0134] In this embodiment, the local feature extraction module is composed of Conv modules;

[0135] The global feature extraction module is composed of a sparse state space model, and it obtains a sparse system matrix through a sparse system parameter identification module.

[0136] In this embodiment, the sparse system parameter identification module is as follows:

[0137] First, obtain the system matrix for constructing the discrete state space model through a linear transformation, and then use soft sparse representation to convert the system matrix of the discrete state space model into a sparse system matrix;

[0138] The sparse state space model is as follows:

[0139] Construct a discrete state space model with the sparse system matrix as the basic parameter.

[0140] In this embodiment, as Figure 1 shown, the specific composition of the Sparse SSM Blocks module is as follows;

[0141] First, two layers of Conv modules are used to initially extract features. Then, adaptive global average pooling is utilized to obtain the information of width and height. The dimensions of the width and height information are concatenated, and then split into two one-dimensional quantities by averaging along the channel dimension. The two one-dimensional quantities are respectively passed through the local feature extraction module and the global feature extraction module to obtain local and global information. The local feature extraction module is composed of Conv modules, and the global feature extraction module is composed of a sparse state space model (Sparse SSM). A sparse system matrix is obtained through the sparse system parameter identification module (SSPI). The sparse state space model is constructed using the sparse system matrix and the construction method of the conventional discrete state space model. The sparse system parameter identification module first obtains the system matrices A, B, C, and D for constructing the discrete state space model through linear transformation. Then, soft sparse representation (SSR) is used to convert A, B, C, and D into a low-rank form to obtain the sparse system matrices SA, SB, SC, and SD. The local information and the global information are concatenated to obtain the fused feature. Finally, the input is connected to its output by adding the residual to obtain the final output of the Sparse SSM Blocks module.

[0142] The specific equations are defined as follows:

[0143] x = σ(Conv2d(σ(BN(Conv2d(I)))));

[0144]

[0145] x c = Cat(x 1 , x 2 );

[0146]

[0147] z 2 = σ(Conv1d(y 2 ));

[0148]

[0149] B, C = Split(Linear(σ(Conv1d(y 1 ))));

[0150] D = I + Δ;

[0151] SA, SB, SC, SD = SSR(A, B, C, D);

[0152] x(k + 1) = SA d (k) · x(k) + SB d (k) · u(k);

[0153] y(k) = SC(k)·x(k) + SD(k)·u(k);

[0154]

[0155] z = Sigmoid(Linear(Cat(z 1 ,z 2 ));

[0156] O = x⊙z + I;

[0157] Wherein, A, B, C, and D are initial system matrices, SA, SB, SC, and SD are sparse system matrices, SA d and SB d are the discretized sparse state transition matrix and sparse input matrix, x(k = 1) represents the state update step, y(k) represents the final output, and ⊙ represents the element-wise product.

[0158] To better illustrate an interpretable infrared image small target detection method described in Embodiments 1-6, it is described in detail through the following embodiments:

[0159] As Figure 2 shown by the yellow dashed box, the existing model is prone to missed detection problems under complex background noise, and no missed detection problem occurs when using this embodiment to detect the image. As Figure 2 shown by the blue solid box, the existing model is prone to false detection problems under complex background noise, and no false detection problem occurs when using this embodiment to detect the image.

[0160] In summary, this embodiment dynamically extracts the low-rank background and sparse features of the image through the low-rank sparse decomposition network model, effectively fuses the two parts of features, and solves the problem of the lack of interpretability of the network model. And it realizes the low-rank state space model and the sparse state space model to further strengthen the low-rankness of the background and the sparsity of the target. At the same time, it realizes the channel selection module to further highlight the channels containing sparse targets, improves the accuracy of sparse target extraction, and solves the problem of difficult separation of target and background under complex backgrounds in the infrared image small target detection technology. Finally, the network ensures lightweight and real-time performance on the premise of ensuring accuracy, and solves the real-time problem of the existing infrared image small target detection network.

[0161] At the same time, it solves the problems of easy missed detection and false detection of the existing model under complex background noise, and this model still has problems of insufficient real-time performance and lightweight.

[0162] The above has introduced in detail a method for detecting small targets in interpretable infrared images proposed by the present invention. In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for detecting small targets in interpretable infrared images, characterized in that: The following steps are involved: Step S1, respectively obtaining a training set and a test set, wherein both the training set and the test set include an original infrared image and a corresponding infrared small target mask; Step S2, decomposing the original infrared image in the training set into low-rank background and sparse targets; Step S3, the original infrared image, the low-rank background and the sparse target are used to update the low-rank background based on the low-rank separation network model; Step S4, extracting the original infrared image, the low-rank background and the sparse target based on the sparse feature extraction network model, and updating the sparse target; Step S5, reconstructing the updated low-rank background and the updated sparse target based on the image reconstruction network model to obtain a reconstructed original infrared image; Step S6, repeating the operations of steps S3 to S5 for multiple times, and constructing an infrared small target detection network model based on a low-rank separation network model, a sparse feature extraction network model, and an image reconstruction network model; Step S7, designing a loss function based on the updated low-rank background, the updated sparse target, the reconstructed original infrared image, and the infrared small target mask in the training set, and using the loss function to train the infrared small target detection network model; Step S8, for the original infrared image in the test set and the corresponding infrared small target mask, based on the trained infrared small target detection network model, perform the operations of steps S3 to S5 to obtain the final detected small target.

2. The method for detecting small targets in interpretable infrared images according to claim 1, characterized in that: In the step S3, the low-rank separation network model is: It is divided into two paths, one of which is connected in series with the Conv module, the Low-Rank SSM Blocks module and the Conv module, and the other is connected to the output of one path by adding the residual; Conv modules are all in the form of Conv+BN+ReLU.

3. The method for detecting small targets in interpretable infrared images according to claim 2, characterized in that: The Low-Rank SSM Blocks modules are: First, features are extracted through a multi-layer Conv module. Then, adaptive global average pooling is used to obtain width and height information. The dimensions of the width and height information are concatenated and split into two one-dimensional quantities according to the channel dimension. The two one-dimensional quantities are passed through a local feature extraction module and a global feature extraction module respectively to obtain local information and global information respectively. The local information and the global information are concatenated to obtain fused features. Finally, the input is connected with its output by adding the residual.

4. The method for detecting small targets in interpretable infrared images according to claim 3, characterized in that: The local feature extraction module is composed of a Conv module; The global feature extraction module is composed of a low-rank state space model, which obtains a low-rank system matrix through a low-rank system parameter identification module.

5. The method for detecting small targets in interpretable infrared images according to claim 4, characterized in that: The low-rank system parameter identification module is: Firstly, the system matrix of the discrete state space model is obtained by linear transformation, and then the system matrix of the discrete state space model is converted into a low-rank system matrix by singular value decomposition. The low-rank state space model is: A discrete state space model is constructed using a low-rank system matrix as the basic parameter.

6. The method for detecting small targets in interpretable infrared images according to claim 1, characterized in that: In the step S4, the sparse target extraction network model is: It is divided into two paths, one of which is a series of Conv module, ChannelSelection module, Sparse SSM Blocks module and Conv module, and the other is connected to the output of one path by adding the residual; Conv modules are all in the form of Conv+BN+ReLU.

7. The method for detecting small targets in interpretable infrared images according to claim 6, characterized in that: The Channel Selection module is: First, use Maxpool to obtain the one-dimensional channel intensity value, and obtain the difference matrix by constructing the directed difference between the one-dimensional channel intensity value and the transposed one-dimensional channel intensity value. Perform similarity measurement on the difference matrix to obtain the standard deviation and mean respectively. Use the standard deviation and mean to construct a Gaussian matrix and normalize it to obtain the similarity weight. Extract the diagonal value of the similarity weight to generate the final weight, and finally multiply it with the input and add the residual to obtain the final output.

8. The method for detecting small targets in interpretable infrared images according to claim 6, characterized in that: The Sparse SSM Blocks modules are: First, features are extracted through a multi-layer Conv module. Then, adaptive global average pooling is used to obtain width and height information. The dimensions of the width and height information are concatenated and split into two one-dimensional quantities according to the channel dimension. The two one-dimensional quantities are respectively passed through a local feature extraction module and a global feature extraction module to obtain local and global information. The local information and the global information are concatenated to obtain fused features. Finally, the input is connected with its output by adding the residual.

9. The method for detecting small targets in interpretable infrared images according to claim 8, characterized in that: The local feature extraction module is composed of a Conv module; The global feature extraction module is composed of a sparse state space model, which obtains a sparse system matrix through a sparse system parameter identification module.

10. The method for detecting small targets in interpretable infrared images according to claim 9, characterized in that: The sparse system parameter identification module is: Firstly, the system matrix of the discrete state space model is obtained by linear transformation, and then the system matrix of the discrete state space model is converted into a sparse system matrix by using soft sparse representation. The sparse state space model is: A discrete state space model is constructed with the sparse system matrix as the basic parameter.

Citation Information

Patent Citations

  • Construction method of infrared small target detection model and infrared small target detection method

    CN119399450A