Infrared weak and small target detection method based on non-convex weighted tensor rank minimization and adaptive space-time modeling
Through adaptive spatiotemporal modeling and non-convex weighted tensor rank minimization, combined with Laplace norm and SCAD penalty terms, the problems of insufficient generalization performance and high computational complexity in infrared small target detection are solved, and more efficient infrared dim small target detection is achieved.
Patent Information
- Application Number
- CN202510566981.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-12
AI Technical Summary
Existing infrared small target detection methods have problems such as insufficient generalization performance, dependence on dataset quality and quantity, high computational complexity, improper handling of spatiotemporal correlations, and biased sparse target estimation when dealing with complex backgrounds and noisy environments.
An infrared dim target detection method based on non-convex weighted tensor rank minimization and adaptive spatiotemporal modeling is adopted. The time step is adaptively set by tensor information entropy. Combined with the Laplace norm and weighted average tensor rank, the smooth truncated absolute deviation penalty term is used to optimize the solution process to improve the reconstruction accuracy of low-rank background and sparse targets.
The accuracy and robustness of infrared dim target detection are improved, the dependence on expert experience is reduced, the computational complexity is reduced, and the ability to suppress complex backgrounds and noise is enhanced.
Smart Images

Figure CN120635402A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to an infrared dim small target detection method based on non-convex weighted tensor rank minimization and adaptive spatiotemporal modeling. Background Art
[0002] Infrared small target detection is a highly complex and crucial technology, crucial for a wide range of military and civilian applications, such as mine detection, nighttime navigation, precision-guided weapons, aerospace development, and early warning systems. However, in long-range imaging tasks, infrared small target detection faces significant challenges. These targets are typically small, typically ranging from 2×2 to 9×9 pixels, and often lack significant texture features. Furthermore, they are often dim and have low signal-to-clutter ratios due to complex backgrounds and noise. Therefore, developing robust and efficient infrared small target detection techniques has attracted widespread attention from both academic and practical communities. Over the past few decades, numerous effective infrared small target detection methods have been developed to address a wide range of scenarios. These methods can be roughly categorized into two categories: data-driven and model-driven. Benefiting from the availability of publicly available infrared image datasets and the development of deep learning methods, data-driven infrared small and dim target detection methods have achieved promising results in a variety of scenarios, effectively reducing the need for complex manual parameter configuration. Among these innovative data-driven methods, two landmark achievements are the false detection versus missed detection (MDvsFA) method and the asymmetric context modulation network (ACMNet) method. Based on the optimization of target detection network models, many related works have been proposed. However, data-driven methods are overly dependent on the quality and quantity of dataset samples, resulting in insufficient generalization performance. In addition, the lack of interpretability of deep learning in many practical target detection systems also limits its application. Model-driven methods can be roughly divided into three categories: background clutter suppression (BCS) methods, human visual system (HVS) methods, and low-rank and sparse decomposition (LRSD) methods. BCS methods mainly use the contrast between the target and the background in spatial information to achieve target detection, such as top-hat filters
[10] and maximum median filters. These methods usually have fast processing efficiency and are widely used in engineering. However, their performance is quite sensitive to background clutter and target size changes. HVS-based methods simulate the human eye's perception process of infrared images and integrate indicators such as contrast, color, and frequency changes to measure the local difference between the target and the background. However, these methods often have difficulty suppressing background clutter in strong clutter scenes.
[0003] Methods based on low-rank and sparse decomposition (LRSD) have attracted much attention due to their excellent detection performance. The LRSD method is based on the assumption that the background and target can be represented as low-rank and sparse components, respectively. A pioneering example of this method is the infrared small target detection (IPI) method, which reformulates the small target detection task as a robust principal component analysis (RPCA) problem. In the IPI model, the nuclear norm and l1 norm are used to approximate the rank and l0 norm of the background component, respectively. In order to improve the accuracy of rank estimation, a variety of methods have been introduced, including the reweighted IPI (ReWIPI) method, the non-convex rank approximation model (NRAM), and other methods. These methods convert infrared data into a two-dimensional matrix and are commonly used due to their effectiveness in single-frame infrared target detection.
[0004] However, raw infrared images from infrared search and tracking (IRST) systems are typically sequential data containing both spatial and temporal information. Therefore, reconstruction methods using high-dimensional data matrices can potentially destroy spatiotemporal correlations. To address this limitation, many researchers have proposed methods based on the tensor domain, reformulating the two-dimensional model as follows:
[0005]
[0006] O represents the original image tensor, and N represents the Gaussian noise tensor.
[0007] where D, B, T∈R m×n×L Represent the original infrared tensor, background tensor and target tensor respectively. L represents the time length dimension. S and λ N are the trade-off parameters for sparse targets and noise, respectively. As can be seen from model (1), the key to improving the detection capability of infrared small targets using the low-rank and sparse decomposition (LRSD) method lies in three aspects: constructing an accurate spatiotemporal tensor model based on the characteristics of the input image data, effectively reducing the reconstruction error of the low-rank background component, and accurately estimating the sparse target component.
[0008] To address the three aforementioned issues, some research has focused on constructing appropriate tensor structures that can exploit spatial and temporal internal correlations. Researchers first proposed the Spatiotemporal Infrared Small Target Tensor (STIPT) model, which stacks L consecutive frames of images into a tensor structure. Some researchers process each frame using a sliding window filter, creating a three-dimensional image block tensor by vertically stacking all image blocks obtained from a sequence of L consecutive frames. However, this approach fails to consider the overlapping areas between image blocks, resulting in redundant information processing and increased computational overhead.
[0009] Other researchers have introduced a three-dimensional spatiotemporal tensor that incorporates high-frequency information. They have also developed a new tensor structure that treats the current frame to be detected as an intermediate slice, thereby better extracting temporal information from the previous L frames and the next L frames. Furthermore, some researchers have proposed a four-dimensional data structure, in which the third and fourth dimensions represent global spatial and temporal information, respectively. Several methods have been proposed to further enhance the construction of spatiotemporal data structures. However, these methods suffer from a significant drawback: the time step L is typically set based on empirical experience and cannot be adaptively adjusted based on the intensity differences between frames in the infrared image sequence. For infrared image sequences with slowly changing backgrounds, an excessively small time step may lead to insufficient exploration of temporal correlation information. Conversely, for sequences with rapidly changing backgrounds, an excessively large time step may destroy the low-rank properties of the LRSD model. Therefore, using a fixed time step for different image sequences is not appropriate. To improve the rank estimation accuracy of the low-rank background component B, various advanced tensor decomposition norms have been developed and improved for this purpose. These methods can be categorized into three main categories: those based on Tucker decomposition, those utilizing tensor chain decomposition (TTD) or tensor ring decomposition (TRD), and those employing tensor singular value decomposition (t-SVD). Tucker decomposition utilizes multilinear algebra techniques to decompose a high-rank tensor into several lower-rank tensors interconnected by a core tensor. However, a significant drawback of Tucker decomposition is its high computational requirements, as the number of parameters grows exponentially with the tensor rank. Tensor chain decomposition (TTD) and tensor ring decomposition (TRD) are complex techniques designed to represent and decompose high-rank tensors into a series of lower-rank tensors. However, the TTD and TRD norms exhibit considerable algorithmic complexity when dealing with high-rank tensors. Furthermore, the TTD norm imposes an additional rank-1 constraint on the boundary factors, lacking a clear physical interpretation, while the TRD norm requires a predefined TR rank, which is impractical in practical applications. The t-SVD-based norm is the most widely used norm in LRSD-based methods. It defines the tensor tube rank via the tensor product (t-product). Many methods have emerged based on t-SVD, including the weighted tensor kernel norm (WTNN) model, the weighted Schatten p-norm, and the modulo-k1k2 extension of the tensor tube rank.
[0010] Recently, several researchers have proposed a method that combines nonconvex tensor low-rank approximation (NTLA) with asymmetric sparse tensor total variation (ASTTV). ASTTV is an improved version of STTV that enhances the influence of temporal correlation information by assigning fixed asymmetric weights to the spatiotemporal difference terms. NTLA can achieve more accurate background estimates by introducing a weighted Laplace norm instead of the traditional nuclear norm. However, the key asymmetric spatiotemporal weights are also fixed values set empirically, similar to the time step parameter mentioned above. These methods aim to better approximate the true rank of the low-rank tensor B by assigning different weights to singular values and integrating other advanced norms into the tensor tube rank framework. The important conjugation property of the tensor product in the frequency domain halves the computational burden of t-SVD, significantly improving the efficiency of the algorithm. However, it should be noted that the tensor tube rank is obtained by applying a discrete Fourier transform (DFT) to the third dimension of the original tensor, which requires a transpose operation on the tensor. Essentially, the singular value decomposition (SVD) results of the transposed tensor show differences, a phenomenon known as transposed variability of tensor recovery (TVTR). Focusing only on a single dimension may lead to abandoning low-rank prior knowledge obtained from multiple perspectives, which is detrimental to the accuracy of the tensor recovery process. For recovering sparse target components, current methods usually use the l1 norm as an approximation of the l0 norm. Although the l1 norm penalty in (1) ensures that the overall optimization problem remains convex, thus facilitating direct solution, it has a significant limitation: it tends to produce biased estimates when dealing with large sparse coefficients.
[0011] In view of the above challenges, the present invention proposes a new method to simultaneously enhance target detection and background suppression capabilities. First, the present invention proposes an adaptive infrared spatiotemporal tensor block model based on tensor information entropy. Then, to solve the problem of low-rank background recovery error caused by the TVTR attribute, the present invention designs a new non-convex norm that combines the advantages of the Laplace norm and the weighted average tensor rank (WTAR) norm. Finally, to improve the estimation accuracy of sparse targets, the present invention replaces the l1 norm with the smoothed truncated absolute deviation (SCAD) norm and extends it to the tensor space. The present invention calls the proposed method the entropy-based adaptive spatiotemporal infrared tensor and non-convex weighted average tensor rank (EASTIT-NWTAR) method for infrared small target detection. Summary of the Invention
[0012] The purpose of the present invention is to solve the above technical problems and propose an infrared dim small target detection method based on non-convex weighted tensor rank minimization and adaptive spatiotemporal modeling, which performs the following steps:
[0013] S1. The server builds the EASTIT model; R represents a real or imaginary number, p represents the total number of image frames, K = m × n, m, n represent the image size, L represents the time parameter of the spatiotemporal domain image tensor block, U i , S i , V i They represent the components of the singular value decomposition.
[0014] S11. The server obtains the infrared image sequence O1,...,O p ∈R m×n×p ;
[0015] Superimpose the subsequent L-1 frames onto the first frame to construct the initial tensor X i ∈R m×n×L ;
[0016] S12. Expand the tensor along the time dimension to form a matrix x i =unfold(X i ,3)∈R K×L ;
[0017] S13. For matrix x i Perform singular value decomposition SVD:
[0018] x i =U i *S i *V i ;
[0019] S14. Get singular values Shannon entropy is defined as follows:
[0020]
[0021] S15. Add the next frame along the time dimension to form the i+1th tensor, and pass the x in steps S13 and S14 i =U i *S i *V i as well as The information entropy H(x i+1 );
[0022] S16. If the increment of the information entropy is less than a predetermined threshold E1, it is determined that after adding the new frame, the i+1th tensor still maintains the low-rank property;
[0023] If the information entropy increment exceeds the threshold E1, the current frame is designated as the starting frame of the subsequent infrared image tensor, and steps S11-S16 are repeated until all images in the sequence are processed;
[0024] S2. Construct the EASTIT-NWTAR norm;
[0025] S3. Optimize the solution process;
[0026] S4. Output the detection results.
[0027] Furthermore, x represents the function variable, α k Represents the weight coefficients of the three tensor transpose dimensions, n k ,k=1,2,3, Represents the background tensor, the transpose of B along the kth dimension
[0028] S21. The Laplace function is defined as φ(x) = 1-e -x / ε ;
[0029] S22. Incorporating the Laplace function into the WTAR norm, we obtain:
[0030]
[0031] in,
[0032] Furthermore, the S23.EASTIT-NWTAR model is defined as follows:
[0033]
[0034] Among them, ||B|| γ,wa is the Laplace function with WTAR norm, λ tv ,λ S and λ N denote the positive trade-off parameters of the ASTTV term, the sparse target component, and the noise component, respectively;
[0035] First item ||B|| γ,wa Used to recover low-rank background components and solve TVTR errors, wa is the weighted average rank
[0036] Abbreviation, γ represents the Laplace function; the overall definition is as follows:
[0037]
[0038] Second item ||B|| AASTTV is incorporated to more accurately preserve sparse structure interference in background regions, thereby reducing the false alarm rate;
[0039] The third SCAD penalty R SCAD (T) is introduced to solve the biased estimation problem of the l0 norm, which can improve the reconstruction accuracy of sparse targets; T represents the sparse target tensor, and the l0 norm represents the non-zero elements in the variable.
[0040] The last Frobenius norm Used to suppress Gaussian noise; N represents the noise tensor, F represents the Frobenius norm, Indicates that the Frobenius norm is used to model Gaussian noise.
[0041] Furthermore, for step S4, after the solution process is optimized, data output is performed according to the optimal solution process.
[0042] The beneficial effects of the present invention are as follows:
[0043] 1. This paper proposes a new method that uses tensor information entropy to measure the temporal variation of an input infrared image sequence. This method has two significant advantages: First, it can adaptively set the time step parameter based on background variations, construct an infrared spatiotemporal block tensor, and ensure that the background component in each sub-block tensor satisfies the low-rank property, thereby reducing the algorithm's performance sensitivity to manual parameter settings based on expert experience. Second, information entropy can also help adaptively determine the contribution of the temporal component in background reconstruction, effectively adjusting the weight coefficient of the temporal component in STTV regularization.
[0044] 2. In order to solve the TVTR problem caused by the DFT operation, the present invention proposes a new non-convex weighted low-rank tensor norm based on the tensor average rank and Laplace norm. The advantages of this new weighted norm are mainly reflected in two aspects. First, it assigns weights to all possible transposed tensors B, providing multiple perspectives to describe low-rank prior information, rather than being limited to a single dimension. Secondly, the weighting of the Laplace norm not only allows differentiated processing of singular values, but also effectively alleviates the over-contraction problem inherent in the traditional nuclear norm minimization (NNM) method. Therefore, the weighted low-rank tensor norm proposed in the present invention can comprehensively improve the accuracy of recovering B, which has been verified by a large number of experiments.
[0045] 3. To effectively overcome the biased estimation problem caused by using the l1 norm as an approximation to the l0 norm, the present invention replaces the l1 norm with the SCAD norm. Compared with the l1 norm, the SCAD norm provides an unbiased estimate, thereby improving the reconstruction accuracy of sparse target components. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0047] Figure 1 It is the overall flow chart of the present invention; DETAILED DESCRIPTION
[0048] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0049] The present invention provides an infrared dim small target detection method based on non-convex weighted tensor rank minimization and adaptive spatiotemporal modeling.
[0050] 1. Information entropy
[0051] Information entropy is a fundamental concept in information theory that measures the uncertainty or information content of a random variable. In image processing, information entropy is often used as a metric to assess the complexity of grayscale distributions. Images with uniform grayscale distributions have low information entropy, while images with complex grayscale distributions have high information entropy.
[0052] For infrared image O∈R m×n , we vectorize the image matrix into a vector Where K = m × n. The probability of each value P(X = i) is defined as:
[0053]
[0054] The Shannon entropy of a random variable X is defined as:
[0055]
[0056] The low-rank property means that the singular values of the data structure are concentrated on a small number of principal components. Therefore, we propose an entropy based on singular value decomposition (SVD) to measure the degree of variation between different frames in an image sequence. Specifically, it quantifies the uncertainty of the pixel intensity distribution between different frames. A higher increase in entropy indicates a larger variation in pixel values and a more significant difference between frames, while a lower increase in entropy indicates a higher similarity between frames. Therefore, by adopting SVD-based entropy, we can adaptively segment the input infrared image sequence into tensor blocks according to the complexity and dynamics of the background, thereby ensuring the effectiveness of low-rank and sparse decomposition.
[0057] 2. Weighted Tensor Average Rank
[0058] We introduce a new non-convex norm to address the transposition error problem in existing t-SVD-based methods. The method is based on the weighted tensor average rank and Laplace norm. The tensor average rank and weighted tensor average rank are defined as follows:
[0059] Definition 1 (Tensor Tube Rank): The average rank of a tensor X is defined as:
[0060]
[0061] Where bcirc(·) represents the block circulant matrix, which is defined as follows:
[0062]
[0063] where X (i) represents the i-th forward slice.
[0064] Definition 2 (weighted tensor average rank): The weighted tensor average rank of a tensor X is defined as:
[0065]
[0066] Among them, α k Denote the weights of all possible transposed matrices.
[0067] It represents the transpose of the background tensor X along the k-th dimension, and T stands for Transport transpose.
[0068] 3.STTV Regularization
[0069] Total variation (TV) regularization has been widely used to describe non-smooth structures and noise in images. Sun first proposed spatial-temporal total variation (STTV) regularization, which can effectively preserve sparse structures, such as strong corners and edges, in low-rank backgrounds. These preserved sparse structures are similar to sparse targets and may lead to a high false alarm rate. STTV regularization is defined as follows:
[0070] ||X|| STTV =||D h X||1+||D v X||1+||D z X||1 (7)
[0071] Where Dh, Dv, and Dz represent the differential operators in the horizontal, vertical, and time directions, respectively, and are defined as follows:
[0072]
[0073] In addition, Liu proposed an asymmetric STTV (ASTTV) method by introducing a parameter δ to assign different weights to the TV regularization term in the time direction. This adjustment allows to adjust the impact of temporal correlation information on the overall performance of the detection framework:
[0074] ||X|| ASTTV =||D h X||1+||D v X||1+δ||D z X||1 (9)
[0075] However, in the ASTTV method, the parameter δ is manually set to 1.25 based on empirical observations, lacking the ability to adaptively adjust to the characteristics of the input data. This fixed setting significantly limits the method's ability to handle diverse scenes, as the optimal value of δ may vary depending on the specific characteristics of the image sequence, such as background complexity, object size, and noise level. To overcome this limitation, further improvements are needed to enhance its adaptability and performance in a wider range of practical applications.
[0076] 4. SCAD Penalty
[0077] In this section, we use the Smoothly Clipped Absolute Deviation (SCAD) penalty to approximate the l0 norm instead of the l1 norm, thereby achieving accurate unbiased estimation and enhanced reconstruction accuracy. The description of the SCAD penalty is as follows:
[0078]
[0079] Among them, μ>0 represents the threshold, a>2, and is usually set to 3.7.
[0080] This paper proposes the EASTIT-NWTAR model to address the shortcomings of existing LRSD-based methods, including an adaptive infrared spatiotemporal tensor construction method based on SVD entropy, a new non-convex norm and SCAD penalty to better recover low-rank components and sparse components. Figure 1 As shown in Figure 3, the proposed EASTIT-NWTAR method consists of three main components: spatiotemporal tensor structure construction, object detection, and image restoration. First, to address the limitations of existing low-rank subspace decomposition (LRSD) methods, this method uses SVD-based information entropy to adaptively segment the image sequence into multiple tensor blocks based on the degree of background variation. These segmented tensor blocks are then input into the NWTAR model for reconstruction and decomposition of the low-rank background tensor and sparse object tensor. Finally, the tensors are restored back to the image sequence.
[0081] S1. Building the EASTIT model
[0082] According to the above definition, as the complexity of the image background increases, the entropy value increases, while when the grayscale distribution is concentrated on a few values, the entropy value decreases. In addition, the low-rank structure of the image background indicates that the singular values are mainly aligned with a few principal components. Inspired by the related work, this paper proposes a tensor entropy based on singular value decomposition for infrared image sequences. For the infrared image sequence O1,...,O p ∈R m×n×p , first superimpose the subsequent L-1 frames onto the first frame to construct the initial tensor X i ∈R m×n×L . Then, the tensor is expanded along the time dimension to form the matrix x i =unfold(X i ,3)∈R K×L , and for the matrix x i Perform a singular value decomposition (SVD) operation:
[0083] x i =U i *S i *V i (11)
[0084] Get singular values Its Shannon entropy is defined as follows:
[0085]
[0086] Next, the next frame is added in sequence along the time dimension to form the i+1th tensor, and the information entropy H(x) based on the singular value is obtained by formulas (11) and (12). i+1 The increment of information entropy can be calculated by H(x i+1 ) and H(x i ) is calculated as a measure of the intensity of image background changes between frames. If the increase in information entropy is less than a predetermined threshold, it is assumed that the i+1th tensor still maintains its low-rank property after adding the new frame. Conversely, if the increase exceeds the threshold, the current frame is designated as the starting frame for the subsequent infrared image tensor, and the above process is repeated until all images in the sequence have been processed.
[0087] S2. Constructing the EASTIT-NWTAR norm
[0088] For the i-th infrared subtensor obtained by the EASTIT method, the target detection problem is transformed into a decomposition of low-rank background and sparse target components. To further improve the approximate accuracy of the low-rank background B, the Laplace function is defined as φ(x) = 1-e -xε. It performs better than the traditional nuclear norm in approximating the rank of low-rank components. However, the Laplace function still has the TVTR problem because it is also optimized based on the t-SVD operation. Therefore, the present invention proposes the NWTAR norm, which incorporates the Laplace function into the WTAR norm and is defined as follows:
[0089]
[0090] in
[0091]
[0092] In addition, this paper proposes an improved ASTTV regularization, namely adaptive asymmetric STTV regularization (AASTTV), to describe the sparse structure in the background region. The adjustment parameter δ can be set dynamically based on information entropy rather than a fixed value, which means that it can be adaptively adjusted based on the changing characteristics of the input infrared image sequence. Therefore, the EASTIT-NWTAR model proposed in this paper is defined as follows:
[0093]
[0094] Among them, ||B|| γ,wa is the Laplace function with WTAR norm, λ tv ,λ S and λ N Denote the positive trade-off parameters of the ASTTV term, the sparse target component, and the noise component, respectively. The first term ||B|| γ,wa It is used to recover the low-rank background component and solve the TVTR error. The second term ||B|| AASTTV is incorporated to more accurately preserve the sparse structure interference in the background area, thereby reducing the false alarm rate. The third SCAD penalty R SCAD (T) is introduced to solve the problem of biased estimation of l0 norm, which can improve the reconstruction accuracy of sparse targets. The last Frobenius norm Used to effectively suppress Gaussian noise. Through (9), model (14) can be transformed as follows:
[0095]
[0096] In summary, the proposed EASTIT-NWTAR model can adaptively construct spatiotemporal infrared tensor blocks and adjust the temporal correlation parameter δ based on the degree of background variation in the input infrared image sequence. Furthermore, this method effectively improves the reconstruction accuracy of the background component and the target tensor by addressing the biased estimation issues of the TVTR error and the l0 norm.
[0097] S3. Optimization solution process
[0098] This paper proposes an effective optimization framework to solve Problem (15) based on ADMM. First, several auxiliary variables are introduced to reduce the optimization difficulty, including Mk and Vk, k = 1, 2, 3. Problem (15) can be rewritten as follows:
[0099]
[0100] The two variables Mk and Vk represent auxiliary multiplier tensors, which are used to simplify the solution difficulty.
[0101] Problem (16) can be transformed into the following form using the inexact augmented Lagrange multiplier (IALM) method:
[0102]
[0103] Where μ is a positive penalty scalar, Qk,k=1,…,3 and Yi,i=1,…,4 represent the Lagrange multipliers. Problem (17) can be decomposed into six optimization subproblems, as follows:
[0104] 1) Update Mk and fix other variables:
[0105]
[0106] Where t represents the number of iterations. H(·,thr) represents the adaptive soft threshold operation, thr = α k / u t , which aims to solve the proposed non-convex Laplace function and WTAR optimization problem.
[0107] 2) Update B and fix other variables:
[0108]
[0109] Taking the derivative of the above equation and setting it to zero, we get:
[0110] (1+μ t +Δ)B t+1 =L t +β1+β2+β3 (20) where:
[0111]
[0112] The closed-form solution to the problem obtained by nFFT operation is as follows:
[0113] 3) Update T and fix other variables:
[0114]
[0115] in
[0116]
[0117] where ξ=λ s / μ, sign(·) represents the sign function
[0118] 4) Update Zi and fix other variables:
[0119]
[0120] The above equation can be solved by element-by-element contraction operation Th(·):
[0121]
[0122] 5) Update N and fix other variables
[0123]
[0124] 6) Update the Lagrange multiplier and fix other variables:
[0125]
[0126] 7) Update μ t+1 :
[0127] μ t+1 =min(ρμ t ,μ max )(29)
[0128] Experimental results verification
[0129] In the real world, infrared small target detection is often susceptible to various noise artifacts, such as sensor interference and environmental clutter. Therefore, target detection methods must possess excellent noise immunity. In this study, simulated background clutter containing Gaussian and streak noise with a standard deviation of 0.004 was added to sequences 1, 3, and 5 to evaluate the performance of the proposed method. The proposed method effectively suppresses both Gaussian and streak noise, highlighting its robustness in noisy environments.
[0130] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A method for infrared dim small target detection based on non-convex weighted tensor rank minimization and adaptive spatiotemporal modeling, characterized in that: S1. The server builds the EASTIT model; R represents a real or imaginary number, p represents the total number of image frames, K = m × n, m, n represent the image size, L represents the time parameter of the spatiotemporal domain image tensor block, U i , S i , V i Respectively represent the components of singular value decomposition; S11. The server obtains the infrared image sequence O1,...,O p ∈R m×n×p ; Superimpose the subsequent L-1 frames onto the first frame to construct the initial tensor X i ∈R m×n×L ; S12. Expand the tensor along the time dimension to form a matrix x i =unfold(X i ,3)∈R K×L ; S13. For matrix x i Perform singular value decomposition SVD: x i =U i *S i *V i ; S14. Get singular values Shannon entropy is defined as follows: S15. Add the next frame along the time dimension to form the i+1th tensor, and pass the x in steps S13 and S14 i =U i *S i *V i as well as The information entropy H(x i+1 ); S16. If the increment of the information entropy is less than a predetermined threshold E1, it is determined that after adding the new frame, the i+1th tensor still maintains the low-rank property; If the information entropy increment exceeds the threshold E1, the current frame is designated as the starting frame of the subsequent infrared image tensor, and steps S11-S16 are repeated until all images in the sequence are processed; S2. Construct the EASTIT-NWTAR norm; S3. Optimize the solution process; S4. Output the detection results.
2. The infrared small target detection method based on non-convex weighted tensor rank minimization and adaptive spatiotemporal modeling according to claim 1 is characterized in that: x represents the function variable, α k Represents the weight coefficients of the three tensor transpose dimensions, n k ,k=1,2,3,B Tk Represents the background tensor, the transpose of B along the kth dimension S21. The Laplace function is defined as φ(x) = 1-e -xε ; S22. Incorporating the Laplace function into the WTAR norm, we obtain: in, 3. The infrared small target detection method based on non-convex weighted tensor rank minimization and adaptive spatiotemporal modeling according to claim 2 is characterized in that: S23. The EASTIT-NWTAR model is defined as follows: Among them, ||B|| γ,wa is the Laplace function with WTAR norm, λ tv ,λ S and λ N denote the positive trade-off parameters of the ASTTV term, the sparse target component, and the noise component, respectively; First item ||B|| γ,wa It is used to recover the low-rank background component and solve the TVTR error. Wa is the abbreviation of the weighted average rank, and γ represents the Laplace function. The overall definition is as follows: Second item ||B|| AASTTV is incorporated to more accurately preserve sparse structure interference in background regions, thereby reducing the false alarm rate; The third SCAD penalty R SCAD (T) is introduced to solve the biased estimation problem of the l0 norm, which can improve the reconstruction accuracy of sparse targets; T represents the sparse target tensor, and the l0 norm represents the non-zero elements in the variable. The last Frobenius norm Used to suppress Gaussian noise; N represents the noise tensor, F represents the Frobenius norm, Indicates that the Frobenius norm is used to model Gaussian noise.
4. The infrared small target detection method based on non-convex weighted tensor rank minimization and adaptive spatiotemporal modeling according to claim 3 is characterized in that: For step S4, after the solution process is optimized, data output is performed according to the optimal solution process.