An Infrared Weak Target Detection Method Based on Adaptive Spatiotemporal Tensor and Weighted Tensor Average Rank Approximation
By using an adaptive spatiotemporal tensor and a weighted tensor average rank approximation method, the number of frames is dynamically adjusted and combined with the weighted Schatten p norm to optimize the low-rank background and sparse target decomposition of infrared image sequences. This solves the problems of detection accuracy and background suppression in infrared weak target detection, and achieves efficient and accurate target detection.
Patent Information
- Application Number
- CN202411805510.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-10
AI Technical Summary
Existing infrared weak target detection methods are difficult to effectively detect and recover targets under conditions of long imaging distance, small target size, complex background and noise interference. In particular, model-driven methods have limitations in terms of resources and generalization performance, and existing tensor domain methods lack physical meaning and efficiency.
An adaptive spatiotemporal tensor and weighted tensor average rank approximation method is adopted. The number of frames is dynamically adjusted by structural similarity index. Combined with weighted Schatten p norm, generalized tensor singular value threshold and alternating direction multiplier method, the low-rank background and sparse target component decomposition of infrared image sequences is optimized.
It improves the accuracy of infrared weak target detection and background clutter suppression, overcomes the problems of improper frame number setting and singular value estimation error in traditional methods, and achieves efficient and accurate target detection in complex scenes.
Smart Images

Figure CN119762752B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, specifically to an infrared weak target detection method based on adaptive spatiotemporal tensor and weighted tensor average rank approximation. Background Technology
[0002] Infrared small target detection is a complex technology crucial in numerous military and civilian applications, including mine detection, night navigation, precision-guided weapons, aerospace technology, and early warning systems. However, due to the long imaging distance, infrared small target detection still faces significant challenges. Infrared targets are extremely small, typically between 2×2 and 9×9 pixels, and lack texture features. Furthermore, targets are usually dark with low signal-to-noise ratios because they are obscured by complex backgrounds and noise. Therefore, the complexity and criticality of robust and efficient infrared target detection methods have attracted widespread academic and practical attention.
[0003] Over the past few decades, many excellent infrared small target detection methods have been proposed for various scenarios. These methods can be divided into data-driven methods and model-driven methods.
[0004] In recent years, with the development of public infrared datasets and deep learning technologies, data-driven methods have been proposed, achieving efficient and satisfactory performance. False negatives versus false positives (MDvsFA) and asymmetric context modulation networks (ACMNet) are pioneering works in data-driven methods. Li et al. improved small target detection through densely nested attention networks (DNANet). However, data-driven methods require a large number of samples for effective feature acquisition, which can be resource-intensive. The inherent limitations of neural networks in terms of physical interoperability and limited generalization performance restrict their application in practical infrared search and track (IRST) systems. Therefore, model-driven methods still have significant practical engineering value and research significance.
[0005] Generally, model-driven methods fall into three categories: Background Clutter Suppression (BCS) methods, Human Visual System (HVS) methods, and Low-Rank Sparse Decomposition (LRSD) methods. BCS methods typically utilize the spatial contrast between the target and the background to design different filters; these methods are effective, including Top-hat filters and Maxmedian filters. However, their detection performance is highly sensitive to target size and complex background clutter. HVS methods utilize contrast mechanisms to detect infrared targets, such as Local Contrast Methods (LCM). These methods cannot suppress strong background clutter similar to the target.
[0006] The LRSD method assumes that the background and target can be modeled as low-rank and sparse components, respectively. The Infrared Patch Image (IPI) model is a pioneering work that transforms the small target detection task into a robust principal component analysis (RPCA) problem. Generally, the LRSD model can be formulated as follows:
[0007]
[0008] Where D, B, T ∈ R m×n Let represent the original image, background image, target image, and noisy image, respectively. λ1 is the positive sparse component parameter. rank(·) denotes the rank estimation operation, ||·|| 0 and ||·|| 0. F Let represent the l0 norm and the Frobenius norm, respectively. The l0 norm represents the number of all non-zero elements, and the Frobenius norm represents the square root of the sum of squares of all elements.
[0009] In the IPI model, the kernel norm and l1 norm are used to approximate the rank(·) and l0 norm, respectively. To improve the accuracy of rank estimation, many methods have been proposed, including the reweighted IPI (ReWIPI) method and the non-convex rank approximation model (NRAM). In summary, model (1) focuses on the two-dimensional matrix domain, and because it can utilize spatial information, it is widely used in single-frame infrared target detection. However, the original infrared image is usually sequential data from the IRST system, containing both spatial and temporal information. Obviously, the reconstruction method of the high-dimensional data matrix will severely destroy the spatiotemporal correlation. Therefore, many researchers have proposed a tensor-domain-based method to solve this problem, and formula (1) can be restated as follows:
[0010]
[0011] Where D, B, T ∈ R m×n×L Let represent the original infrared tensor, background tensor, and target tensor, respectively. L represents the time dimension. The Reweighted Infrared Patch Tensor (RIPT) method was initially developed to convert a single infrared image into a higher-order tensor. However, the RIPT method constructs the tensor structure through sliding windows and stacking overlapping image patches, which is time-consuming and lacks physical meaning. Furthermore, the sum of kernel norms (SNN) used in RIPT is not optimal because it expands the tensor into a matrix along different directions, destroying the internal spatial information.
[0012] To address the aforementioned limitations of existing LRSD methods, this invention proposes a novel infrared small target detection method based on weighted Schatten p-norm and tensor average rank adaptive spatiotemporal infrared tensor (WPTAR-ASTIT). This method focuses on comprehensively improving the accuracy of approximating low-rank background components and sparse target components. Summary of the Invention
[0013] To address the problems of the existing technologies, this invention provides an infrared weak target detection method based on adaptive spatiotemporal tensor and weighted tensor average rank approximation. By manually setting the key parameter L to construct the constraint of the spatiotemporal infrared tensor, we propose an adaptive metric that combines the structural similarity index (SSIM) to describe the degree of change between different frames from three perspectives: pixel intensity, contrast, and structure. By calculating the SSIM index of the first and last frames of the sub-tensor, the number of frames L is dynamically adjusted to ensure that the degree of change between frames within each sub-tensor meets the requirements of low-rank background prior.
[0014] The technical solution adopted in this invention is as follows:
[0015] An infrared weak target detection method based on adaptive spatiotemporal tensor and weighted tensor average rank approximation, comprising the following steps:
[0016] S1. Obtain infrared sequence variation data from the server and establish an ASTIT model. The SSIM is defined as follows:
[0017]
[0018] Where x and y represent two images, μ x ,μ y ,σ x ,σ y Let x and y represent the mean and variance of the graph, respectively. The coefficient C1 = (K1H) 2 C2 = (K2H) 2 K1 and K2 are usually set to K1 = 0.01, K2 = 0.03, and H represents the dynamic range of the pixel, which is 255 for infrared images;
[0019] S2. Establish the WPTAR-ASTIT model;
[0020] For the infrared subtensor D∈R m×n×L It is decomposed into low-rank background and sparse target components:
[0021]
[0022] S3. Optimize the solution process;
[0023] Will Simplified to:
[0024]
[0025] S4. Solve the data to obtain the results.
[0026] Furthermore, S21. Constructing the ASTIT structure: This involves combining the SSIM metric to describe the degree of change in background clutter, using the original input infrared image sequence O1,...,O P Adaptively converted into multiple infrared subtensors D∈R m×n×L This ensures that the low-rank prior conditions in each subtensor are satisfied.
[0027] Furthermore, in step S2, the number of frames L in the proposed model is dynamically adjusted according to the data characteristics, rather than being a pre-fixed value.
[0028] Furthermore, S22. Decompose background and target components: The transformed subtensor is decomposed into low-rank background components and sparse target components using the WPTAR-ATSIT model.
[0029] Furthermore, S23. Reconstruct the target image: Reconstruct the target image from the obtained target tensor through the inverse operation.
[0030] Furthermore, S24. Separate the real target: Further separate the real target by using an adaptive threshold.
[0031] The present invention has the following beneficial effects:
[0032] 1. To overcome the limitations of constructing spatiotemporal infrared tensors by manually setting key parameters L, we propose an adaptive metric that combines the Structural Similarity Index (SSIM) to describe the degree of variation between different frames from three perspectives: pixel intensity, contrast, and structure. We can dynamically adjust the number of frames L by calculating the SSIM exponents of the first and last frames of the sub-tensor, ensuring that the degree of variation between frames within each sub-tensor satisfies the requirements of a low-rank background prior.
[0033] 2. To address the TVTR problem caused by DFT operations in the t-product norm, we propose a novel weighted low-rank tensor norm based on the tensor average rank and the weighted Schatten p-norm. This novel weighted norm offers two advantages. First, assigning weights to all possible transposes of B provides different perspectives for describing low-rank prior information, rather than focusing solely on a single dimension. Second, assigning weights to the Schatten p-norm not only allows for different treatment of singular values but also effectively alleviates the over-shrinkage problem inherent in traditional nuclear norm minimization (NNM) methods, as verified in our previous work. Therefore, the proposed weighted low-rank tensor norm comprehensively improves the accuracy of B-recovery, as validated in extensive experiments.
[0034] 3. To improve the accuracy of the l0-norm approximation of the sparse target component T, we first use the Schatten p-norm to describe the sparsity instead of the l1-norm. Furthermore, the Schatten p-norm alleviates the strict recovery condition in the l1-norm, which represents the sum of the absolute values of all elements and is more applicable in the real world.
[0035] 4. By combining the generalized tensor singular value thresholding method with the alternating direction multiplier method (ADMM), an optimization framework was developed to accurately and efficiently recover background and target components. Attached Figure Description
[0036] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0037] Figure 1 It is a process framework diagram.
[0038] Figure 2 This is a scene diagram for multi-target detection.
[0039] Figure 3 This is a scene diagram of the detection in a noisy environment. Detailed Implementation
[0040] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0041] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.
[0042] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0043] This invention provides an infrared weak target detection method based on adaptive spatiotemporal tensor and weighted tensor average rank approximation, including: adaptive spatiotemporal domain infrared tensor block construction, low-rank and sparse decomposition model, and image reconstruction component, such as... Figure 1 As shown. This invention can simultaneously improve target detection and background suppression capabilities to meet the urgent need for infrared (IR) target detection technology in complex scenarios.
[0044] As mentioned above, methods based on the t-SVD norm suffer from transpose errors on low-rank tensors. In this section, we primarily introduce the weighted tensor average rank to explore different perspectives on low-rank prior information.
[0045] Definition 1 (Tensor Mean Rank, WTTR): For tensor X∈R m×n×L The tensor average rank of X is defined as follows:
[0046] in X (i) Let X be the i-th slice along the third dimension. Definition 2 (Weighted Tensor Average Rank, WTAR): The weighted tensor average rank of X is defined as follows:
[0047] Where α k Let represent the weighting coefficients of different transpose components, and satisfy . This represents the tensor after the tensor X is transposed along the k-th dimension.
[0048] Based on WTTR and WTAR, we can solve the TVTR problem caused by the transpose operation in the t-SVD method.
[0049] The proposed target detection method comprises two stages. First, a spatiotemporal tensor model is constructed, utilizing the inherent correlation of the input infrared image sequence; its low-rank property has been validated in the LRSD method. Subsequently, the proposed method efficiently and accurately decomposes the original tensor into background and target components.
[0050] S1. Obtain infrared sequence variation data from the server and establish a STIT model.
[0051] A STIT model is constructed by combining SSIM (Sensory Measurement Metric) to adaptively set the frame number L based on the degree of change in the infrared sequence. SSIM is a perceptual metric widely used in image processing and computer vision to evaluate the similarity or difference between two images based on their brightness, contrast, and structural information.
[0052] SSIM provides a value between -1 and 1, with values closer to 1 indicating higher similarity between images. Conversely, values closer to -1 indicate lower similarity, suggesting significant differences between images. SSIM is defined as follows:
[0053]
[0054] Where x and y represent two images, μ x ,μ y ,σ x ,σ y Let x and y represent the mean and variance of the graph, respectively. The coefficient C1 = (K1H) 2 C2 = (K2H) 2 K1 and K2 are usually set to K1 = 0.01, K2 = 0.03, and H represents the dynamic range of the pixel, which is 255 for infrared images;
[0055] In existing LRSD methods, a fixed time step is typically used to construct the infrared patch tensor. However, in scenes with rapidly changing backgrounds, a fixed time step may result in an infrared patch tensor containing two or more frames with significant background differences, thus compromising the low-rank property of the tensor in the time dimension.
[0056] Given an IR image sequence O1,...,O P Calculate the first frame image O1 and the subsequent i-th frame image O i The SSIM between the two frames is calculated and compared to a threshold. If the SSIM is higher than the threshold, structural similarity is considered to exist between the two frames, fully satisfying the low-rank prior condition, and therefore it can be constructed as an infrared subtensor. Otherwise, the last frame image that does not meet the threshold requirement is used as the first slice of the next subtensor, and structural similarity calculation is recalculated for subsequent images until the last frame image in the sequence. Thus, multiple infrared subtensors D∈R can be obtained. m×n×L .
[0057] By applying an SSMI threshold, infrared images with similar structural features are clustered into a single infrared subtensor. These images exhibit strong spatiotemporal correlations, thereby improving the accuracy of recovering the low-rank background tensor. In the data construction method of this invention, the frame number L is dynamically adjusted based on the intensity of background changes between sequential images. If the background changes smoothly, L is larger; conversely, if the changes are more abrupt, L is smaller, to ensure the low-rank characteristics of the background tensor. This method effectively solves the problem of previous methods relying too heavily on manually set frame numbers based on experience, which often leads to suboptimal performance.
[0058] S2. Establish the WPTAR-ASTIT model.
[0059] For the infrared subtensor D∈R m×n×L It is decomposed into low-rank background and sparse target components, as shown below:
[0060]
[0061] The proposed model is a continuous version of Equation (2). Based on WTAR and the weighted Schatten p-norm, a new weighted low-rank tensor norm is proposed to solve the TVTR problem, as shown below:
[0062]
[0063] Where bcirc represents the block cyclic matrix, σ i Represents singular values, p is the key parameter of the Schatten p-norm, w i The definitions assigned to different singular values are as follows:
[0064]
[0065] Furthermore, the Schatten p-norm is used to approximate the l0 norm of the sparse components instead of the l1 norm to improve the accuracy of the recovery, as defined below:
[0066]
[0067] Therefore, the novel weighted low-rank tensor norm can effectively solve the singular value estimation error caused by the transpose operation in the t-SVD method, thereby integrating low-rank relevant information from all possible perspectives. Furthermore, the p-norm can overcome the shortcomings of the l1-norm in practical scenarios. Thus, the WPTAR-ASTIT model proposed in this invention comprehensively improves spatiotemporal data tensor construction, low-rank background estimation, and sparse target recovery. This model significantly improves the detection accuracy of small infrared targets and the suppression of background clutter.
[0068] S3. Optimize the solution process
[0069] In this invention, we propose an efficient optimization framework based on ADMM to solve equation (6). Equation (6) can be simplified to:
[0070]
[0071] To reduce the difficulty of solving formula (10), multiple auxiliary tensors Z are introduced. k To replace the transpose tensor BTk with k = 1, 2, 3, the following expression is used:
[0072]
[0073] Equation (11) can be solved using the Inexact Augmented Lagrange Multiplier (IALM) method, as detailed below:
[0074]
[0075] v, where μk is the positive penalty scalar and Y represents the Lagrange multiplier. Equation (12) can be decomposed into four optimization subproblems, as detailed below:
[0076] 1) Update Z while keeping other variables fixed:
[0077]
[0078] Where t represents the number of iterations, and F(·) represents the version of the improved tensor singular value thresholding (t-SVT) method, defined as Algorithm 1 for solving the weighted Schatten p and tensor average rank norm (WPTAR) optimization problem.
[0079] 2) Update B while keeping other variables fixed:
[0080] Differentiate the above equation and set its value to zero. We can obtain:
[0081]
[0082] 3) Update T while keeping other variables fixed:
[0083]
[0084] 4) Update Yk while keeping other variables fixed:
[0085]
[0086] A. The process of the WPTAR-ASTIT method
[0087] The specific process for establishing the WPTAR-ASTIT model for S2 is as follows.
[0088] 1) Constructing the ASTIT structure: The degree of background clutter variation is described by combining the SSIM metric, with the original input infrared image sequence O1,...,O P Adaptively converted into multiple infrared subtensors D∈R m×n×L This ensures that the low-rank prior conditions in each subtensor are satisfied. It is important to note that the number of frames L in the proposed model is dynamically adjusted based on the data characteristics, rather than being a pre-fixed value.
[0089] 2) Decompose background and target components: The transformed subtensor is decomposed into low-rank background components and sparse target components using the WPTAR-ATSIT model.
[0090] 3) Reconstruct the target image: Reconstruct the target image from the obtained target tensor through inverse operation.
[0091] 4) Separate the real target: Further separate the real target using an adaptive threshold, defined as follows:
[0092] Thr = max(v) min ,μ+kσ) (18)
[0093] Where μ and σ represent the mean and standard deviation of the target image, and k is an empirically determined constant. min =0.75 multiplied by the maximum gray value is an adaptive value.
[0094] B. Verification of Experimental Results
[0095] Robustness in multi-target and noisy scenarios
[0096] This section verifies the robustness of the proposed method in various scenarios, including multi-target scenarios, Gaussian noise pollution, and stripe noise interference.
[0097] 1) Multi-target scenarios: In the field of infrared target detection, multi-target detection capability is particularly important, especially in scenarios such as dense fire attacks and drone swarms. Therefore, we conducted tests to evaluate the multi-target detection performance of the proposed method, where multiple targets were synthesized using a similar method. For example... Figure 2 As shown, the test backgrounds include cloud scenes, ocean and sky scenes, and terrain scenes. Targets are marked with red ellipses to enhance visibility. Detection results are displayed... Figure 2 The second line indicates that all targets were detected accurately.
[0098] 2) Noise Scenarios: Since the detection process is easily affected by various noise interferences from sensors and environmental clutter, robustness to noise is crucial for target detection methods. In this invention, Gaussian noise and stripe noise with a standard deviation of 20 are added to typical background clutter, and the target detection results are as follows: Figure 3 As shown in the second and fourth rows, it can be concluded that Gaussian noise and fringe noise are clearly suppressed by the proposed method.
[0099] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
[0100] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0101] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting weak infrared targets based on adaptive spatiotemporal tensor and weighted tensor average rank approximation, characterized in that, Perform the following steps: S1. Obtain infrared sequence variation data from the server and establish an ASTIT model. The SSIM is defined as follows: Where x and y represent two images, μ x ,μ y ,σ x ,σ y Let x and y represent the mean and variance of the image, respectively. Coefficient C1 = (K1H) 2 C2 = (K2H) 2 K1 and K2 are set to K1 = 0.01, K2 = 0.03, and H represents the dynamic range of the pixel, which is 255 for infrared images; S2. Establish the WPTAR-ASTIT model; For the infrared subtensor D∈R m×n×L It is decomposed into low-rank background and sparse target components: S3. Optimize the solution process; Will Simplified to: Where D, B, T ∈ R m×n×L These represent the original infrared tensor, background tensor, and target tensor, respectively; L represents the time dimension; a k Let represent the weighting coefficients of different transpose components, and satisfy . p is the key parameter of the Schatten p-norm; S4. Solve the data to obtain the results.
2. The infrared weak target detection method based on adaptive spatiotemporal tensor and weighted tensor average rank approximation as described in claim 1, characterized in that, S21. Constructing the ASTIT structure: This involves combining the SSIM metric to describe the degree of change in background clutter, using the original input infrared image sequence O1,...,O P Adaptively converted into multiple infrared subtensors D∈R m×n×L This ensures that the low-rank prior conditions in each subtensor are satisfied.
3. The infrared weak target detection method based on adaptive spatiotemporal tensor and weighted tensor average rank approximation as described in claim 2, characterized in that, In step S2, L in the proposed model is dynamically adjusted based on data characteristics, rather than being a pre-fixed value.
4. The infrared weak target detection method based on adaptive spatiotemporal tensor and weighted tensor average rank approximation as described in claim 3, characterized in that, S22. Decompose background and target components: The transformed subtensor is decomposed into low-rank background components and sparse target components using the WPTAR-ATSIT model.
5. The infrared weak target detection method based on adaptive spatiotemporal tensor and weighted tensor average rank approximation as described in claim 4, characterized in that, S23. Reconstruct the target image: Reconstruct the target image from the obtained target tensor through inverse operations.
6. The infrared weak target detection method based on adaptive spatiotemporal tensor and weighted tensor average rank approximation as described in claim 5, characterized in that, S24. Separate the real target: Further separate the real target using an adaptive threshold.
Citation Information
Patent Citations
Thermal infrared small target detection method based on non-overlapping block space-time tensor model
CN115690381A
Learning a truncation rank of singular value decomposed matrices representing weight tensors in neural networks
US20190332941A1