Infrared small target detection method and system based on depth-guided low-rank sparse decomposition

Through the deep guided low-rank sparse decomposition method, combined with the hierarchical nonlinear tensor ring background module and the fusion attention mechanism, the accuracy and robustness problems of infrared small target detection in complex backgrounds are solved, and efficient and accurate small target detection is achieved.

CN120747481AActive Publication Date: 2025-10-03SOUTHWEST JIAOTONG UNIV

Patent Information

Application Number
CN202511149408.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-10-03
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing infrared small target detection technology has low detection accuracy under complex backgrounds, the background tensor structure is easily disturbed, the high-dimensional singular value decomposition has high computational overhead, the sparse prior generalization ability is insufficient, and it is difficult to accurately characterize nonlinear target characteristics.

Method used

A depth-guided low-rank sparse decomposition method is adopted. Through the hierarchical nonlinear tensor ring background module and the fusion attention mechanism, combined with a deep neural network, low-rank sparse tensor decomposition is performed to improve detection accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and robustness of infrared small target detection, can effectively handle complex backgrounds, reduce computational overhead, and enhance the ability to characterize nonlinear target features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747481A_ABST
    Figure CN120747481A_ABST
Patent Text Reader

Abstract

The invention provides an infrared small target detection method and system based on depth-guided low-rank sparse decomposition, and relates to the technical field of remote sensing image processing, and the method comprises the steps: obtaining all original infrared images shot by remote sensing equipment, and sequentially stacking the original infrared images according to an obtaining time sequence, and obtaining an infrared original tensor; performing low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model; a low-rank background tensor containing nonlinear transformation is obtained through processing of a constructed hierarchical nonlinear tensor ring background module; processing through a sparse target module fused with an attention mechanism to obtain a sparse feature tensor of the infrared small target area; and reconstructing a low-rank sparse tensor decomposition model guided by the deep neural network, and carrying out solving processing to obtain a final infrared small target detection result. According to the invention, accurate, robust and rapid detection can be carried out on a small target under a complex background.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a method and system for detecting small infrared targets based on depth-guided low-rank sparse decomposition. Background Art

[0002] In recent years, infrared small target detection technology has garnered widespread attention in a variety of fields, including military, security, environmental monitoring, medical diagnosis, and industrial surveillance. Compared to traditional visible light imaging, infrared imaging offers stronger anti-interference capabilities and all-weather operation, making it particularly suitable for target detection in complex environments. However, small infrared targets often lack texture information, have blurred outlines, and have a very low pixel count. Furthermore, these small infrared targets often suffer from complex and ever-changing backgrounds, making them easily obscured by the background or structured interference, posing significant challenges to accurate detection.

[0003] In the existing technology, low-rank sparse tensor decomposition models have been widely used in infrared small target detection tasks due to their good mathematical interpretability and structural separation capabilities. However, the tensor decomposition method still has certain limitations in practical applications: first, the background tensor structure is easily affected by interference and destroyed, reducing the separation accuracy; second, the high-dimensional singular value decomposition has a large computational overhead, which limits the real-time and scalability of the model; third, it relies on The sparse prior constructed by the norm still has insufficient generalization ability when facing diverse target distributions, and it is difficult to accurately characterize nonlinear target characteristics.

[0004] Therefore, there is an urgent need for an infrared small target detection method and system based on depth-guided low-rank sparse decomposition to solve the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for infrared small target detection based on depth-guided low-rank sparse decomposition to improve the above-mentioned problems. To achieve the above-mentioned purpose, the technical solutions adopted by the present invention are as follows: In a first aspect, the present application provides an infrared small target detection method based on depth-guided low-rank sparse decomposition, comprising: Obtain all original infrared images captured by the remote sensing device, and stack the original infrared images in order of acquisition time to obtain an infrared original tensor; Performing low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model; Based on the low-rank sparse tensor decomposition model, a low-rank background tensor including nonlinear transformation is obtained by constructing a hierarchical nonlinear tensor ring background module; Based on the low-rank background tensor, a sparse target module fused with an attention mechanism is used to process the low-rank background tensor to obtain a sparse feature tensor of the infrared small target area. Based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing nonlinear transformation and the sparse feature tensor of the infrared small target area, the low-rank sparse tensor decomposition model guided by the deep neural network is reconstructed and solved to obtain the final infrared small target detection result.

[0006] In a second aspect, the present application also provides an infrared small target detection system based on depth-guided low-rank sparse decomposition, comprising: An acquisition unit is used to acquire all original infrared images captured by the remote sensing device and stack the original infrared images in the order of acquisition time to obtain an infrared original tensor; A first processing unit is configured to perform low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model; A second processing unit is configured to obtain a low-rank background tensor including nonlinear transformation by constructing a hierarchical nonlinear tensor ring background module based on the low-rank sparse tensor decomposition model; A third processing unit is configured to obtain a sparse feature tensor of the infrared small target area based on the low-rank background tensor through a sparse target module fused with an attention mechanism; The fourth processing unit is used to reconstruct the low-rank sparse tensor decomposition model guided by the deep neural network based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing nonlinear transformation, and the sparse feature tensor of the infrared small target area, and perform solution processing to obtain the final infrared small target detection result.

[0007] The beneficial effects of the present invention are: The present invention proposes a deep neural network-guided low-rank sparse decomposition model for infrared small target detection, which improves the detection accuracy and robustness of infrared small targets by combining model-driven prior domain knowledge and data-driven learning capabilities. Among them, the hierarchical nonlinear tensor ring background module designed by the present invention avoids the challenges of tensor structure destruction and high-dimensional singular value decomposition by combining the compact structure of tensor ring decomposition and the nonlinear representation capability of tensor singular value decomposition, and enhances the nonlinear low-rank representation of complex backgrounds. Moreover, the multi-scale white / black attention module designed by the present invention can display encoded sparse features and quickly focus on infrared small target areas. At the same time, the 3D convolution fusion module promotes feature fusion and temporal information interaction at different scales, thereby improving the accuracy of small target detection.

[0008] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the embodiments of the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0010] Figure 1 A schematic flow chart of an infrared small target detection method based on depth-guided low-rank sparse decomposition according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of the infrared small target detection system based on depth-guided low-rank sparse decomposition described in an embodiment of the present invention.

[0011] In the figure: 701, acquisition unit; 702, first processing unit; 703, second processing unit; 704, third processing unit; 705, fourth processing unit. DETAILED DESCRIPTION

[0012] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0013] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.

[0014] Example 1: This embodiment provides an infrared small target detection method based on depth-guided low-rank sparse decomposition.

[0015] See also Figure 1 , the figure shows that the method includes step S1, step S2, step S3, step S4 and step S5.

[0016] Step S1: Obtain all original infrared images captured by the remote sensing device, and stack the original infrared images in the order of acquisition time to obtain an infrared original tensor; It can be understood that this step obtains the original infrared images taken continuously from the remote sensing equipment and stacks them in chronological order to construct an infrared original tensor with consistency in spatial and temporal dimensions. This process is not only the starting point of the entire detection chain, but also provides high-dimensional, structured data input for subsequent modeling. Compared with traditional single-frame image processing, tensor modeling can simultaneously retain the continuity and correlation of the image in space (such as target position) and time (such as target movement trajectory), which helps to enhance the recognizability of small targets with weak features in the background. Specifically, this "stacking" operation actually constructs an infrared original tensor. ,in, and are the image length and width, is the number of stacked frames. This tensor structure allows for the overall modeling of the dynamic changes of the target and the time-varying characteristics of the background without destroying the continuity of the image sequence. This design is particularly suitable for target detection in complex backgrounds (such as sea surface, cloud layer, thermal interference, etc.), because the temporal consistency of small targets is often more stable than single-frame features. In practical applications, such as border monitoring, drone reconnaissance or air warning tasks, small targets are often hidden in complex backgrounds and are easily missed by relying on single-frame images. Therefore, the infrared raw tensor formed by stacking time series not only lays a solid foundation for subsequent low-rank sparse modeling, but also significantly enhances the detection capability of targets with consistent motion. In terms of technical effect, this step significantly improves the structural expression capability of the data, provides more comprehensive input information for the subsequent deep modeling process, and is conducive to more effective separation of background and target.

[0017] Step S2: performing low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model; It's understandable that the background in raw infrared images often exhibits a low-rank structure due to its stability, repetitiveness, and strong structure. Small targets, on the other hand, are highly localized, low-energy, and irregular, and can therefore be viewed as sparse perturbations in the global data. Based on this prior, the system constructs a low-rank sparse tensor decomposition model. In practical applications, this step effectively suppresses low-frequency interference in the background (such as sky and surface thermal radiation) while preserving target features, even if the target pixels only occupy a very small proportion of the image.

[0018] Among them, according to the infrared original tensor It can be decomposed into a linear combination of low-rank background and sparse target, and a low-rank sparse tensor decomposition model is constructed. The formula is as follows: ; ; in, is the infrared raw tensor, is the background tensor, is the target tensor, The tensor nuclear norm regularization of the background, is the sparse constraint of the target, and is the model balance parameter.

[0019] Step S3: Based on the low-rank sparse tensor decomposition model, a hierarchical nonlinear tensor ring background module is constructed to obtain a low-rank background tensor including nonlinear transformation; It can be understood that this step introduces an innovative "hierarchical nonlinear tensor ring background module" to further model the background tensor on the basis of the low-rank sparse tensor decomposition model obtained in the previous stage, so as to generate a low-rank background tensor with nonlinear transformation capabilities. This step significantly enhances the adaptability to complex backgrounds. By introducing nonlinear expression and tensor ring structure, the module can not only better preserve the overall structural characteristics of the image, but also effectively filter out those small target similar structural background information that may lead to missed detection or false detection, providing a purer and more clearly structured background reference for the next step of sparse target extraction, thereby improving the overall detection accuracy and robustness. In this step, step S3 includes step S31, step S32, step S33 and step S34.

[0020] Step S31: Initialize the low-rank sparse tensor decomposition model into a third-order factor tensor, wherein the third-order factor tensor is obtained through a predefined tensor ring rank and a third-order factor tensor initialization process; It can be understood that this step initializes the third-order factor tensor ,in is the predefined tensor ring rank. In this step, the tensor ring is an efficient tensor network structure that represents a high-dimensional tensor as a cyclic product sequence of multiple third-order tensor factors, so that the dimension of each factor tensor is This initialization process provides an optimizable basic representation structure for the subsequent nonlinear forward transformation network, allowing the network to model low-dimensional factors instead of directly operating on high-dimensional tensors, significantly reducing computational overhead. Furthermore, because the decomposition form of the tensor ring itself has compression properties, it can effectively suppress high-frequency noise and small target interference while maintaining the main structural information of the background tensor, laying a good foundation for the robustness and compressibility of background modeling.

[0021] Step S32: performing a transformation process on the third-order factor tensor using a nonlinear forward transformation network, wherein a transformed latent factor tensor is obtained through a nonlinear transformation combining a multilayer perceptron and a sinusoidal activation function; It can be understood that in this step, a nonlinear forward transformation network combining a multi-layer perceptron and a sinusoidal activation function is constructed. and inverse transform network As shown below, the forward transformation network The formula is: ; ; in, To construct a nonlinear forward transformation network combining multi-layer perceptron and sinusoidal activation function, Combining layers for network blocks, express For the The nonlinear layer of the layer, For the The nonlinear layer of the layer, For the -1 nonlinear layer, For input After the output of layer 1, For the +1 layer input, For the The output of the layer, For the input tensor In the third dimension with the weight matrix Do multiplication, is the learnable bias term.

[0022] Inverse Transformation Network The formula is: ; ; in, To construct a nonlinear forward transformation network combining multi-layer perceptron and sinusoidal activation function, Combining layers for network blocks, express For the The nonlinear layer of the layer.

[0023] This step generalizes traditional low-rank modeling from a static linear structure to a dynamic nonlinear structure through a nonlinear forward transformation network, enabling background modeling to dynamically learn complex background patterns from the original factor level. In practical applications, this nonlinear processing can clearly distinguish between background and target, reduce pseudo-target interference, and provide a high-quality, information-rich latent expression tensor for subsequent singular value constraints and target extraction, thereby improving the accuracy and robustness of small target detection.

[0024] Step S33: performing a tensor singular value decomposition process on the transformed latent factor tensor, wherein the factor tensor is subjected to a tensor nuclear norm constraint by the tensor singular value decomposition induced by the nonlinear transformation to obtain a constrained factor tensor; It can be understood that this step uses the tensor singular value decomposition induced by nonlinear transformation to constrain the tensor nuclear norm of each factor tensor to enhance the low-rank subspace representation of the local factor. The formula is: ; in, For the A third-order factor tensor, is the inverse transformation network, is the forward transformation network, The tensor nuclear norm regularization of the background, is the first of the potential factor tensors after positive transformation A frontal slice, is the nuclear norm regularization of the background, is the front slice dimension size, is the factor tensor Frontal slice, For slices No. singular values, is the number of non-zero singular values ​​of the slice, For the Singular values.

[0025] Step S34: performing low-rank background tensor reconstruction processing on the constrained factor tensor, wherein a compact low-rank background tensor is reconstructed by tensor ring decomposition to obtain a low-rank background tensor.

[0026] It can be understood that this step uses tensor ring decomposition to reconstruct a compact nonlinear low-rank background tensor, the formula of which is: ; in, represents the value of the optimal solution of the formula, It is a hierarchical nonlinear tensor ring decomposition module. are the parameters of the forward and inverse transformation network, is the model balance parameter.

[0027] It can be understood that this step achieves efficient restoration of background information through tensor ring reconstruction, improves the fitting ability of nonlinear scenes while maintaining low-rank compression performance, and provides a solid and reliable basis support for subsequent sparse target extraction, thereby improving the robustness of the entire system to complex background interference and the sensitivity of target detection.

[0028] Step S4: Based on the low-rank background tensor, the sparse target module of the fusion attention mechanism is processed to obtain a sparse feature tensor of the infrared small target area; It can be understood that this step first performs a subtraction operation on the original infrared tensor and the reconstructed low-rank background tensor to generate a sparse tensor. This subtraction operation can initially remove the main background components and highlight small target areas, especially those point-like targets that exhibit local fluctuations or motion characteristics in the time series. Because the background tensor is highly compressed and purified, this subtraction should theoretically retain mainly structural information related to small targets.

[0029] Next, the system introduces a sparse object module that incorporates an attention mechanism. This module consists of two parallel but complementary submodules: a multi-scale white attention module and a multi-scale black attention module. The former uses max-pooling, min-pooling, normalization, and convolution operations to develop a response mechanism sensitive to bright white target regions, making it particularly suitable for detecting small, high-temperature targets (such as aircraft and engines in infrared images). The latter suppresses dark background features, guiding the model to further suppress residual background noise or pseudo-target interference from the sparse tensor. This bidirectional attention design effectively addresses complex scenarios in infrared images where there is strong brightness contrast or inversion between the target and background.

[0030] Finally, the two attention response maps are fed into a 3D convolutional fusion module, which fuses temporal information with spatial scale features and retains key sparse target features through skip connections and normalization strategies. The module then outputs a fused third feature map, which is then further normalized to form a sparse feature tensor representing the infrared small target region.

[0031] In this step, step S4 includes step S41, step S42, step S43 and step S44.

[0032] Step S41: Processing is performed based on the infrared original tensor and the low-rank background tensor, wherein a sparse tensor is obtained by performing tensor difference calculation and noise suppression processing; It can be understood that this step performs tensor-level difference calculation, that is, for each time frame and its corresponding pixel point, the difference between the original tensor and the low-rank background tensor is calculated. In this process, the main structural information in the background has been modeled in the low-rank decomposition, and theoretically should be close to zero after the difference. However, small targets are not fully fitted by the background model, so they show high residual responses after the difference.

[0033] However, the interpolation operation itself tends to retain some non-target information, such as high-frequency noise and local thermal disturbances. Therefore, the system further incorporates a noise suppression mechanism. This process can employ spatial filtering (such as wavelet threshold filtering or non-local mean filtering), temporal smoothing, or statistical threshold-based methods (such as standard deviation-based sparse threshold suppression) to remove noise signals that are discontinuous in time or space, have low energy, or have meaningless patterns. This process aims to eliminate non-target "pseudo-strong responses" from the sparse response, thereby avoiding subsequent false detections.

[0034] Step S42: Processing the sparse tensor with a first multi-scale white attention module composed of a combination of maximum pooling, minimum pooling, activation function, normalization, and convolution operations, wherein a first output feature map is obtained by enhancing the response of bright target areas; It's understandable that this step incorporates several classic deep vision operations, including max pooling, min pooling, activation functions (such as ReLU or Leaky ReLU), normalization (such as BatchNorm or LayerNorm), and convolution. These operations are combined in a specific order to form a scale-aware attention-enhancing pathway: max pooling emphasizes the locations with the strongest response in the sparse tensor, effectively preserving the main peak features of bright targets; min pooling complements max pooling by preserving edges and enhancing local contrast to a certain extent; activation functions enhance nonlinear feature responses, making the model more sensitive to small target edges and local changes; normalization improves the numerical stability of features across channels or spatial locations, suppressing abnormal fluctuations; and convolution extracts local spatial context, enhancing the coherence and shape integrity of the target area.

[0035] Furthermore, the introduction of multi-scale processing is a key feature of this module. The system performs the aforementioned operations at multiple scales (e.g., sliding windows of varying sizes or convolutions of varying step lengths) to simultaneously capture bright objects of varying sizes. This structural design covers a wide range of scales, from pixel-level hotspots to medium-sized object regions, ensuring that small objects are not overlooked due to low resolution or weak features.

[0036] Step S43: Processing the sparse tensor with a second multi-scale black attention module consisting of a combination of minimum pooling, maximum pooling, activation function, normalization, and convolution operations, wherein a second output feature map is obtained by suppressing background interference information; It's understandable that this step further suppresses background interference information from the sparse map, particularly misleading regions with localized high responses that don't represent true targets. This module design complements the white attention module in the previous step: while the former focuses on enhancing bright targets, this step focuses on suppressing spurious bright spots and noisy background structures that could be mistaken for targets, resulting in a cleaner and more refined second output feature map.

[0037] The core operation sequence of the black attention module is similar to that of the white attention module, but its strategy is more inclined towards information suppression and noise filtering. Specific operations include: minimum pooling first: prioritizing the extraction of local minimum values, so that the model pays more attention to the stability of dark areas in the image, which helps to determine which areas have low responses but stable background features; maximum pooling auxiliary: assisting in identifying local activation peaks, and then forming local contrast enhancement with the minimum pooling results, so as to better identify "pseudo-targets"; activation function: used for nonlinear transformation to suppress the amplification effect of edge areas on the overall response; normalization operation: reducing the numerical deviation between different scales to make the attention response smoother and more consistent; convolution operation: introducing local context, so that the model can judge whether the area is an interference item based on texture continuity, spatial consistency and other aspects.

[0038] Step S44: Process the first output feature map and the second output feature map through a three-dimensional convolution fusion module consisting of a three-dimensional convolution, an activation function, and a normalization operation. By introducing a jump connection mechanism and a feature normalization strategy, a fused third feature map is obtained, and the fused third feature map is subjected to a jump connection and normalization operation to obtain a sparse feature tensor of the infrared small target area.

[0039] It can be understood that this step first performs a three-dimensional convolution operation. Unlike two-dimensional convolution, which can only process local spatial information, 3D convolution can simultaneously capture the interactive characteristics between time, space, and scale. This has great advantages for small infrared targets. Especially when the target has short movements or disappears / reappears in time sequence, 3D convolution can learn its temporal dynamic changes, significantly improving the model's recognition ability for "unstable small targets."

[0040] Subsequently, activation functions (such as ReLU or GELU) are used to enhance nonlinear representation capabilities, enabling the model to more accurately distinguish object boundaries from residual noise. Normalization operations (such as Batch Normalization or Layer Normalization) improve the balance of feature distribution, avoiding training instability caused by large differences in feature amplitudes. This is particularly important when fusing multi-scale attention outputs.

[0041] Furthermore, the system introduces a skip connection mechanism, a strategy derived from deep network structures such as ResNet. Its purpose is to preserve the original or low-level features from shallow networks, ensuring that important details of small objects in the local space are not lost during the fusion process. Skip connections also enhance gradient flow and improve training efficiency, making them particularly suitable for restoring details of sparse objects.

[0042] Finally, the resulting "fused third feature map" is further refined through skip connections and normalization, outputting the final sparse feature tensor for the infrared small target region. This tensor offers the following key advantages: a significant target region response, strong suppression of background response, and a clear spatial and temporal structure. It serves as the direct basis for the system's subsequent small target deconstruction and precise positioning.

[0043] Step S5: Based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing nonlinear transformations, and the sparse feature tensor of the infrared small target area, reconstruct the low-rank sparse tensor decomposition model guided by the deep neural network, and perform a solution process to obtain the final infrared small target detection result.

[0044] It can be understood that this step organically integrates structured modeling (low-rank sparse decomposition) with the nonlinear expressive power of deep learning. This significantly enhances the model's generalization capabilities for complex scenes and fine-grained perception of small target features without sacrificing mathematical interpretability. This strategy effectively addresses issues inherent in traditional tensor decomposition methods, such as insufficient modeling of complex backgrounds, insufficient characterization of target diversity, and slow optimization convergence. It is a key step in achieving precise infrared small target detection. The infrared small target detection results output by the model exhibit accurate positioning, clear response, and strong interference rejection, making them suitable for a variety of practical scenarios, including air-to-ground reconnaissance, boundary monitoring, and surface target recognition. In this step, step S5 comprises steps S51, S52, S53, and S54.

[0045] Step S51: Based on the low-rank background tensor, a low-rank sparse tensor decomposition model guided by a deep neural network is reconstructed through a sparse target module fused with an attention mechanism; It can be understood that this step uses the low-rank background tensor extracted in the previous stage as input, combined with the sparse target module integrated with the attention mechanism, to jointly construct a new low-rank sparse tensor decomposition model with deep structure guidance capabilities. The goal of this step is to merge the traditional low-rank sparse decomposition concept based on priors with the data-driven capabilities of deep neural networks, improving the model's performance in modeling the differences between small objects and complex backgrounds.

[0046] Specifically, the new model is as follows: ; ; in, is the target tensor, For the factor tensors, is the infrared raw tensor, is the background tensor, The tensor nuclear norm regularization of the background, is the sparse constraint of the target, Reconstruct constraints for the target tensor and background tensor, and is the model balance parameter, is an attention-guided sparse target module, and are the learnable parameters in the background module and the target module, is the model balance parameter, It is a hierarchical nonlinear tensor ring decomposition module. is a constraint.

[0047] Step S52: iteratively optimize the variable subproblem in the low-rank sparse tensor decomposition model guided by the deep neural network based on the alternating direction multiplier method to obtain an intermediate optimization solution that satisfies the constraints; It can be understood that this step proposes an ADMM and Adam hybrid optimization strategy to solve the model. The specific steps are: Rewrite the deep neural network guided low-rank sparse tensor decomposition model and introduce auxiliary variables , the augmented Lagrangian function of the new model is as follows: ; in, is the Lagrange multiplier, is the penalty parameter, It is a hierarchical nonlinear tensor ring decomposition module. For the factor tensors, is the infrared raw tensor, The tensor nuclear norm regularization of the background, is the sparse constraint of the target, Reconstruct constraints for the target tensor and background tensor, and are the learnable parameters in the background module and the target module, , and is the model balance parameter.

[0048] Step 5.2.1: For , , Subproblem, fix other variables in the tth iteration, and use Adam optimizer to update the learnable network parameters in the target module and background module , , , the loss function is as follows: ; ; ; ; in, represents the minimum loss of tensor nuclear norm, represents the background reconstruction loss, Represents the background tensor and target tensor separation loss.

[0049] for Subproblem, fix other variables in the tth iteration and update the t+1th iteration Variable, its formula is: ; ; in represents a tensor soft thresholding operation, is the threshold, Indicates the sign of the element, is a tensor Index is , , The element value of is the absolute value symbol.

[0050] for Subproblem, fix other variables in the tth iteration and update the t+1th iteration Variable, its formula is: ; in, represents the Lagrange multiplier variable, represents the penalty parameter, represents auxiliary variables, represents the original observation tensor, Represents the background tensor.

[0051] Step S53: performing parameter update processing driven by an Adam optimizer on the target module and background module in the model based on the intermediate optimization solution that satisfies the constraints, to obtain optimized learnable network parameters; It's understandable that the "intermediate optimized solutions" relied upon in this step come from the approximate background and target tensor solutions obtained in the previous stage using the Alternating Direction Method of Multipliers (ADMM). These solutions, subject to constraints such as low rank, sparsity, and residual balance, already possess preliminary target-background discrimination capabilities. The system then uses these intermediate results as supervisory signals, combined with the input infrared tensor, to optimize the parameters of each submodule in the model (such as the tensor ring nonlinear background modeling network and the multi-scale attention-guided sparse target module) using a loss function.

[0052] The above iteration stopping conditions are: when the relative error is less than the minimum iteration error, the iteration stops; when the number of non-zero values ​​of the target tensor no longer changes, the iteration stops; when the maximum number of iterations is reached, the iteration stops. , stop the iteration; get the target tensor, deconstruct the target tensor to get the infrared sequence image, which is the final target result image.

[0053] Step S54: bring the infrared original tensor and the optimized learnable network parameters into a low-rank sparse tensor decomposition model guided by a deep neural network for processing to obtain the final infrared small target detection result.

[0054] It is understandable that the initial input infrared raw tensor, along with the learnable network parameters obtained through gradient training optimization, is input into the constructed deep neural network-guided low-rank sparse tensor decomposition model for final joint reasoning and target extraction processing, producing the system's final detection results—the clearly marked small target areas in the infrared image. This stage not only concludes all previous modeling and optimization, but also marks the key transition point from model training logic to practical reasoning. It maintains high detection accuracy while ensuring reasoning speed and applicability, and has great potential for practical promotion. Whether in real-time infrared monitoring, drone reconnaissance, border warning, or shipborne infrared recognition, this result can provide accurate, robust, and fast support for small target detection.

[0055] Example 2: like Figure 2 As shown, this embodiment provides an infrared small target detection system based on deep guided low rank sparse decomposition, see Figure 2The system includes an acquisition unit 701 , a first processing unit 702 , a second processing unit 703 , a third processing unit 704 and a fourth processing unit 705 .

[0056] An acquisition unit 701 is configured to acquire all original infrared images captured by a remote sensing device and stack the original infrared images in the order of acquisition time to obtain an infrared original tensor; A first processing unit 702 is configured to perform low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model; The second processing unit 703 is configured to obtain a low-rank background tensor including nonlinear transformation by constructing a hierarchical nonlinear tensor ring background module based on the low-rank sparse tensor decomposition model; The third processing unit 704 is configured to obtain a sparse feature tensor of the infrared small target area based on the low-rank background tensor through a sparse target module fused with an attention mechanism; The fourth processing unit 705 is used to reconstruct the low-rank sparse tensor decomposition model guided by the deep neural network based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing nonlinear transformation, and the sparse feature tensor of the infrared small target area, and perform solution processing to obtain the final infrared small target detection result.

[0057] It should be noted that, regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.

[0058] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

[0059] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for infrared small target detection based on depth-guided low-rank sparse decomposition, characterized in that: include: Obtain all original infrared images captured by the remote sensing device, and stack the original infrared images in order of acquisition time to obtain an infrared original tensor; Performing low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model; Based on the low-rank sparse tensor decomposition model, a low-rank background tensor including nonlinear transformation is obtained by constructing a hierarchical nonlinear tensor ring background module; Based on the low-rank background tensor, a sparse target module fused with an attention mechanism is used to process the low-rank background tensor to obtain a sparse feature tensor of the infrared small target area. Based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing nonlinear transformation and the sparse feature tensor of the infrared small target area, the low-rank sparse tensor decomposition model guided by the deep neural network is reconstructed and solved to obtain the final infrared small target detection result.

2. The infrared small target detection method based on depth-guided low-rank sparse decomposition according to claim 1 is characterized in that: A low-rank background and sparse target decomposition process is performed based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model, including: According to the infrared original tensor as a linear combination of low-rank background and sparse target, a low-rank sparse tensor decomposition model is constructed, wherein the formula of the low-rank sparse tensor decomposition model is as follows: ; ; in, is the background tensor, is the target tensor, The tensor nuclear norm regularization of the background, is the sparse constraint of the target, and is the model balance parameter.

3. The infrared small target detection method based on depth-guided low-rank sparse decomposition according to claim 1 is characterized in that: Based on the low-rank sparse tensor decomposition model, a hierarchical nonlinear tensor ring background module is constructed, including: Initializing the low-rank sparse tensor decomposition model into a third-order factor tensor, wherein the third-order factor tensor is obtained through a predefined tensor ring rank and a third-order factor tensor initialization process; The third-order factor tensor is subjected to a transformation process using a nonlinear forward transformation network, wherein a transformed latent factor tensor is obtained by a nonlinear transformation combining a multilayer perceptron and a sinusoidal activation function; Performing a tensor singular value decomposition process on the transformed latent factor tensor, wherein the factor tensor is subjected to a tensor nuclear norm constraint by the tensor singular value decomposition induced by the nonlinear transformation to obtain a constrained factor tensor; The constrained factor tensor is subjected to low-rank background tensor reconstruction processing, wherein a compact low-rank background tensor is reconstructed by tensor ring decomposition to obtain a low-rank background tensor.

4. The infrared small target detection method based on depth-guided low-rank sparse decomposition according to claim 1, characterized in that: Based on the low-rank background tensor, the sparse target module of the fusion attention mechanism is processed to obtain the sparse feature tensor of the infrared small target area, including: Processing is performed based on the infrared original tensor and the low-rank background tensor, wherein a sparse tensor is obtained by performing tensor difference calculation and noise suppression processing; Processing the sparse tensor through a first multi-scale white attention module composed of a combination of max pooling, min pooling, activation function, normalization, and convolution operations, wherein a first output feature map is obtained by enhancing the response of bright object areas; Processing the sparse tensor with a second multi-scale black attention module consisting of a combination of minimum pooling, maximum pooling, activation function, normalization, and convolution operations, wherein a second output feature map is obtained by suppressing background interference information; The first output feature map and the second output feature map are processed by a three-dimensional convolution fusion module consisting of three-dimensional convolution, activation function and normalization operation. By introducing a skip connection mechanism and feature normalization strategy, a fused third feature map is obtained. The fused third feature map is then subjected to skip connection and normalization operations to obtain a sparse feature tensor of the infrared small target area.

5. The infrared small target detection method based on depth-guided low-rank sparse decomposition according to claim 1 is characterized in that: Based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing nonlinear transformations, and the sparse feature tensor of the infrared small target area, a low-rank sparse tensor decomposition model guided by a deep neural network is reconstructed and solved, including: Based on the low-rank background tensor, a sparse target module fused with an attention mechanism reconstructs a low-rank sparse tensor decomposition model guided by a deep neural network; Based on the alternating direction multiplier method, the variable subproblems in the low-rank sparse tensor decomposition model guided by deep neural networks are iteratively optimized to obtain intermediate optimal solutions that meet the constraints. Based on the intermediate optimization solution that satisfies the constraints, an Adam optimizer-driven parameter update process is performed on the target module and the background module in the model to obtain optimized learnable network parameters; The infrared original tensor and the optimized learnable network parameters are brought into a low-rank sparse tensor decomposition model guided by a deep neural network for processing to obtain the final infrared small target detection result.

6. An infrared small target detection system based on depth-guided low-rank sparse decomposition, characterized in that: include: An acquisition unit is used to acquire all original infrared images captured by the remote sensing device and stack the original infrared images in the order of acquisition time to obtain an infrared original tensor; A first processing unit is configured to perform low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model; A second processing unit is configured to obtain a low-rank background tensor including nonlinear transformation by constructing a hierarchical nonlinear tensor ring background module based on the low-rank sparse tensor decomposition model; A third processing unit is configured to obtain a sparse feature tensor of the infrared small target area based on the low-rank background tensor through a sparse target module fused with an attention mechanism; The fourth processing unit is used to reconstruct the low-rank sparse tensor decomposition model guided by the deep neural network based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing nonlinear transformation, and the sparse feature tensor of the infrared small target area, and perform solution processing to obtain the final infrared small target detection result.

7. The infrared small target detection system based on depth-guided low-rank sparse decomposition according to claim 6, characterized in that: The first processing unit includes: The first processing subunit is used to construct a low-rank sparse tensor decomposition model based on the infrared original tensor as a linear combination of the low-rank background and the sparse target, wherein the formula of the low-rank sparse tensor decomposition model is as follows: ; ; in, is the background tensor, is the target tensor, The tensor nuclear norm regularization of the background, is the sparse constraint of the target, and is the model balance parameter.

8. The infrared small target detection system based on depth-guided low-rank sparse decomposition according to claim 6, characterized in that: The second processing unit includes: A second processing subunit is configured to perform initialization of a third-order factor tensor on the low-rank sparse tensor decomposition model, wherein the third-order factor tensor is obtained through a predefined tensor ring rank and a third-order factor tensor initialization process; a third processing subunit, configured to perform a transformation process on the third-order factor tensor using a nonlinear forward transformation network, wherein a transformed latent factor tensor is obtained by a nonlinear transformation combining a multilayer perceptron and a sinusoidal activation function; a fourth processing subunit, configured to perform a tensor singular value decomposition (TSD) on the transformed latent factor tensor, wherein the factor tensor is subjected to a tensor nuclear norm constraint by the tensor singular value decomposition induced by the nonlinear transformation to obtain a constrained factor tensor; The fifth processing sub-unit is used to perform low-rank background tensor reconstruction processing on the constrained factor tensor, wherein a compact low-rank background tensor is reconstructed by tensor ring decomposition to obtain a low-rank background tensor.

9. The infrared small target detection system based on depth-guided low-rank sparse decomposition according to claim 6, characterized in that: The third processing unit includes: a sixth processing subunit, configured to perform processing based on the infrared original tensor and the low-rank background tensor, wherein a sparse tensor is obtained by performing tensor difference calculation and noise suppression processing; a seventh processing subunit, configured to process the sparse tensor through a first multi-scale white attention module composed of a combination of maximum pooling, minimum pooling, an activation function, normalization, and a convolution operation, wherein a first output feature map is obtained by enhancing a response of a bright target area; an eighth processing subunit, configured to process the sparse tensor using a second multi-scale black attention module composed of a combination of minimum pooling, maximum pooling, an activation function, normalization, and a convolution operation, wherein a second output feature map is obtained by suppressing background interference information; The ninth processing sub-unit is used to process the first output feature map and the second output feature map through a three-dimensional convolution fusion module composed of three-dimensional convolution, activation function and normalization operation, and obtain a fused third feature map by introducing a jump connection mechanism and feature normalization strategy. The fused third feature map is subjected to jump connection and normalization operations to obtain a sparse feature tensor of the infrared small target area.

10. The infrared small target detection system based on depth-guided low-rank sparse decomposition according to claim 8, characterized in that: The fourth processing subunit includes: A first optimization subunit is configured to reconstruct a low-rank sparse tensor decomposition model guided by a deep neural network based on the low-rank background tensor through a sparse target module fused with an attention mechanism; The second optimization subunit is used to iteratively optimize the variable subproblems in the low-rank sparse tensor decomposition model guided by the deep neural network based on the alternating direction multiplier method to obtain an intermediate optimization solution that meets the constraints; A third optimization subunit is configured to perform parameter update processing driven by an Adam optimizer on the target module and the background module in the model based on the intermediate optimization solution that satisfies the constraints, so as to obtain optimized learnable network parameters; The fourth optimization subunit is used to bring the infrared original tensor and the optimized learnable network parameters into the low-rank sparse tensor decomposition model guided by the deep neural network for processing to obtain the final infrared small target detection result.

Citation Information

Patent Citations

  • Image significance object detection method based on multiscale low-rank decomposition and with sensitive structural information

    CN103700091A

  • Infrared image target detection method and device, computing equipment and storage medium

    CN113538296A

  • Infrared small target detection method, device and equipment and readable storage medium

    CN117392378A

  • Infrared image small target detection method and system based on transform domain tensor depth expansion

    CN118172543A

Cited By

  • Learnable tensor low-rank enhancement method for weak-illumination near-infrared image

    CN121746266A

  • Rapid direct positioning method and device based on deep structured tensor completion

    CN121937669A

  • Fast direct positioning method and device based on deep structured tensor completion

    CN121937669B

  • Efficient video understanding method based on low-rank key tensor residual decomposition

    CN122137972A