Infrared small target detection method and system based on deep guided low-rank sparse decomposition

By employing a deep-guided low-rank sparse decomposition method, combined with a hierarchical nonlinear tensor ring background module and a fusion attention mechanism, the accuracy and real-time performance issues of infrared small target detection in complex backgrounds are addressed, achieving efficient and accurate small target detection.

CN120747481BActive Publication Date: 2025-12-12SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511149408.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-12-12
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing infrared small target detection technologies suffer from low detection accuracy in complex backgrounds, are susceptible to interference with the background tensor structure, have high computational overhead due to Gaussian outlier decomposition, and lack sufficient generalization ability of sparse priors, making it difficult to accurately characterize nonlinear target features.

Method used

We employ a deep-guided low-rank sparse decomposition method, which combines a hierarchical nonlinear tensor ring background module and a fusion attention mechanism with a deep neural network to perform low-rank sparse tensor decomposition, thereby improving the nonlinear representation capability of complex backgrounds. Furthermore, we enhance feature fusion and temporal information interaction through a 3D convolution fusion module.

Benefits of technology

It significantly improves the accuracy and robustness of infrared small target detection, effectively handles complex backgrounds, and enhances the detection accuracy and real-time performance of small targets. It is applicable to fields such as military, security, environmental monitoring, medical diagnosis, and industrial monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747481B_ABST
    Figure CN120747481B_ABST
Patent Text Reader

Abstract

The application provides an infrared small target detection method and system based on depth-guided low-rank sparse decomposition, and relates to the technical field of remote sensing image processing. The method comprises the following steps: acquiring all original infrared images shot by a remote sensing device, and stacking the original infrared images in order of acquisition time to obtain an infrared original tensor; performing low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model; performing processing through a constructed hierarchical nonlinear tensor ring background module to obtain a low-rank background tensor containing nonlinear transformation; performing processing through a sparse target module with a fusion attention mechanism to obtain a sparse feature tensor of an infrared small target region; reconstructing a depth neural network guided low-rank sparse tensor decomposition model and performing solving processing to obtain a final infrared small target detection result. The application can accurately, robustly and quickly detect small targets in a complex background.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing technology, and more specifically, to an infrared small target detection method and system based on depth-guided low-rank sparse decomposition. Background Technology

[0002] In recent years, infrared small target detection technology has received widespread attention in various fields such as military, security, environmental monitoring, medical diagnosis, and industrial surveillance. Compared with traditional visible light imaging technology, infrared imaging has the advantages of strong anti-interference capability and all-weather operation, making it particularly suitable for target detection tasks in complex environments. However, infrared small targets are usually characterized by missing texture information, blurred outlines, and extremely low pixel ratio. In addition, the complex and varied background scenes make it easy for targets to be submerged in the background or obscured by structurally similar interference information in the image, thus posing a great challenge to their accurate detection.

[0003] In existing technologies, low-rank sparse tensor decomposition models have been widely applied to infrared small target detection tasks due to their good mathematical interpretability and structural separation capabilities. However, tensor decomposition methods still have certain limitations in practical applications: firstly, the background tensor structure is easily affected by interference and destroyed, reducing separation accuracy; secondly, Gaussian heterogeneous decomposition has a large computational cost, limiting the model's real-time performance and scalability; and thirdly, it relies on... The sparse prior constructed by norm still has insufficient generalization ability when facing diverse target distributions, and it is difficult to accurately characterize the features of nonlinear targets.

[0004] Therefore, there is an urgent need for an infrared small target detection method and system based on depth-guided low-rank sparse decomposition to solve the above problems. Summary of the Invention

[0005] The purpose of this invention is to provide an infrared small target detection method and system based on depth-guided low-rank sparse decomposition, so as to improve the above-mentioned problems. To achieve the above objective, the technical solution adopted by this invention is as follows:

[0006] Firstly, this application provides an infrared small target detection method based on depth-guided low-rank sparse decomposition, including:

[0007] All raw infrared images captured by the remote sensing device are acquired and stacked sequentially according to the acquisition time to obtain the raw infrared tensor.

[0008] Based on the original infrared tensor, low-rank background and sparse target decomposition processing is performed to obtain a low-rank sparse tensor decomposition model.

[0009] Based on the low-rank sparse tensor decomposition model, a low-rank background tensor containing nonlinear transformations is obtained by constructing a hierarchical nonlinear tensor ring background module.

[0010] The third processing unit is configured to obtain a sparse feature tensor of an infrared small target region based on the low-rank background tensor and through sparse target module processing with a fusion attention mechanism.

[0011] The fourth processing unit is configured to reconstruct a deep neural network guided low-rank sparse tensor decomposition model based on the low-rank sparse tensor decomposition model, the low-rank background tensor with nonlinear transformation and the sparse feature tensor of the infrared small target region, and perform solving processing to obtain a final infrared small target detection result.

[0012] In a second aspect, the present application further provides an infrared small target detection system based on deep guided low-rank sparse decomposition, comprising:

[0013] The acquisition unit is configured to acquire all original infrared images photographed by a remote sensing device, and stack the original infrared images in order of acquisition time to obtain an infrared original tensor.

[0014] The first processing unit is configured to perform low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model.

[0015] The second processing unit is configured to obtain a low-rank background tensor with nonlinear transformation based on the low-rank sparse tensor decomposition model and through hierarchical nonlinear tensor ring background module processing.

[0016] The third processing unit is configured to obtain a sparse feature tensor of an infrared small target region based on the low-rank background tensor and through sparse target module processing with a fusion attention mechanism.

[0017] The fourth processing unit is configured to reconstruct a deep neural network guided low-rank sparse tensor decomposition model based on the low-rank sparse tensor decomposition model, the low-rank background tensor with nonlinear transformation and the sparse feature tensor of the infrared small target region, and perform solving processing to obtain a final infrared small target detection result.

[0018] The present application has the following beneficial effects:

[0019] The application provides a deep neural network guided low-rank sparse decomposition model for infrared small target detection, which improves the detection accuracy and robustness of the infrared small target by combining model driven prior domain knowledge and data driven learning ability.

[0020] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art upon examination of the following or can be learned by practice of the present application. The objects and other advantages of the application can be realized and attained by means of the instrumentalities particularly pointed out in the written description and claims hereof as well as the appended drawings. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be considered as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0022] Figure 1 A flowchart of the infrared small target detection method based on deep guided low-rank sparse decomposition described in the embodiments of the present application;

[0023] Figure 2 A structure diagram of the infrared small target detection system based on deep guided low-rank sparse decomposition described in the embodiments of the present application.

[0024] In the figure: 701, acquisition unit; 702, first processing unit; 703, second processing unit; 704, third processing unit; 705, fourth processing unit. DETAILED DESCRIPTION

[0025] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present application.

[0026] It should be noted that similar reference numerals and letters represent similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Meanwhile, in the description of the present application, the terms "first", "second", and the like are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0027] Embodiment 1

[0028] The embodiment provides an infrared small target detection method based on deep guided low-rank sparse decomposition.

[0029] Referring to Figure 1 , the method includes steps S1, S2, S3, S4 and S5.

[0030] In step S1, all original infrared images captured by a remote sensing device are acquired, and the original infrared images are stacked in time sequence to obtain an infrared original tensor;

[0031] It can be understood that the original infrared images captured continuously by the remote sensing device are acquired in this step, and are stacked in time sequence, so as to construct an infrared original tensor with consistent spatial and time dimensions. This process is not only the starting point of the entire detection chain, but also provides high-dimensional and structured data input for subsequent modeling. Compared with traditional single-frame image processing, tensor modeling can simultaneously retain the continuity and correlation of images in space (such as target position) and time (such as target moving track), which is helpful to enhance the distinguishability of weak features of small targets in the background. Specifically, the "stacking" operation essentially constructs an infrared original tensor wherein, and is the image length and width, The number of stacked frames is. The tensor structure allows the dynamic changes of the target and the time-varying characteristics of the background to be modeled as a whole without destroying the continuity of the image sequence. This design is particularly suitable for target detection in complex backgrounds (such as sea surface, cloud layer, thermal interference, etc.), because the consistency of small targets in time sequence is often more stable than single-frame features. In practical applications, such as border monitoring, unmanned aerial vehicle reconnaissance or air early warning, small targets are often hidden in complex backgrounds, and relying on single-frame images is easy to miss detection. Therefore, the infrared raw tensor formed by stacking the time sequence not only lays a solid foundation for subsequent low-rank sparse modeling, but also significantly enhances the detection ability of motion-consistent targets. In terms of technical effects, this step significantly improves the structural expression ability of data, providing more comprehensive input information for the subsequent deep modeling process, which is conducive to more effective separation of background and target.

[0032] Step S2, performing low-rank background and sparse target decomposition processing based on the infrared raw tensor to obtain a low-rank sparse tensor decomposition model;

[0033] It can be understood that the background in the original infrared image often exhibits a low-rank structure due to its stability, repeatability and strong structure. Small targets have high locality, low energy and irregularity, and therefore can be regarded as sparse disturbances in global data. Based on this priori, a low-rank sparse tensor decomposition model is constructed. In practical application scenarios, this step can effectively suppress low-frequency interference (such as sky, ground thermal radiation, etc.) in the background while preserving target features, even if the target pixels only account for a very small proportion of the image.

[0034] wherein the infrared raw tensor can be decomposed into a linear combination of a low-rank background and a sparse target, and a low-rank sparse tensor decomposition model is constructed, and the formula is as follows:

[0035] ;

[0036] ;

[0037] wherein, is the infrared raw tensor, is the background tensor, is the target tensor, the tensor nuclear norm of the background is regularized, the sparsity constraint of the target is, and are model balancing parameters.

[0038] Step S3, based on the low-rank sparse tensor decomposition model, processing through a constructed hierarchical nonlinear tensor ring background module to obtain a low-rank background tensor containing nonlinear transformation;

[0039] It can be understood that this step is based on the low-rank sparse tensor decomposition model obtained in the previous stage, and introduces an innovative "hierarchical nonlinear tensor ring background module" to further model the background tensor, so as to generate a low-rank background tensor with nonlinear transformation capability. This step significantly enhances the adaptability to complex background. By introducing nonlinear expression and tensor ring structure, this module not only better preserves the overall structural features of the image, but also effectively filters out small target similar structure background information that may cause missed detection or false detection, providing a purer and more structured background reference for the next step of sparse target extraction, thereby improving the overall detection accuracy and robustness. In this step, step S3 includes step S31, step S32, step S33 and step S34.

[0040] Step S31, initializing a third-order factor tensor for the low-rank sparse tensor decomposition model, wherein a third-order factor tensor is obtained through a pre-defined tensor ring rank and an initialization process of the third-order factor tensor;

[0041] It can be understood that in this step, the third-order factor tensor is initialized , wherein is a pre-defined tensor ring rank. In this step, the tensor ring is an efficient tensor network structure that represents a high-dimensional tensor as a ring-like product sequence of multiple third-order tensor factors, so that the dimension of each factor tensor is This initialization process provides an optimizable basic representation structure for the subsequent nonlinear forward transformation network, so that the network does not have to operate directly on the high-dimensional tensor, but models on the low-dimensional factor, significantly reducing the computational overhead. In addition, since the decomposition form of the tensor ring itself has compression properties, it can effectively suppress high-frequency noise and small target interference while preserving the main structural information of the background tensor, laying a good foundation for the robustness and compressibility of background modeling.

[0042] Step S32, transforming the third-order factor tensor through a nonlinear forward transformation network, wherein a transformed latent factor tensor is obtained through a nonlinear transformation of a combination of a multilayer perceptron and a sine activation function;

[0043] It can be understood that in this step, a nonlinear forward transformation network of a combination of a multilayer perceptron and a sine activation function is constructed and an inverse transformation network as follows, wherein the forward transformation network The formula is:

[0044] ;

[0045] ;

[0046] wherein, a nonlinear forward transformation network composed of a multilayer perceptron and a sine activation function, a network block combination layer, denotes a nonlinear layer of the i-th layer, a nonlinear layer of the i-th layer, a nonlinear layer of the i-th -1 layer, an input output of the i-th layer, an input of the i-th +1 layer, an output of the i-th layer, an input tensor is multiplied by a weight matrix in the third dimension, is a learnable bias term.

[0047] an inverse transformation network is given by:

[0048]

[0049]

[0050] wherein, a nonlinear forward transformation network composed of a multilayer perceptron and a sine activation function, a network block combination layer, denotes a nonlinear layer of the i-th layer.

[0051] This step generalizes the traditional low-rank modeling from a static linear structure to a dynamic nonlinear structure through a nonlinear forward transformation network, so that the background modeling has the ability to dynamically learn complex background patterns from the original factor level. Especially in practical applications, when the target is submerged in a background area with similar structure, this nonlinear processing can make the boundary between the background and the target more clear, reduce the interference of false targets, and provide a high-quality and information-rich latent representation tensor for the subsequent singular value constraint and target extraction process, thereby improving the accuracy and robustness of the final small target detection.

[0052] In step S33, the transformed latent factor tensor is subjected to tensor singular value decomposition processing, wherein the tensor singular value decomposition induced by the nonlinear transformation performs tensor kernel norm constraint on the factor tensor to obtain a constrained factor tensor;

[0053] ​​It is understandable that this step utilizes tensor singular value decomposition induced by nonlinear transformation to impose tensor nuclear norm constraints on each factor tensor, thereby enhancing the low-rank subspace representation of local factors. The formula is as follows:

[0054] ;

[0055] in, For the first A third-order factor tensor It is an inverse transform network. For positive transform network, Tensor nuclear norm regularization of the background It is the first latent factor tensor after positive transformation. A frontal slice, Nuclear norm regularization with background The size of the front slice dimension. The first factor tensor Frontal slice, For slices The A singular value, The number of non-zero singular values ​​in the slice. For the first A singular value.

[0056] Step S34: Perform low-rank background tensor reconstruction processing on the constrained factor tensor, wherein a compact low-rank background tensor is reconstructed by tensor ring decomposition to obtain the low-rank background tensor.

[0057] Understandably, this step utilizes tensor ring decomposition to reconstruct a compact, nonlinear, low-rank background tensor, the formula of which is:

[0058] ;

[0059] in, This represents the value of the optimal solution to the formula. It is a hierarchical nonlinear tensor ring decomposition module. These are the parameters of the forward and inverse transform network. These are the model equilibrium parameters.

[0060] It is understandable that this step achieves efficient restoration of background information through tensor ring reconstruction, which improves the fitting ability to nonlinear scenes while maintaining low-rank compression performance, providing a solid and reliable foundation for subsequent sparse target extraction, thereby improving the robustness of the entire system to complex background interference and the sensitivity of target detection.

[0061] Step S4: Based on the low-rank background tensor, the sparse target module with fusion attention mechanism is processed to obtain the sparse feature tensor of the infrared small target region.

[0062] It can be understood that this step first performs difference operation on the original infrared tensor and the reconstructed low-rank background tensor to generate a sparse tensor. The difference operation can preliminarily remove the background principal component and highlight the small target region, especially the point target that presents local fluctuation or motion characteristics in the time sequence. Since the background tensor has been highly compressed and purified, the difference should theoretically mainly retain the structural information related to the small target.

[0063] Next, the system introduces a sparse target module with a fusion attention mechanism. The module is composed of two sets of parallel but complementary sub-modules: a multi-scale white attention module and a multi-scale black attention module. The former constructs a response mechanism sensitive to white bright target regions through maximum pooling, minimum pooling, normalization and convolution operations, and is particularly suitable for detecting high-temperature small targets (such as aircraft, engines, etc. in infrared images); the latter suppresses dark background features to guide the model to further suppress residual background noise or false target interference from the sparse tensor. This two-way attention design can effectively cover the complex situations where there is strong brightness contrast or inversion between targets and backgrounds in infrared images.

[0064] Finally, the two sets of attention response maps are sent to a 3D convolution fusion module, which fuses time information and spatial scale features, and preserves key sparse target features through skip connection and normalization strategy, and outputs a fused third feature map. This map is finally further standardized to form a sparse feature tensor representing the infrared small target region.

[0065] In this step, step S4 includes step S41, step S42, step S43 and step S44.

[0066] Step S41, processing according to the infrared original tensor and the low-rank background tensor, wherein the sparse tensor is obtained by performing tensor difference calculation and noise suppression processing;

[0067] It can be understood that this step performs tensor-level difference calculation, that is, for each time frame and its corresponding pixel point, the difference between the original tensor and the low-rank background tensor is calculated. In this process, the main structural information in the background should be close to zero after the difference because it has been modeled in the low-rank decomposition, while the small target shows a high residual response after the difference because it is not completely fitted by the background model.

[0068] However, the difference operation itself is also prone to retain non-target information such as partial high-frequency noise and local thermal disturbance, so the system further introduces a noise suppression processing mechanism. This processing can use spatial filtering (such as wavelet threshold filtering, non-local mean filtering), time domain smoothing, or statistical threshold-based methods (such as sparse threshold suppression based on standard deviation) to remove noise signals that are not continuous in time or space, have too small energy, or have meaningless patterns. This process aims to eliminate non-target "false strong responses" from sparse responses to avoid subsequent false detections.

[0069] Step S42, the sparse tensor is processed by a first multi-scale white attention module composed of a combination of maximum pooling, minimum pooling, activation function, normalization and convolution operation, wherein a first output feature map is obtained by enhancing the response of the bright target region;

[0070] It can be understood that this step integrates a variety of classical deep vision operations, including maximum pooling, minimum pooling, activation function (such as ReLU or Leaky ReLU), normalization (such as BatchNorm or LayerNorm), and convolution operation. These operations are combined in a specific order to form an attention enhancement path with scale perception capability: maximum pooling: emphasizes the position with the strongest response in the sparse tensor, effectively preserving the main peak features of the bright target; minimum pooling: to a certain extent, it plays a role in edge preservation and local contrast enhancement, complementing the maximum pooling; activation function: enhances the nonlinear feature response, making the model more sensitive to small target edges and local changes; normalization operation: improves the numerical stability of features in each channel or spatial position, and suppresses abnormal fluctuations; convolution operation: extracts local spatial context, enhances the coherence and shape integrity of the target region.

[0071] In addition, the introduction of multi-scale processing is an important feature of this module. The system performs the above operations on multiple scales (such as different sizes of sliding windows or different step lengths of convolution) to capture bright targets with large size differences at the same time. This structural design can cover different scales from pixel-level hotspots to medium target regions, ensuring that small targets are not ignored due to low resolution or weak features.

[0072] Step S43, the sparse tensor is processed by a second multi-scale black attention module composed of a combination of minimum pooling, maximum pooling, activation function, normalization and convolution operation, wherein a second output feature map is obtained by suppressing background interference information;

[0073] It can be understood that this step further suppresses background interference information from the sparse map, especially those misleading areas with local high response but not belonging to real targets. This module design is complementary to the white attention module in the previous step: the former focuses on enhancing bright targets, while this step focuses on suppressing false bright spots and noise structures in the background that may be mistaken for targets, resulting in a cleaner and more refined second output feature map.

[0074] The core operation sequence of this black attention module is similar to that of white attention, but its strategy is more biased towards information suppression and noise filtering. The specific operations include: minimum pooling first: preferentially extracting local minimum values, so that the model pays more attention to the stability of dark areas in the image, which helps to determine which areas have lower response but have stable background features; maximum pooling assistance: auxiliary identification of local activation peaks, and then local contrast enhancement with the minimum pooling result, so as to better identify "false targets"; activation function: for nonlinear transformation, to suppress the amplification effect of edge areas on the overall response; normalization operation: to reduce the numerical deviation between different scales, making the attention response more smooth and consistent; convolution operation: to introduce local context, so that the model can judge whether the area is an interference item from the aspects of texture continuity and spatial consistency.

[0075] Step S44, the first output feature map and the second output feature map are processed by a three-dimensional convolution fusion module composed of three-dimensional convolution, activation function and normalization operation, a third feature map is obtained after fusion by introducing a skip connection mechanism and a feature normalization strategy, and a sparse feature tensor of an infrared small target area is obtained after the third feature map after fusion is subjected to skip connection and normalization operation.

[0076] It can be understood that first, a three-dimensional convolution (3D Convolution) operation is performed, which is different from a two-dimensional convolution that can only process spatial local information. 3D convolution can capture interactive features between time, space and scale at the same time, which has great advantages for infrared small targets - especially when the target has short-term movement or disappears / reappears in time series, 3D convolution can learn its temporal dynamic changes, significantly improving the model's ability to recognize "non-stable small targets".

[0077] Immediately after, an activation function (such as ReLU or GELU) is used to enhance the non-linear expression ability, so that the model can more accurately distinguish target boundaries from residual noise. And the normalization operation (such as Batch Normalization or Layer Normalization) improves the balance of feature distribution, avoiding the problem of unstable training process caused by too large difference in feature amplitude, especially when multiple scale attention outputs are fused.

[0078] Further, the system introduces a skip connection mechanism, which is a strategy derived from deep network structures such as ResNet, which serves to retain original or low-level features from shallow networks, so that important details of small targets in local space are not lost in the fusion process. Skip connection can also enhance gradient flow and improve training efficiency, especially for restoring details of sparse targets.

[0079] Finally, the "fused third feature map" obtained through the above fusion operation is again processed through a skip connection and a normalization operation, and output as the final infrared small target region sparse feature tensor. This tensor has the following key advantages: target region response is significant, background response is strongly suppressed, and spatiotemporal structure is clear, which is the direct basis for subsequent small target deconstruction and accurate positioning of the system.

[0080] Step S5, based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing a nonlinear transformation, and the sparse feature tensor of the infrared small target region, reconstructing the deep neural network guided low-rank sparse tensor decomposition model and performing solving processing to obtain the final infrared small target detection result.

[0081] It can be understood that this step organically combines structured modeling (low-rank sparse decomposition) with the nonlinear expression ability of deep learning, greatly enhancing the model's generalization ability to complex scenes and fine-grained perception ability to small target features without sacrificing mathematical interpretability. This strategy effectively solves the problems of insufficient modeling of complex backgrounds, insufficient description of target diversity, slow optimization convergence, and other problems in traditional tensor decomposition methods, and is a key step to achieve accurate infrared small target detection. The infrared small target detection result output by the model has the advantages of accurate positioning, clear response, and strong interference suppression, and is suitable for various practical scenarios such as air-ground reconnaissance, border monitoring, and sea surface target identification. In this step, step S5 includes steps S51, S52, S53, and S54.

[0082] Step S51, based on the low-rank background tensor, reconstructing a deep neural network guided low-rank sparse tensor decomposition model through a sparse target module with a fusion attention mechanism;

[0083] It can be understood that this step takes the low-rank background tensor extracted in the previous stage as the input basis, and combines a sparse target module with a fusion attention mechanism to jointly construct a new low-rank sparse tensor decomposition model with deep structure guidance capability. The goal of this step is to integrate the traditional low-rank sparse decomposition idea based on priori with the data-driven ability of deep neural networks, and improve the modeling performance of the model on the differences between small targets and complex backgrounds.

[0084] Specifically, the new model is as follows:

[0085] ;

[0086] ;

[0087] in, For the target tensor, For the first A factor tensor, For the original infrared tensor, It is the background tensor. Tensor nuclear norm regularization of the background For the sparse constraint of the objective, Reconstruct constraints for the target tensor and the background tensor. and These are the model equilibrium parameters. For attention-guided sparse target modules, and For learnable parameters in the background and target modules, These are the model equilibrium parameters. This is a hierarchical nonlinear tensor ring decomposition module. These are constraints.

[0088] Step S52: Iteratively optimize the variable subproblems in the low-rank sparse tensor decomposition model guided by the deep neural network based on the alternating direction multiplier method to obtain intermediate optimized solutions that satisfy the constraints.

[0089] It is understandable that this step proposes a hybrid optimization strategy of ADMM and Adam to solve the model. The specific steps are as follows:

[0090] Rewrite the low-rank sparse tensor decomposition model guided by deep neural networks and introduce auxiliary variables. The augmented Lagrangian function of the new model is as follows:

[0091] ;

[0092] in, It is a Lagrange multiplier. It is a penalty parameter. This is a hierarchical nonlinear tensor ring decomposition module. For the first A factor tensor, For the original infrared tensor, Tensor nuclear norm regularization of the background For the sparse constraint of the objective, Reconstruct constraints for the target tensor and the background tensor. and For learnable parameters in the background and target modules, , and are model balancing parameters.

[0093] Step 5.2.1: For the , , subproblem, fix other variables at their t-th iteration values, update the learnable network parameters in the target module and background module using Adam optimizer , , , the loss function is as follows:

[0094] ;

[0095] ;

[0096] ;

[0097] ;

[0098] where denotes the tensor nuclear norm minimization loss, denotes the background reconstruction loss, denotes the background tensor and target tensor separation loss.

[0099] For the subproblem, fix other variables at their t-th iteration values, update the variable at its (t+1)-th iteration, whose formula is:

[0100] ;

[0101] ;

[0102] where denotes the tensor soft threshold operation, is the threshold value, denotes the element sign, is the tensor element value with index , , , is the absolute value symbol.

[0103] For the subproblem, fix other variables at their t-th iteration values, update the variable at its (t+1)-th iteration, whose formula is:

[0104] ;

[0105] where denotes a Lagrange multiplier variable, denotes a penalty parameter, denotes an auxiliary variable, denotes an original observation tensor, denotes a background tensor.

[0106] Step S53, based on the intermediate optimization solution satisfying the constraints, the Adam optimizer driven parameter update process is performed for the target module and the background module in the model, and the optimized learnable network parameters are obtained;

[0107] It can be understood that the "intermediate optimization solution" relied on in this step comes from the background and target tensor approximate solution obtained based on the alternating direction multiplier method (ADMM) processing in the previous stage, which has a preliminary target and background distinguishing ability under the constraint conditions (such as low rank, sparsity, residual balance). At this time, the system takes these intermediate results as a supervision signal, combines the input infrared tensor, and optimizes the parameters in each sub-module (such as the tensor ring nonlinear background modeling network and the multi-scale attention guided sparse target module) in the model based on the loss function.

[0108] Wherein, the stop condition of the above iteration is that when the relative error is less than the minimum iteration error, the iteration is stopped; when the number of non-zero of the target tensor no longer changes, the iteration is stopped; when the maximum iteration number , the iteration is stopped; the target tensor is obtained, the infrared sequence image is obtained by deconstructing the target tensor, which is the final target result image.

[0109] Step S54, the infrared original tensor and the optimized learnable network parameters are input into the deep neural network guided low-rank sparse tensor decomposition model for processing, and the final infrared small target detection result is obtained.

[0110] It can be understood that the initial input infrared original tensor and the learned network parameters optimized by gradient training are input into the constructed deep neural network guided low-rank sparse tensor decomposition model for final joint inference and target extraction processing, and the final detection result of the system is output, that is, the small target region in the infrared image is clearly labeled. This stage not only marks the end of all previous modeling and optimization, but also is a key conversion point from training logic to practical inference of the model. It can not only maintain high detection accuracy, but also ensure inference speed and applicability, and has good practical popularization potential. Whether in real-time infrared monitoring, unmanned aerial vehicle reconnaissance, border early warning or shipborne infrared identification and other application scenarios, the result can provide accurate, robust and fast small target detection support.

[0111] Embodiment 2:

[0112] As Figure 2As shown, this embodiment provides an infrared small target detection system based on depth-guided low-rank sparse decomposition. See [link to documentation]. Figure 2 The system includes an acquisition unit 701, a first processing unit 702, a second processing unit 703, a third processing unit 704, and a fourth processing unit 705.

[0113] Acquisition unit 701 is used to acquire all the original infrared images captured by the remote sensing device, and stack the original infrared images in the order of acquisition time to obtain the original infrared tensor.

[0114] The first processing unit 702 is used to perform low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model.

[0115] The second processing unit 703 is used to obtain a low-rank background tensor containing nonlinear transformations by processing the low-rank sparse tensor decomposition model through the construction of a hierarchical nonlinear tensor ring background module.

[0116] The third processing unit 704 is used to obtain the sparse feature tensor of the infrared small target region by processing the low-rank background tensor through the sparse target module with the fusion attention mechanism.

[0117] The fourth processing unit 705 is used to reconstruct the low-rank sparse tensor decomposition model guided by the deep neural network based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing nonlinear transformation and the sparse feature tensor of the infrared small target region, and to solve the model to obtain the final infrared small target detection result.

[0118] It should be noted that the specific methods by which each module performs operations in the system described in the above embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0119] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0120] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for infrared small target detection based on deep guided low-rank sparse decomposition, characterized in that, The method comprises the following steps: acquire all original infrared images taken by a remote sensing device, and stack the infrared images in sequence according to the acquisition time to obtain an infrared original tensor; perform low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model; based on the low-rank sparse tensor decomposition model, perform hierarchical nonlinear tensor ring background module processing to obtain a low-rank background tensor containing nonlinear transformation; based on the low-rank background tensor, perform sparse target module processing with a fusion attention mechanism to obtain a sparse feature tensor of an infrared small target region; based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing nonlinear transformation, and the sparse feature tensor of the infrared small target region, reconstruct a deep neural network guided low-rank sparse tensor decomposition model and perform solving processing to obtain a final infrared small target detection result; wherein, based on the low-rank sparse tensor decomposition model, performing hierarchical nonlinear tensor ring background module processing comprises: performing initialization third-order factor tensor processing on the low-rank sparse tensor decomposition model, wherein a third-order factor tensor is obtained through a pre-defined tensor ring rank and an initialization process of the third-order factor tensor; performing transformation processing on the third-order factor tensor through a nonlinear positive transformation network, wherein a transformed latent factor tensor is obtained through nonlinear transformation of a combination of a multilayer perceptron and a sine activation function; performing tensor singular value decomposition processing on the transformed latent factor tensor, wherein a constrained factor tensor is obtained through tensor kernel norm constraint of the factor tensor induced by nonlinear transformation through tensor singular value decomposition; performing low-rank background tensor reconstruction processing on the constrained factor tensor, wherein a compact low-rank background tensor is reconstructed through tensor ring decomposition to obtain the low-rank background tensor.

2. The infrared small target detection method based on deep guided low-rank and sparse decomposition according to claim 1, characterized in that performing low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model comprises: constructing a low-rank sparse tensor decomposition model according to the infrared original tensor as a linear combination of low-rank background and sparse target, wherein the formula of the low-rank sparse tensor decomposition model is as follows: ; wherein, is a background tensor, is a target tensor, tensor core norm regularization of the background, sparsity constraint for the target, and are model balancing parameters.

3. The infrared small target detection method based on deep guided low-rank and sparse decomposition according to claim 1, characterized in that based on the low-rank background tensor, performing sparse target module processing with a fusion attention mechanism to obtain a sparse feature tensor of an infrared small target region comprises: processing the infrared original tensor and the low-rank background tensor, wherein a sparse tensor is obtained through tensor difference calculation and noise suppression processing; performing processing on the sparse tensor through a first multi-scale white attention module composed of a combination of maximum pooling, minimum pooling, an activation function, normalization, and convolution operation, wherein a first output feature map is obtained by enhancing the response of bright target regions; performing processing on the sparse tensor through a second multi-scale black attention module composed of a combination of minimum pooling, maximum pooling, an activation function, normalization, and convolution operation, wherein a second output feature map is obtained by suppressing background interference information; The first output feature map and the second output feature map are processed by a three-dimensional convolution fusion module composed of three-dimensional convolution, an activation function and a normalization operation, a jump connection mechanism and a feature normalization strategy are introduced, a third feature map after fusion is obtained, and the third feature map after fusion is subjected to jump connection and normalization operation to obtain a sparse feature tensor of an infrared small target region.

4. The infrared small target detection method based on deep guided low-rank and sparse decomposition according to claim 1, characterized in that Based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing nonlinear transformation and the sparse feature tensor of the infrared small target region, a deep neural network guided low-rank sparse tensor decomposition model is reconstructed and solved, including: Based on the low-rank background tensor, a sparse target module with fusion attention mechanism is used to reconstruct the deep neural network guided low-rank sparse tensor decomposition model; Based on the alternating direction multiplier method, the variable sub-problems in the deep neural network guided low-rank sparse tensor decomposition model are iteratively optimized to obtain an intermediate optimization solution that satisfies the constraints; Based on the intermediate optimization solution that satisfies the constraints, parameter update processing driven by the Adam optimizer is performed on the target module and the background module in the model to obtain optimized learnable network parameters; The infrared original tensor and the optimized learnable network parameters are input into the deep neural network guided low-rank sparse tensor decomposition model for processing to obtain the final infrared small target detection result.

5. An infrared small target detection system based on deep guided low-rank sparse decomposition, characterized in that, It includes: An acquisition unit is configured to acquire all original infrared images captured by a remote sensing device, and stack the infrared images in order of acquisition time to obtain an infrared original tensor; A first processing unit is configured to perform low-rank background and sparse target decomposition processing based on the infrared original tensor to obtain a low-rank sparse tensor decomposition model; A second processing unit is configured to process the low-rank sparse tensor decomposition model through a hierarchical nonlinear tensor ring background module to obtain a low-rank background tensor containing nonlinear transformation; A third processing unit is configured to process the low-rank background tensor through a sparse target module with fusion attention mechanism to obtain a sparse feature tensor of an infrared small target region; A fourth processing unit is configured to reconstruct a deep neural network guided low-rank sparse tensor decomposition model based on the low-rank sparse tensor decomposition model, the low-rank background tensor containing nonlinear transformation and the sparse feature tensor of the infrared small target region, and perform solving processing to obtain the final infrared small target detection result; The second processing unit includes: A second processing subunit is configured to initialize a third-order factor tensor by processing the low-rank sparse tensor decomposition model, wherein a third-order factor tensor is obtained through a pre-defined tensor ring rank and an initialization process of the third-order factor tensor; A third processing subunit is configured to transform the third-order factor tensor through a nonlinear positive transformation network, wherein a transformed latent factor tensor is obtained through nonlinear transformation of a multilayer perceptron and a sine activation function combination; A fourth processing subunit is configured to perform tensor singular value decomposition processing on the transformed latent factor tensor, wherein a constrained factor tensor is obtained by performing tensor kernel norm constraint on the factor tensor through tensor singular value decomposition induced by nonlinear transformation; The fifth processing subunit is configured to perform low-rank background tensor reconstruction processing on the constrained factor tensor, and a compact low-rank background tensor is reconstructed through tensor ring decomposition to obtain a low-rank background tensor.

6. The infrared dim small target detection system based on deep guided low-rank sparse decomposition according to claim 5, characterized in that, The first processing unit comprises: The first processing subunit is configured to construct a low-rank sparse tensor decomposition model according to the infrared original tensor as a linear combination of a low-rank background and a sparse target, and the formula of the low-rank sparse tensor decomposition model is as follows: ; wherein, is a background tensor, is a target tensor, tensor core norm regularization of the background, sparse constraint for the target, and is a model balancing parameter.

7. The infrared dim small target detection system based on deep guided low-rank sparse decomposition according to claim 5, characterized in that, The third processing unit comprises: The sixth processing subunit is configured to process the infrared original tensor and the low-rank background tensor, and a sparse tensor is obtained through tensor difference calculation and noise suppression processing. The seventh processing subunit is configured to process the sparse tensor through a first multi-scale white attention module composed of a maximum pooling operation, a minimum pooling operation, an activation function, a normalization operation and a convolution operation, and a first output feature map is obtained by enhancing the response of a bright target region. The eighth processing subunit is configured to process the sparse tensor through a second multi-scale black attention module composed of a minimum pooling operation, a maximum pooling operation, an activation function, a normalization operation and a convolution operation, and a second output feature map is obtained by suppressing background interference information. The ninth processing subunit is configured to process the first output feature map and the second output feature map through a three-dimensional convolution fusion module composed of a three-dimensional convolution operation, an activation function and a normalization operation, and a third feature map after fusion is obtained by introducing a skip connection mechanism and a feature normalization strategy, and a sparse feature tensor of an infrared small target region is obtained through skip connection and normalization operation on the third feature map after fusion.

8. The infrared dim small target detection system based on deep guided low-rank sparse decomposition according to claim 5, characterized in that, The fourth processing subunit comprises: The first optimization subunit is configured to reconstruct a deep neural network guided low-rank sparse tensor decomposition model through a sparse target module with a fusion attention mechanism based on the low-rank background tensor. The second optimization subunit is configured to perform iterative optimization on a variable sub-problem in the deep neural network guided low-rank sparse tensor decomposition model based on an alternating direction multiplier method to obtain an intermediate optimization solution that satisfies constraints. The third optimization subunit is configured to perform Adam optimizer driven parameter update processing on the target module and the background module in the model based on the intermediate optimization solution that satisfies constraints to obtain optimized learnable network parameters. The fourth optimization subunit is configured to input the infrared original tensor and the optimized learnable network parameters into the deep neural network guided low-rank sparse tensor decomposition model for processing to obtain a final infrared small target detection result.

Citation Information

Patent Citations

  • Image significance object detection method based on multiscale low-rank decomposition and with sensitive structural information

    CN103700091A

  • Infrared image target detection method and device, computing equipment and storage medium

    CN113538296A