Video SAR depth imaging method and system based on interpretable sparse tensor recovery
The sparse tensor recovery method for video SAR imaging addresses inefficiencies by establishing a tensor model, incorporating low-rank and sparse characteristics, and using deep learning to optimize network parameters, resulting in efficient and adaptable high-performance imaging.
Patent Information
- Application Number
- CN202510471041.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-15
AI Technical Summary
The existing video SAR imaging methods have poor imaging performance under sparse sampling conditions and poor adaptability in complex scenarios, resulting in high frame rate data acquisition and calculation overhead, low imaging efficiency, and insufficient interpretability of deep networks.
Based on the video SAR depth imaging method that can explain sparse tensor recovery, a tensor signal model is established, combined with low-rank sparse characteristics, a sparse reconstruction model is constructed, and a deep imaging network is deconstructed through deep expansion iterative deconstruction, and the network parameters are optimized by backpropagation to achieve high-performance imaging under sparse sampling.
Effectively reduce data redundancy, reduce data acquisition and storage overhead, improve imaging efficiency, is suitable for complex geographic scenarios, and ensure high-performance imaging under sparse sampling conditions.
Smart Images

Figure CN120314941A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of synthetic aperture radar imaging, and particularly to a video SAR depth imaging method and system based on interpretable sparse tensor recovery. Background Art
[0002] Currently, in the related technologies of video SAR imaging, existing imaging methods generally include traditional frame-by-frame imaging methods (e.g., backprojection, range-Doppler algorithm), imaging methods based on compressive sensing, tensor-based video SAR imaging methods, and video SAR imaging methods based on joint low-rank sparse recovery.
[0003] Among them, for traditional frame-by-frame imaging methods, in practical applications, to obtain the dynamic information of the observed scene, video SAR usually requires a high frame rate, which will lead to huge data acquisition, storage, and calculation overheads, restricting its application in practical scenarios.
[0004] For imaging methods based on compressive sensing, they usually use tensor sparse decomposition for sparse reconstruction of video SAR. However, their solution process requires multiple iterations, resulting in limited usage scenarios and low processing efficiency for such methods.
[0005] For tensor-based video SAR imaging methods, although they can utilize the inter-frame correlation of multi-frame SAR data to improve the video SAR imaging performance under sparse sampling conditions, the imaging results of such methods are not only affected by the selection of iteration parameters, but also the iteration process will lead to a decrease in imaging efficiency.
[0006] For video SAR imaging methods based on joint low-rank sparse recovery, although their iteration parameters can be autonomously obtained through learning, which can improve imaging performance and processing efficiency. However, the performance of such methods is restricted by the artificially set low-rank and sparse priors. When the actually observed video SAR data is inconsistent with the prior, the imaging performance will be greatly reduced, unable to meet the actual application requirements. Moreover, the black-box processing of the deep network reduces the interpretability of the imaging process.
[0007] Therefore, how to effectively ensure imaging performance and practicality for complex scenarios under sparse sampling conditions has become an urgent problem to be solved. Summary of the Invention
[0008] To solve the technical problems in the related technologies, the present invention provides a video SAR depth imaging method and system based on interpretable sparse tensor recovery.
[0009] To achieve the above object, the technical solutions adopted by the present invention include:
[0010] According to a first aspect of the present invention, there is provided a video SAR depth imaging method based on interpretable sparse tensor recovery, including the following steps:
[0011] Step S1: Establish a tensor signal model according to the working process of video SAR and the observation model of single-frame SAR signals;
[0012] Step S2: According to the established tensor signal model, combined with the low-rank sparse characteristics of the SAR image video, establish a video SAR sparse reconstruction model;
[0013] Step S3: Iteratively solve the established video SAR sparse reconstruction model;
[0014] Step S4: By deeply unfolding the iterative solution of the video SAR sparse reconstruction model, construct a preliminary video SAR depth imaging network, and perform SAR video image correlation learning and sparse tensor recovery on it;
[0015] Step S5: Through backpropagation for network training, optimize the parameters of the preliminary video SAR depth imaging network to obtain the final video SAR depth imaging network.
[0016] Optionally, the specific steps of step S1 include:
[0017] Step S1-1: According to the signal data of single-frame SAR, establish an observation model of single-frame SAR signals:
[0018] y t = ψAx t + n
[0019] In the formula, y t represents the echo data of the t-th frame of SAR signal, ψ represents the signal downsampling matrix, A represents the echo data observation matrix determined by the SAR system, x t represents the scattering coefficient of the t-th frame of the observed scene, and n represents the observation noise;
[0020] Step S1-2: According to the working process of video SAR, combined with the observation model of single-frame SAR signals, establish a tensor signal model of video SAR observation signals:
[0021]
[0022] In the formula, Y represents the multi-frame echo signals of video SAR, Φ represents the downsampling operator corresponding to the downsampling matrix ψ, ⊙ a,r represents the Hadamard product along the azimuth and range directions, represents the echo generation operator corresponding to the observation matrix A, X represents the SAR image video, and N represents the observation noise of the multi-frame echo signals.
[0023] Optionally, step S2 specifically includes:
[0024] According to the established tensor signal model, combined with the low-rank and sparse characteristics of the SAR image video X, the SAR image video X is decomposed into a low-rank part and a sparse part, and a video SAR sparse reconstruction model is established:
[0025]
[0026] In the formula, represents the minimization optimization problem with respect to X, ‖d·X‖1 represents taking the L1 norm of d·X, d represents the sparse feature tensor, and λ is the regularization coefficient. represents the correlation function for extracting the low-rank property of the SAR image video, and s.t. represents the constraint condition of the minimization optimization problem.
[0027] Optionally, step S3 specifically includes:
[0028] Step S3-1: According to the relationship between the video SAR single-frame signal and the tensor signal model, rewrite the video SAR sparse reconstruction model as:
[0029]
[0030] In the formula, x1 represents the first frame image of the video SAR, x T represents the T-th frame image of the video SAR, x t represents the t-th frame image of the video SAR, t represents the frame number of the video SAR image, and t = 1, …, T. D represents the sparse matrix corresponding to the sparse feature tensor d, and C(x1, …, x T ) represents the correlation function corresponding to , and its expression is:
[0031]
[0032] In the formula, τ represents the frame number of the video SAR image, and x τ represents the τ-th frame image of the video SAR;
[0033] Step S3-2: Using the augmented Lagrangian method, transform the video SAR sparse reconstruction model in step S3-1 into an unconstrained optimization problem, and its augmented Lagrangian function is:
[0034]
[0035] In the formula, L(·) represents the augmented Lagrangian function, u1 represents the introduced auxiliary variable corresponding to x1, and u T represents the introduced auxiliary variable corresponding to x T corresponding, and u t represents the introduced auxiliary variable corresponding to xt The corresponding auxiliary variable, and ρ represents the coefficient of the augmented Lagrangian penalty function;
[0036] Step S3-3: According to the alternating direction multiplier method, transform the unconstrained optimization problem into the iterative solution of two sub-problems for x1,…,x T and u1,…,u T respectively.
[0037] Optionally, the specific steps of step S3-3 include:
[0038] Step S3-3-1: Initialize and set the regularization coefficient λ, the maximum number of iterations K, and the current iteration number k = 0;
[0039] Step S3-3-2: When the iteration number k ≤ K, solve
[0040]
[0041] In the formula, represents the (k + 1)-th iterative solution value of x1, represents the (k + 1)-th iterative solution value of x T respectively.
[0042] According to the approximate alternating inexact minimization method, the solution process of
[0043] Z k+1 = λ c X k softmax β ((X k )) T D T DX k )
[0044] X k+1 / 2 = W b + W a Z k+1
[0045]
[0046] In the formula, Z k+1 represents the (k + 1)-th iterative solution value of Z, λ c represents the weight coefficient of the (k + 1)-th iteration, X k represents the k-th iterative solution value of the intermediate variable X, softmax β (·) represents the soft threshold function with the weighted value of β, (X k ) T represents the transpose of X k and DT denotes the transpose of D, X k+1 / 2 denotes the intermediate variable of the (k + 1)-th iteration, W b and W a is the weight matrix, denotes x t the solution value of the (k + 1)-th iteration of, L ADMM (·) represents the sparse reconstruction process based on the alternating direction method of multipliers;
[0047] Step S3-3-3: According to the least squares method, solve through the following formula
[0048]
[0049] In the formula, denotes u t the solution value of the (k + 1)-th iteration of, denotes u t the solution value of the k-th iteration of, α t denotes the update coefficient of the SAR image of the t-th frame of video, denotes x t the solution value of the (k + 1)-th iteration of;
[0050] Step S3-3-4: Let the iteration number k = k + 1, and repeat Step S3-3-2 to Step S3-3-3 until the iteration solution stop condition is satisfied.
[0051] Optionally, the specific steps of Step S4 include:
[0052] Step S4-1: According to the calculation methods of Z k+1 , X k+1 / 2 and in Step S3-3-2, design a network structure for solving :
[0053] Design a weighted attention network module to solve Z k+1 , the input of this layer of network is the output of the previous layer of network and output to the next layer of network;
[0054] Design a fully connected network module to solve X k+1 / 2 , the input of this layer of network is the output of the previous layer of network and output to the next layer of network;
[0055] Design an alternating direction method of multipliers network module to solve The input of this layer of network is the output of the previous layer of network and output To the next layer network;
[0056] Step S4-2: According to the calculation method in step S3-3-3, design a network structure for solving The input of this layer network is the output of the previous layer network and the output of the previous iteration Output the calculation result of this layer network to the next layer network;
[0057] Step S4-3: Repeat step S4-1 to step S4-2 for K m -1 times, and input the output of the last step S4-2 into step S4-1 to complete the construction of the preliminary video SAR depth imaging network, where K m represents the number of network iteration modules composed of step S4-1 to step S4-2.
[0058] Optionally, the parameters of the weighted attention network module include: the feature dimension is 128, the number of attention heads is 16, the weight coefficient is 0.5, the spatial convolution kernel size is 5, the number of channels is 3, the stride is 1, and the padding number is 4; the temporal convolution kernel size is 3, the number of channels is 3, the stride is 1, and the padding number is 2.
[0059] Optionally, the parameters of the fully connected network module include: the input layer feature dimension is 1024, the intermediate layer feature dimension is 256, and the output feature dimension is 1024.
[0060] Optionally, the parameters of the alternating direction multiplier method network module include: the number of iterations is 6, and the update step size is 0.6.
[0061] According to the second aspect of the present invention, there is also provided a video SAR depth imaging system based on interpretable sparse tensor recovery, which is applied to the video SAR depth imaging method based on interpretable sparse tensor recovery described in any one of the technical solutions in the first aspect of the present invention. The video SAR depth imaging system based on interpretable sparse tensor recovery includes:
[0062] A tensor signal model establishment module, configured to establish a tensor signal model according to the video SAR working process and the observation model of the single-frame SAR signal;
[0063] A video SAR sparse reconstruction model establishment module, configured to establish a video SAR sparse reconstruction model according to the established tensor signal model and in combination with the low-rank sparse characteristics of the SAR image video;
[0064] An iterative solution module, configured to iteratively solve the established video SAR sparse reconstruction model;
[0065] A depth unfolding module, which is used to construct a preliminary video SAR depth imaging network by unfolding the iterative solution of a video SAR sparse reconstruction model, and perform SAR video image correlation learning and sparse tensor recovery on it;
[0066] A network training and optimization module, which is used to train the network through backpropagation and optimize the parameters of the preliminary video SAR depth imaging network to obtain the final video SAR depth imaging network.
[0067] Beneficial effects:
[0068] 1. Through the above technical solutions, first, the tensor signal model established in step S1 of the present invention can model multi-frame SAR echo signals as a tensor structure, breaking through the two-dimensional signal processing framework of traditional frame-by-frame imaging. At the same time, combined with the low-rank sparse joint in step S2 to establish a video SAR sparse reconstruction model, it can utilize the inter-frame correlation (low-rankness) of SAR videos in the time dimension and the scattering sparsity in the spatial dimension to construct a reconstruction model with stronger physical constraints. In this way, data redundancy can be effectively reduced, the repeated calculations caused by high frame rates in traditional frame-by-frame imaging methods can be avoided, and the data acquisition volume and storage overhead can be reduced.
[0069] Second, in step S3 of the present invention, the video SAR sparse reconstruction model is iteratively solved. At the same time, combined with the depth unfolding technology in step S4, the iterative process is mapped into a trainable neural network layer. In this way, compared with the existing compression-sensing-based imaging methods in related technologies, the convergence speed can be effectively improved.
[0070] Third, in step S5 of the present invention, the coefficients are autonomously optimized through backpropagation, which can effectively overcome the problem of poor scene adaptability caused by manually setting priors in the existing video SAR imaging methods for joint low-rank sparse recovery, and can be effectively applied to complex ground object scenes.
[0071] Fourth, through the synergistic effect of the video SAR sparse reconstruction model in step S2 and the depth unfolding in step S4, high-performance imaging is achieved by utilizing spatio-temporal joint sparsity under the condition of azimuth / range sparse sampling. That is to say, under the condition of sparse sampling, the imaging performance can be effectively guaranteed.
[0072] In summary, the method of the present invention establishes a tensor signal model according to the working process of video SAR, then combines the low-rank sparse characteristics of video SAR images to establish a video SAR sparse reconstruction model to obtain the iterative solution of the imaging problem. Subsequently, by depth unfolding this iterative solution, a preliminary video SAR depth imaging network is constructed, and in this network, SAR video image correlation learning and sparse tensor recovery are carried out. Finally, the imaging network is trained through backpropagation to optimize the parameters in the network to achieve high-performance video SAR imaging under sparse sampling.
[0073] 2. Other beneficial effects or advantages of the present invention will be described in detail in the specific implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0075] Among them:
[0076] Figure 1 is a schematic flow chart of the video SAR depth imaging method based on interpretable sparse tensor recovery provided by an exemplary embodiment of the present invention;
[0077] Figure 2 is a schematic diagram of the network structure of the video SAR depth imaging method based on interpretable sparse tensor recovery provided by an exemplary embodiment of the present invention;
[0078] Figure 3 is a schematic diagram of the geometric configuration of the SAR system provided by an exemplary embodiment of the present invention;
[0079] Figure 4 is a schematic diagram of the original scene for simulation verification provided by an exemplary embodiment of the present invention;
[0080] Figure 5 is a schematic diagram of the simulation imaging result when the data sampling rate is 50% provided by an exemplary embodiment of the present invention;
[0081] Figure 6 is a schematic diagram of the simulation imaging result when the data sampling rate is 20% provided by an exemplary embodiment of the present invention;
[0082] Figure 7 is a schematic diagram of the measured data acquisition system provided by an exemplary embodiment of the present invention;
[0083] Figure 8 is a schematic diagram of the imaging result of the measured data 1 when the data sampling rate is 30% provided by an exemplary embodiment of the present invention;
[0084] Figure 9 is a schematic diagram of the imaging result of the measured data 2 when the data sampling rate is 30% provided by an exemplary embodiment of the present invention. SPECIFIC IMPLEMENTATION MANNERS
[0085] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0086] Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0087] To facilitate a clearer and more accurate understanding of the technical solutions of the present invention by relevant technical personnel, the following will first provide a more detailed description of the existing related technologies and the problems they have.
[0088] SAR (Synthetic Aperture Radar) is a microwave imaging technology that can obtain electromagnetic images of an observed scene by transmitting, receiving, and processing electromagnetic wave signals. Through sub-aperture overlap or continuous illumination, video SAR can obtain consecutive multiple-frame SAR images of the observed scene at a certain frame rate. Therefore, video SAR can sense the dynamic information of the observed scene, and this ability plays an important role in the long-term monitoring of ground areas and the perception of dynamic information.
[0089] Similar to optical video, the output of video SAR is a video stream composed of frame-by-frame SAR images. Therefore, the direct method of video SAR imaging is frame-by-frame reconstruction. According to a specific video frame rate, the SAR echo data can be divided into multiple overlapping echo data frames.
[0090] For example, the traditional frame-by-frame imaging method (such as backprojection, range-Doppler, etc.) is used to perform frame-by-frame imaging on the echo data to obtain the final SAR video image. However, in practical applications, in order to obtain the dynamic information of the observed scene, video SAR usually requires a high frame rate, which will result in huge data acquisition, storage, and calculation overheads, limiting its application in practical scenarios.
[0091] Therefore, it is necessary to study the video SAR imaging method under sparse sampling conditions. However, the imaging method based on matched filtering requires the sampling data to meet the Nyquist sampling rate, making it inapplicable to video SAR imaging under sparse sampling. With the development of the Compressed Sensing (CS) theory, the SAR imaging method based on compressed sensing has been widely studied. In video SAR imaging, tensor sparse decomposition is usually used for sparse reconstruction of video SAR. For example, the research in the literature "Joint Low-Rank and Sparse Tensors Recovery for Video Synthetic Aperture Radar Imaging" shows that the low-rank sparse video SAR imaging method can achieve super-resolution imaging under sparse sampling. However, this type of method is based on the sparse assumption of the observed scene, and the solution process requires multiple iterations, resulting in low processing efficiency in its first and foremost usage scenarios.
[0092] Based on this, the tensor-based video SAR imaging method can utilize the inter-frame correlation of multi-frame SAR data to improve the video SAR imaging performance under sparse sampling conditions. However, the problem of tensor-based video SAR imaging is usually solved by iterative optimization methods. The imaging results are not only affected by the selection of iterative parameters, but also the iterative process will lead to a decrease in imaging efficiency.
[0093] To address this problem, the literature "LRSR-ADMM-Net: A Joint Low-Rank and Sparse Recovery Network for SAR Imaging" proposed a video SAR imaging network for joint low-rank and sparse recovery. In this network, the iterative parameters can be autonomously obtained through learning, thereby improving the imaging performance and processing efficiency. However, the performance of this type of video SAR imaging network is limited by the artificially set low-rank and sparse priors. When the actually observed video SAR data is inconsistent with these priors, the imaging performance will be greatly reduced, unable to meet the actual application requirements, and the black-box processing of the deep network reduces the interpretability of the imaging process.
[0094] In summary, how to effectively ensure the imaging performance and the practicality for complex scenes under sparse sampling conditions has become an urgent problem to be solved.
[0095] Based on this, the present invention provides a brand-new solution, that is, the video SAR depth imaging method based on interpretable sparse tensor recovery of the present invention. The design idea of the technical solution of the present invention is as follows: First, establish a tensor signal model according to the working process of video SAR. Then, establish a video SAR sparse reconstruction model in combination with the low-rank and sparse characteristics of video SAR images to obtain an iterative solution to this imaging problem. Subsequently, by deeply unfolding this iterative solution, construct a preliminary video SAR depth imaging network. Finally, train the imaging network through backpropagation to optimize the parameters in the network to achieve high-performance imaging of video SAR under sparse sampling.
[0096] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.
[0097] As Figures 1 to 9 shown, according to the first aspect of the present invention, there is provided a video SAR depth imaging method based on interpretable sparse tensor recovery, including the following steps:
[0098] Step S1: Establish a tensor signal model according to the working process of video SAR and the observation model of single-frame SAR signals;
[0099] Step S2: According to the established tensor signal model, in combination with the low-rank and sparse characteristics of SAR video images, establish a video SAR sparse reconstruction model;
[0100] Step S3: Perform iterative solution on the established video SAR sparse reconstruction model;
[0101] Step S4: By deeply unfolding the iterative solution of the video SAR sparse reconstruction model, construct a preliminary video SAR depth imaging network, and perform SAR video image correlation learning and sparse tensor recovery on it;
[0102] Step S5: Through backpropagation for network training, optimize the parameters of the preliminary video SAR depth imaging network to obtain the final video SAR depth imaging network.
[0103] Through the above technical solution, first, the tensor signal model established in step S1 of the present invention can model multi-frame SAR echo signals as a tensor structure, breaking through the two-dimensional signal processing framework of traditional frame-by-frame imaging. At the same time, in combination with the low-rank and sparse joint in step S2 to establish a video SAR sparse reconstruction model, it can utilize the inter-frame correlation (low-rankness) in the time dimension and the scattering sparsity in the space dimension of SAR video to construct a reconstruction model with stronger physical constraints. In this way, data redundancy can be effectively reduced, the repeated calculations caused by high frame rates in traditional frame-by-frame imaging methods can be avoided, and the data acquisition volume and storage overhead can be reduced.
[0104] Second, step S3 of the present invention iteratively solves the video SAR sparse reconstruction model. At the same time, in combination with the deep unfolding technology in step S4, the iterative process is mapped into a trainable neural network layer. In this way, compared with the existing compression-sensing-based imaging methods in related technologies, the convergence speed can be effectively improved.
[0105] Third, step S5 of the present invention autonomously optimizes the coefficients through backpropagation, which can effectively overcome the problem of poor scene adaptability caused by artificially setting priors in the existing video SAR imaging methods for joint low-rank sparse recovery, and can be effectively applied to complex ground object scenes.
[0106] Fourth, through the synergistic effect of the video SAR sparse reconstruction model in step S2 and the deep unfolding in step S4, under the condition of azimuth / range sparse sampling, high-performance imaging is achieved by using spatio-temporal joint sparsity. That is to say, under the condition of sparse sampling, the imaging performance can be effectively guaranteed.
[0107] In summary, the method of the present invention establishes a tensor signal model according to the working process of video SAR, and then combines the low-rank sparse characteristics of video SAR images to establish a video SAR sparse reconstruction model to obtain an iterative solution to the imaging problem. Subsequently, by deeply unfolding this iterative solution, a preliminary video SAR deep imaging network is constructed, and in this network, SAR video image correlation learning and sparse tensor recovery are carried out. Finally, the imaging network is trained through backpropagation to optimize the parameters in the network to achieve high-performance video SAR imaging under sparse sampling.
[0108] It can be understood that in the step of training the network, that is, before step S5, it may further include: preparing a training and test data set for the preliminary video SAR deep imaging network, and an echo tensor and its corresponding SAR image video in the data set are used as a training sample.
[0109] In an embodiment of the present invention, step S1 of the present invention may specifically include:
[0110] Step S1-1: According to the signal data of a single-frame SAR, establish an observation model of the single-frame SAR signal:
[0111] y t =ψAx t +n
[0112] In the formula, y t represents the echo data of the t-th frame SAR signal, ψ represents the signal downsampling matrix, A represents the echo data observation matrix determined by the SAR system, x t represents the scattering coefficient of the t-th frame observation scene, and n represents the observation noise;
[0113] Step S1-2: According to the working process of video SAR and combined with the observation model of a single-frame SAR signal, establish a tensor signal model for the observed signal of video SAR:
[0114]
[0115] In the formula, Y represents the multi-frame echo signal of video SAR, Φ represents the downsampling operator corresponding to the downsampling matrix ψ, and ⊙ a,r represents the Hadamard product along the azimuth and range directions, represents the echo generation operator corresponding to the observation matrix A, X represents the SAR image video, and N represents the observation noise of the multi-frame echo signal.
[0116] In this embodiment, the observation model of the single-frame SAR signal can accurately describe the SAR echo characteristics and provide a basis for subsequent tensor modeling. At the same time, the tensor signal model uniformly represents multi-frame signals through tensors, and can effectively reduce the data redundancy requirement compared with the traditional frame-by-frame imaging method in the existing related technologies.
[0117] In an embodiment of the present invention, step S2 of the present invention may specifically include:
[0118] According to the established tensor signal model and combined with the low-rank and sparse characteristics of the SAR image video X, decompose the SAR image video X into two parts: low-rank and sparse, and establish a sparse reconstruction model for video SAR:
[0119]
[0120] In the formula, represents the minimization optimization problem with respect to X, ‖d·X‖1 represents taking the L1 norm of d·X, d represents the sparse feature tensor, and λ is the regularization coefficient, represents the relevant function for extracting the low-rank property of the SAR image video, and s.t. represents the constraint condition of the minimization optimization problem.
[0121] In this embodiment, the low-rank term utilizes the motion continuity between video frames, restricts the tensor rank through nuclear norm minimization, suppresses background clutter, and enhances the target contrast. At the same time, the sparse term (‖d·X‖1) extracts the sparse scattering features of the scene through the sparse feature tensor d, and can effectively reduce the sidelobe effect.
[0122] In an embodiment of the present invention, step S3 of the present invention may specifically include:
[0123] Step S3-1: According to the relationship between the single-frame signal of video SAR and the tensor signal model, rewrite the sparse reconstruction model of video SAR as:
[0124]
[0125] In the formula, x1 represents the first frame image of the video SAR, x T represents the T-th frame image of the video SAR, x t represents the t-th frame image of the video SAR, t represents the frame number of the video SAR image, and t = 1, …, T. D represents the sparse matrix corresponding to the sparse feature tensor d, C(x1, …, x T ) represents the correlation function corresponding to , and its expression is:
[0126]
[0127] In the formula, τ represents the frame number of the video SAR image, x τ represents the τ-th frame image of the video SAR;
[0128] Step S3-2: Using the augmented Lagrangian method, transform the video SAR sparse reconstruction model in Step S3-1 into an unconstrained optimization problem, and its augmented Lagrangian function is:
[0129]
[0130] In the formula, L(·) represents the augmented Lagrangian function, u1 represents the introduced auxiliary variable corresponding to x1, u T represents the introduced auxiliary variable corresponding to x T corresponding to, u t represents the introduced auxiliary variable corresponding to x t corresponding to, and ρ represents the coefficient of the augmented Lagrangian penalty function;
[0131] Step S3-3: According to the alternating direction multiplier method, transform the unconstrained optimization problem into iterative solutions of two sub-problems of x1, …, x T and u1, …, u T .
[0132] Step S3-3-1: Initialize and set the regularization coefficient λ, the maximum number of iterations K, and the current number of iterations k = 0;
[0133] Step S3-3-2: When the number of iterations k ≤ K, solve
[0134]
[0135] In the formula, represents the (k + 1)-th iterative solution value of x1, represents the (k + 1)-th iterative solution value of x T ;
[0136] According to the approximate alternating inexact minimization method, the solution process of
[0137] Z k+1 = λ c X k softmax β ((X k )) T D T DX k )
[0138] X k+1 / 2 = W b + W a Z k+1
[0139]
[0140] In the formula, Z k+1 represents the (k + 1)-th iterative solution value of Z, λ c represents the weight coefficient of the (k + 1)-th iteration, X k represents the k-th iterative solution value of the intermediate variable X, softmax β (·) represents the soft threshold function with the weighted value of β, (X k )) T represents the transpose of X k D T represents the transpose of D, X k+1 / 2 represents the intermediate variable of the (k + 1)-th iteration, W b and W a are weight matrices, represents the (k + 1)-th iterative solution value of x t , L ADMM (·) represents the sparse reconstruction process based on the alternating direction multiplier method;
[0141] Step S3-3-3: According to the least squares method, solve
[0142]
[0143] In the formula, represents the (k + 1)-th iterative solution value of u t , represents the k-th iterative solution value of u t , α t represents the update coefficient of the t-th frame video SAR image, represents the (k + 1)-th iterative solution value of x t ;
[0144] Step S3-3-4: Let the iteration number k = k + 1, and repeat Steps S3-3-2 to S3-3-3 until the iteration solution stop condition is met.
[0145] In this embodiment, the global optimization problem is decomposed into T parallel sub-problems, and the alternating optimization is implemented using ADMM (Alternating Direction Method of Multipliers), which can effectively reduce the computational complexity. At the same time, the penalty term introduced in the augmented Lagrangian function can effectively accelerate the convergence speed. In addition, by introducing the dual variable u t , it can effectively avoid the divergence problem caused by the matrix ill-conditioning in the existing related imaging methods based on compressive sensing, and further accelerate the convergence speed.
[0146] In an embodiment of the present invention, Step S4 of the present invention may specifically include:
[0147] Step S4-1: According to the calculation methods of Z k+1 , X k+1 / 2 and in Step S3-3-2, design a network structure for solving :
[0148] Design a weighted attention network module to solve Z k+1 , the input of this layer of network is the output of the previous layer of network and output to the next layer of network;
[0149] Design a fully connected network module to solve X k+1 / 2 , the input of this layer of network is the output of the previous layer of network and output to the next layer of network;
[0150] Design an alternating direction multiplier method network module to solve , the input of this layer of network is the output of the previous layer of network and output to the next layer of network;
[0151] Step S4-2: According to the calculation method of in Step S3-3-3, design a network structure for solving , the input of this layer of network is the output by the previous layer of network and the output of the previous iteration, and output the calculation result of this layer of network to the next layer of network;
[0152] Step S4-3: Repeat Km Perform steps S4-1 to S4-2 for -1 times, and input the output of the last step S4-2 into step S4-1 to complete the construction of the preliminary video SAR depth imaging network, where K m represents the number of network iteration modules composed of steps S4-1 to S4-2.
[0153] In this embodiment, by mapping the iterative steps into network layers and retaining the convergence of the mathematical model, the uncertainty of the black-box network can be effectively avoided. At the same time, the network autonomously extracts low-rank and sparse features through SAR video image correlation learning, which can effectively improve the reconstruction accuracy compared with manually designed features.
[0154] In an embodiment of the present invention, the parameters of the weighted attention network module of the present invention may include: the feature dimension is 128 dimensions, the number of attention heads is 16, the weight coefficient is 0.5, the spatial convolution kernel size is 5, the number of channels is 3, the stride is 1, and the padding number is 4; the temporal convolution kernel size is 3, the number of channels is 3, the stride is 1, and the padding number is 2.
[0155] Through the parameter design of this embodiment, the traceability of network behavior can be improved, and the real-time performance of calculation can be guaranteed.
[0156] In an embodiment of the present invention, the parameters of the fully connected network module of the present invention may include: the input layer feature dimension is 1024, the intermediate layer feature dimension is 256, and the output feature dimension is 1024.
[0157] In this embodiment, through the design of the bottleneck structure (1024 dimensions → 256 dimensions → 1024 dimensions), feature dimensionality reduction and information compression can be achieved, reducing memory occupancy and improving calculation efficiency.
[0158] In an embodiment of the present invention, the parameters of the alternating direction method of multipliers network module of the present invention may include: the number of iterations is 6, and the update step size is 0.6.
[0159] Through the parameter design of this embodiment, the number of iterations is compressed to 6 equivalent network layers through deep unfolding, reducing the computational complexity. Moreover, the update step size of 0.6 can balance the convergence speed and stability in gradient descent, avoiding oscillations caused by too large a step size or slow convergence speed caused by too small a step size.
[0160] For the convenience of those skilled in the relevant art to understand, the method of the present invention will be described below by taking an exemplary embodiment as an example.
[0161] First, please refer to Figure 2 , the network structure of the method of the present invention is as Figure 2 shown. Figure 3Schematic diagram of the geometric configuration of the SAR system according to this exemplary embodiment, and its basic parameters are shown in Table 1 below:
[0162] Table 1 Basic parameters of the SAR system in an exemplary embodiment
[0163]
[0164]
[0165] In this embodiment, the high-resolution SAR image of MiniSAR is used as the original scene, and the network training dataset and test dataset are generated in combination with the system parameters in Table 1. The specific process is as follows:
[0166] Step 1, generate the original scene. Crop the high-resolution MiniSAR image into images with a size of 1024×1024 at an inter-frame overlap rate of 0.8, and use these images as the original scene for generating echo data;
[0167] Step 2, generate echo data. Specifically as follows:
[0168] After setting the operation parameters, generate echoes; demodulate the obtained echoes, and the baseband echo after demodulation can be expressed as:
[0169]
[0170] where τ is the range time variable, and the number of its discrete points is N rg = 1024, η is the azimuth time variable, and the number of its discrete points is N az = 1024, Ω is the SAR observation scene, σ(i,j) is the backscattering coefficient value at the observation scene (i,j), w r (·), w a (·) respectively represent the range window function and the azimuth window function, which are taken as rectangular windows in this embodiment, K r represents the chirp rate of the chirp signal, η c is the moment when the target is crossed by the beam center, T a is the synthetic aperture time size of the point target, c is the speed of light, i is the azimuth coordinate of the observation scene, j is the range coordinate of the observation scene, and R(η) is the target range history, and its calculation formula is:
[0171]
[0172] In the formula, R is the distance between the SAR platform and the center of the scene, and η0 is the center point of the azimuth sampling time;
[0173] Step 3, set the parameters of the video SAR imaging network.
[0174] Figure 2The total number of iterative layers K of the middle network is set to 8, and the number of iterative layers of the alternating direction multiplier method network module is set to 6.
[0175] Step 4: Set the training parameters of the SAR imaging network.
[0176] The number of samples in the training dataset is 3000, the number of samples in the test dataset is 300, the batch size for network training is 30, the total number of epochs is 20, the learning rate is set to 0.0001, and the Adam optimizer is used.
[0177] Step 5: Train the video SAR imaging network using backpropagation.
[0178] Using the network training parameters and training dataset set above, train and learn the video SAR imaging network. The configuration of the network training platform is: Intel Xeon Gold 6128 CPU and NVIDIA Tesla P100 GPU (16G video memory).
[0179] Step 6: Network testing: Use the test dataset to test the trained imaging network. Send the echo data in the training dataset into the trained network, output the video SAR imaging results of the network, and calculate the corresponding performance metrics.
[0180] The simulation data test results at different data sampling rates are shown in Table 2 below, Figure 5 and Figure 6 as shown. The measured data test results are shown in Table 3 and Figure 8 , Figure 9 as shown. Among them, the calculation formulas for the evaluation metrics peak signal-to-noise ratio PSNR, structural similarity SSIM, and image entropy ENT are as follows:
[0181]
[0182]
[0183] Among them, X is the output result of the imaging network, is its corresponding original scene, E t = ∑ i,j |x t | 2 is the total energy of the t-th frame image, C1 is a parameter for calculating SSIM, and its value is 0.01. C2 is a parameter for calculating SSIM, and its value is 0.03. P t is the occurrence probability of the t-th pixel value in X.
[0184] Table 2 Simulation data test results at different data sampling rates
[0185]
[0186] Table 3 Test Results of Measured Data
[0187]
[0188] Through this embodiment, the present invention establishes a video SAR imaging network for sparse sampling conditions. The established video SAR imaging model based on low-rank sparse tensor recovery can utilize the inter-frame correlation of video SAR to solve the problem of sparse sampling. At the same time, the video SAR imaging model problem is solved through a deep unfolding network. In the proposed imaging network, a weighted attention module is used to extract the inter-frame correlation of video SAR, improving the imaging performance of the proposed method under sparse sampling conditions and its applicability to complex scenes, and meeting the actual application requirements. It can be seen from the simulation and measured results that the method of the present invention has the characteristics of low data sampling requirements and excellent imaging performance.
[0189] According to the second aspect of the present invention, there is also provided a video SAR depth imaging system based on interpretable sparse tensor recovery, which is applied to the video SAR depth imaging method based on interpretable sparse tensor recovery in any one of the technical solutions of the first aspect of the present invention. The video SAR depth imaging system based on interpretable sparse tensor recovery includes a tensor signal model establishment module, a video SAR sparse reconstruction model establishment module, an iterative solution module, a deep unfolding module, and a network training and optimization module.
[0190] Among them, the tensor signal model establishment module is used to establish a tensor signal model according to the working process of video SAR and the observation model of single-frame SAR signals; the video SAR sparse reconstruction model establishment module is used to establish a video SAR sparse reconstruction model according to the established tensor signal model in combination with the low-rank sparse characteristics of SAR image videos; the iterative solution module is used to iteratively solve the established video SAR sparse reconstruction model; the deep unfolding module is used to construct a preliminary video SAR depth imaging network by deeply unfolding the iterative solution of the video SAR sparse reconstruction model, and perform SAR video image correlation learning and sparse tensor recovery on it; the network training and optimization module is used to train the network through backpropagation and optimize the parameters of the preliminary video SAR depth imaging network to obtain the final video SAR depth imaging network.
[0191] The video SAR depth imaging system based on interpretable sparse tensor recovery of the present invention can effectively ensure the imaging performance under sparse sampling conditions and has applicability to complex scenes.
[0192] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A video SAR depth imaging method based on interpretable sparse tensor recovery, characterized in that, It includes the following steps: Step S1: Establish a tensor signal model according to the working process of video SAR and the observation model of single-frame SAR signals; Step S2: Based on the established tensor signal model and combined with the low-rank and sparse characteristics of the SAR image video, establish a video SAR sparse reconstruction model; Step S3: Perform iterative solution on the established video SAR sparse reconstruction model; Step S4: By deeply unfolding the iterative solution of the video SAR sparse reconstruction model, construct a preliminary video SAR depth imaging network, and perform SAR video image correlation learning and sparse tensor recovery on it; Step S5: Through backpropagation for network training, optimize the parameters of the preliminary video SAR depth imaging network to obtain the final video SAR depth imaging network.
2. The video SAR depth imaging method based on interpretable sparse tensor recovery according to claim 1, characterized in that, The specific content of step S1 includes: Step S1-1: According to the signal data of single-frame SAR, establish an observation model of single-frame SAR signals: y t = ψAx t + n where y t represents the echo data of the t-th frame SAR signal, ψ represents the signal downsampling matrix, A represents the echo data observation matrix determined by the SAR system, x t represents the scattering coefficient of the t-th frame observation scene, and n represents the observation noise; Step S1-2: According to the working process of video SAR and combined with the observation model of single-frame SAR signals, establish a tensor signal model of video SAR observation signals: Wherein, Y represents the multi-frame echo signal of video SAR, Φ represents the downsampling operator corresponding to the downsampling matrix ψ, and ⊙ a,r represents the Hadamard product along the azimuth and range directions, represents the echo generation operator corresponding to the observation matrix A, X represents the SAR image video, and N represents the observation noise of the multi-frame echo signal.
3. The video SAR depth imaging method based on interpretable sparse tensor recovery according to claim 2, wherein, The specific content of step S2 includes: According to the established tensor signal model and combined with the low-rank and sparse characteristics of the SAR image video X, decompose the SAR image video X into two parts, low-rank and sparse, and establish a video SAR sparse reconstruction model: In the formula, represents the minimization optimization problem with respect to X, ‖d·X‖1 represents taking the L1 norm of d·X, d represents the sparse feature tensor, and λ is the regularization coefficient. represents the related function for extracting the low-rank property of the SAR image video, and s.t. represents the constraint condition of the minimization optimization problem.
4. The video SAR depth imaging method based on interpretable sparse tensor recovery according to claim 3, wherein The specific content of step S3 includes: Step S3-1: According to the relationship between single-frame signals of video SAR and the tensor signal model, rewrite the video SAR sparse reconstruction model as: where \(x_1\) represents the first frame image of the video SAR, \(x\) T represents the \(T\)-th frame image of the video SAR, \(x\) t represents the \(t\)-th frame image of the video SAR, \(t\) represents the frame number of the video SAR image, and \(t = 1,\ldots,T\), \(D\) represents the sparse matrix corresponding to the sparse feature tensor \(d\), \(C(x_1,\ldots,x\) T ) represents the correlation function corresponding to , and its expression is: where τ represents the frame number of the video SAR image, and x τ represents the τ-th frame image of the video SAR. Step S3-2: Using the augmented Lagrangian method, transform the video SAR sparse reconstruction model in step S3-1 into an unconstrained optimization problem, and its augmented Lagrangian function is: Where \(L(\cdot)\) represents the augmented Lagrangian function, \(u_1\) represents the introduced auxiliary variable corresponding to \(x_1\), and \(u\) T represents the introduced auxiliary variable corresponding to \(x\) T , and \(u\) t represents the introduced auxiliary variable corresponding to \(x\) t , and \(\rho\) represents the coefficient of the augmented Lagrangian penalty function; Step S3-3: According to the alternating direction method of multipliers, transform the unconstrained optimization problem into the iterative solution of two sub-problems for x1, …, x T and u1, …, u T respectively.
5. The video SAR depth imaging method based on interpretable sparse tensor recovery according to claim 4, wherein The specific content of step S3-3 includes: Step S3-3-1: Initialization And set the regularization coefficient λ, the maximum number of iterations K, and the current number of iterations k = 0; Step S3-3-2: When the iteration number k ≤ K, solve In the formula, represents the (k + 1)-th iterative solution value of x1, represents x T 's (k + 1)-th iterative solution value; According to the approximate alternating inexact minimization method, The solution process is as follows: Z k+1 = λ c X k softmax β ((X k ) T D T DX k ) X k+1 / 2 = W b + W a Z k+1 where Z k+1 represents the (k + 1)-th iterative solution value of Z, λ c represents the weight coefficient of the (k + 1)-th iteration, X k represents the k-th iterative solution value of the intermediate variable X, softmax β (·) represents the soft threshold function with the weighted value of β, (X k ) T represents the transpose of X k , D T represents the transpose of D, X k+1 / 2 represents the intermediate variable of the (k + 1)-th iteration, W b and W a are weight matrices, represents the (k + 1)-th iterative solution value of x t , L ADMM (·) represents the sparse reconstruction process based on the alternating direction method of multipliers; Step S3-3-3: Solve according to the least squares method using the following formula In the formula, represents the (k + 1)-th iterative solution value of u t , represents the k-th iterative solution value of u t , and α t represents the update coefficient of the SAR image of the t-th frame video; represents the (k + 1)-th iterative solution value of x t . Step S3-3-4: Let the iteration number k = k + 1, and repeat steps S3-3-2 to S3-3-3 until the iteration solution stop condition is met.
6. The video SAR depth imaging method based on interpretable sparse tensor recovery according to claim 5, wherein The specific content of step S4 includes: Step S4-1: According to the calculation methods of Z k+1 and X k+1 / 2 in step S3-3-2, design a network structure for solving : Design a weighted attention network module to solve for Z k+1 , the input of this layer of the network is the output of the previous layer of the network and output to the next layer of the network; Design a fully connected network module to solve for X k+1 / 2 , the input of this layer of the network is the output of the previous layer of the network and output to the next layer of the network; Design an alternating direction multiplier method network module to solve The input of this layer of network is the output of the previous layer of network and output to the next layer of network; Step S4-2: According to the calculation method in Step S3-3-3, design a network structure for solving . The input of this layer of network is the output by the previous layer of network and the output of the previous iteration, and output the calculation result of this layer of network to the next layer of network; Step S4-3: Repeat K m -1 times the steps from S4-1 to S4-2, and input the output of the last time of step S4-2 into step S4-1 to complete the construction of the preliminary video SAR depth imaging network, where K m represents the number of network iteration modules composed of steps S4-1 to S4-2.
7. The video SAR depth imaging method based on interpretable sparse tensor recovery according to claim 6, characterized in that, The parameters of the weighted attention network module include: the feature dimension is 128, the number of attention heads is 16, the weight coefficient is 0.5, the spatial convolution kernel size is 5, the number of channels is 3, the stride is 1, the padding number is 4, the temporal convolution kernel size is 3, the number of channels is 3, the stride is 1, and the padding number is 2.
8. The video SAR depth imaging method based on interpretable sparse tensor recovery according to claim 6, wherein The parameters of the fully connected network module include: the input layer feature dimension is 1024, the middle layer feature dimension is 256, and the output feature dimension is 1024.
9. The video SAR depth imaging method based on interpretable sparse tensor recovery according to claim 6, wherein The parameters of the alternating direction method of multipliers network module include: the number of iterations is 6, and the update step size is 0.
6.
10. A video SAR depth imaging system based on interpretable sparse tensor recovery, characterized in that Applied to the video SAR depth imaging method based on interpretable sparse tensor recovery described in any one of claims 1-9, characterized in that the video SAR depth imaging system based on interpretable sparse tensor recovery includes: A tensor signal model establishment module, used to establish a tensor signal model according to the working process of video SAR and the observation model of single-frame SAR signals; A video SAR sparse reconstruction model establishment module, used to establish a video SAR sparse reconstruction model according to the established tensor signal model and combined with the low-rank and sparse characteristics of the SAR image video; An iterative solution module for iteratively solving the established video SAR sparse reconstruction model; A deep unfolding module for constructing a preliminary video SAR depth imaging network by unfolding the iterative solution of the video SAR sparse reconstruction model, and performing SAR video image correlation learning and sparse tensor recovery on it; A network training and optimization module for training the network through backpropagation and optimizing the parameters of the preliminary video SAR depth imaging network to obtain the final video SAR depth imaging network.
Citation Information
Cited By
SAR three-dimensional enhanced imaging method based on non-local tensor decomposition
CN121049906A