Hyperspectral snapshot compression imaging reconstruction method combining memory and depth expansion
By combining memory and deep unfolding methods, introducing a degradation learning module and a memory-enhanced spatial-spectral attention mechanism, the problems of noise interference and information loss in hyperspectral snapshot compressed imaging reconstruction are solved, achieving high-quality and stable reconstruction results.
Patent Information
- Application Number
- CN202511108678.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Existing hyperspectral snapshot compression imaging reconstruction methods are unstable under noise interference or changes in measurement conditions, and fail to fully exploit the spatial-spectral features of hyperspectral data, which may lead to the loss of important information during the reconstruction process.
A hyperspectral snapshot compression imaging reconstruction method combining memory and depth unfolding is adopted. By introducing a degradation learning module and a memory-enhanced spatial-spectral attention mechanism, the optimal solution is gradually approximated, thereby enhancing the stability and accuracy of the reconstruction process.
It improves the accuracy and stability of hyperspectral image reconstruction, enabling more complete capture of complex spatial structures and spectral characteristics, and enhancing the quality and robustness of reconstructed images.
Smart Images

Figure CN120976434A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of deep learning technology and relates to a hyperspectral snapshot compression imaging reconstruction method that combines memory and depth unfolding. Background Technology
[0002] Traditional hyperspectral imaging methods typically require scanning or multiple exposures to acquire a complete spectral cube. While these methods can produce high-quality images, their complex system structure, high cost, and slow imaging speed make them unsuitable for real-time monitoring of dynamic scenes. In contrast, hyperspectral snapshot compression imaging technology compresses three-dimensional hyperspectral data into a two-dimensional measurement containing both spatial and spectral information, acquiring complete information in a single exposure. This significantly improves imaging speed and system compactness. This technology is particularly suitable for applications where targets change rapidly and real-time performance is critical.
[0003] In snapshot compressed imaging systems, coded aperture snapshot spectral imaging systems are widely used due to their efficient optical modulation mechanisms. This system captures modulated two-dimensional measurements on a sensor using a pre-designed optical mask and spectroscopic elements, and then reconstructs three-dimensional hyperspectral data using subsequent algorithms. Since the measurement process is essentially a compressed sensing process, a robust reconstruction algorithm is needed to infer the complete hyperspectral cube from incomplete observation data. However, because this inverse problem is highly ill-posed and sensitive to noise, achieving efficient and robust reconstruction remains a key challenge in this field.
[0004] Traditional reconstruction methods typically regularize the solution process by constructing artificial priors (such as low-rank, sparsity, etc.). While these methods alleviate the ill-conditioned nature of the inverse problem to some extent, they often fail to fully exploit the rich spatial-spectral features of hyperspectral data due to limitations in prior design and dependence on specific data distributions. In recent years, deep learning methods, with their powerful data-driven characteristics, have been introduced into hyperspectral reconstruction tasks and have achieved significant progress. Among them, deep networks based on end-to-end training can automatically extract complex features and reconstruct high-quality hyperspectral data. However, these methods often lack physical interpretability and struggle to explicitly model the relationship between the imaging system and the reconstruction process, leading to instability when faced with noise interference or changes in measurement conditions. To address this, deep unfolding networks proposed in recent years combine the iterative process of traditional optimization algorithms with deep learning, effectively improving reconstruction results while preserving physical interpretability and leveraging the powerful representational capabilities of deep learning. However, existing deep unfolding networks often involve shallow unfolding or independent stage processing, failing to fully consider information transfer and memory effects between stages. This can lead to the loss of important contextual information and historical features during reconstruction, hindering the modeling of long-term dependencies. Therefore, how to introduce an efficient memory mechanism into the deep unfolding framework to enhance the synergy and information retention between stages has become an important research direction for further improving the reconstruction quality and stability of hyperspectral snapshot compressed imaging. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a hyperspectral snapshot compressed imaging reconstruction method that combines memory and deep unfolding. A holistic framework consisting of a degradation learning module, a robust unfolding inference network, and a memory-enhanced attention mechanism is proposed to guide high-quality reconstruction of hyperspectral images. In each stage of deep unfolding inference, this method introduces a degradation learning module to model and correct the differences between the imaging model and actual observations. Combined with robust linear projection and nonlinear denoising operations, it effectively reduces error accumulation during reconstruction and improves reconstruction stability. Addressing the complex spatial-spectral dependencies of hyperspectral images, a memory-enhanced spatial-spectral attention module is designed in the denoising network, comprising a gated memory unit, a surrogate spatial attention, and a windowed spectral attention. Specifically, the gated memory unit utilizes historical memory states to assist current feature learning to enhance long-term dependencies; the surrogate spatial attention combines local window partitioning and multi-scale fusion to effectively model local and global spatial information; and the windowed spectral attention highlights key band features through spectral-dimensional window partitioning and interaction. In this way, the present invention can more fully capture the complex spatial structure and spectral characteristics in hyperspectral images, and while reconstructing image details, it has stronger interpretability and robustness compared with other methods.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A hyperspectral snapshot compressed imaging reconstruction method combining memory and depth unrolling is disclosed. The method includes the following steps: S1: acquiring two-dimensional compressed images of the target scene; S2: preprocessing the hyperspectral image training set to generate sample data for training; S3: inputting the preprocessed training samples into a memory-enhanced spatial-spectral attention robust unrolling network for training; S4: after training, inputting test samples into the network to reconstruct the hyperspectral image and output the result.
[0008] 1. Further, in step S1, a two-dimensional compressed image is acquired using a coded aperture snapshot spectral imaging system, specifically including: S11, selecting the spectral range, setting the working wavelength to 450nm to 650nm, and dividing it into 28 channels; S12, designing the coded aperture, introducing different optical codes at each pixel; S13, setting the parameters of the coded aperture snapshot spectral imaging system, including but not limited to the number of dispersive elements and the dispersive shift step size; S14, capturing a compressed measurement image of the entire scene through an optical sensor.
[0009] 2. Further, in step S2, the hyperspectral images used as the training set are preprocessed, specifically including: S21, dividing the hyperspectral images into two parts; the first part is divided into samples using a 256×256 non-overlapping window, and the second part is divided into samples using a 128×128 non-overlapping window, and randomly combined into several 256×256 samples; S22, performing data augmentation on the two parts of samples respectively, including random rotation and flipping; S23, using a simulated imaging system, the hyperspectral images are modulated by a pre-designed physical mask, and then cropped, offset, and summed to generate corresponding two-dimensional compressed measurements. The data-augmented hyperspectral images and their corresponding two-dimensional compressed measurements together constitute training sample pairs.
[0010] 3. Further, in step S3, the two-dimensional compressed measurement training samples processed by the simulated imaging system are initialized. First, sliding extraction is performed in the channel dimension to rearrange each column in the compressed measurement and restore it to the corresponding position in the output hyperspectral image, thereby converting the original two-dimensional data into a coarse three-dimensional hyperspectral image. Then, the physical mask of the imaging system is stitched with the coarse hyperspectral image in the channel dimension, and the number of channels is restored to the number of channels in the original hyperspectral image through a 1×1 convolution, thus obtaining the initial input of the network. The specific formula is as follows:
[0011]
[0012] Where X is the input feature map, M is the physical mask, and I(X) is the obtained initial input. The symbol represents a convolution operation with a filter size of N×N, and "[]" represents a concatenation operation along the channel dimension.
[0013] Furthermore, the recovery of a three-dimensional hyperspectral image from two-dimensional measurement data is reduced to the following optimization problem:
[0014]
[0015] Where argmin represents the variable value that minimizes the objective function, y is the two-dimensional compressed measurement, Φ is the sensing matrix, and s is the original three-dimensional spectral data. For the predicted three-dimensional spectral data, τ is the regularization parameter, and R(s) is the regularization term.
[0016] Then, an auxiliary variable v is introduced, and variable v is updated stage by stage through alternating optimization. k With x k This allows the reconstruction results to gradually converge to the target hyperspectral image, as shown in the following formula:
[0017]
[0018] Each stage first uses the intermediate variable x k Reconstruction estimate v from the previous stage k-1 A linear projection operation is performed to ensure that the estimation results are consistent with the physical measurement model. This projection step can be represented as a constrained least squares optimization problem:
[0019]
[0020] Furthermore, the closed-form solution to the constrained least squares optimization problem is obtained using the following formula:
[0021]
[0022] in, This is used to project the current measurement residual back to a linear manifold that satisfies the measurement constraints, thereby correcting the estimation error, (y-Φv) k ) represents the current estimated value v k The difference between v and the actual measured data y k This represents the reconstruction estimate at the current stage, and through a projection step, a new auxiliary variable x is obtained that better conforms to the constraints of the measurement data. k+1 .
[0023] A degradation learning module (including a 1×1 convolutional layer and a 3×3 deep convolutional layer) is used to learn and model the degradation residual between the actual observation matrix and the ideal matrix. A 3×3 deep convolutional layer is also used to implement gradient correction, thereby optimizing the solution process of linear projection. The formula is as follows:
[0024]
[0025] in, This is the corrected observation matrix estimated by the degradation learning module in this stage. This indicates a deep convolutional layer used for gradient correction.
[0026] The next goal is to make result x k+1 To more closely approximate the true hyperspectral image distribution, the optimization objective becomes:
[0027]
[0028] This process employs a training denoising network. To mitigate potential information loss during multi-stage iterations of the deep unfolded network, this method introduces a memory mechanism into each denoising network, specifically by incorporating the hidden state h from the previous stage. k And the observation matrix information Φ, the formula is as follows:
[0029] v k+1 ,h k+1 =D k+1 (x k+1 ,h k ,Φ)
[0030] Among them, D k+1 h represents the deep denoising network used in stage k+1. k The set of multi-scale memory features generated by the k-th stage denoising network, i.e. It contains hidden states at different levels in a U-shaped network, where Φ represents measurement matrix information.
[0031] Furthermore, the denoising network consists of a U-shaped network, comprising three stages: downsampling, intermediate layers, and upsampling. First, for x... k A 1×1 convolution is applied to adjust the channel dimensions and perform preliminary feature extraction. Next, a 3×3 convolutional layer is used to fuse the encoded aperture mask M, where M is copied N times along the channel dimensions to ensure its shape aligns with the input features, ultimately yielding the mask-embedded feature x. ME The formula is as follows:
[0032]
[0033] Among them, M r T represents the result of repeating M N times along the channel dimension. I For a full 1 tensor, This represents a convolution operation with a filter size of N×N.
[0034] The downsampling module performs a 4×4 convolution operation, as shown in the following formula:
[0035]
[0036] Where X is the input feature map, and U(X) is the feature map obtained after upsampling. This represents a convolution operation with a filter size of N×N.
[0037] The upsampling module uses a 2×2 deconvolution operation, as shown in the following formula:
[0038]
[0039] Where X is the input feature map, and D(X) is the feature map obtained after downsampling. This represents a deconvolution operation with a filter size of N×N.
[0040] After two downsampling operations and attention module processing, the feature map output by the U-shaped network is obtained after passing through the intermediate layer and then undergoing two more upsampling operations and attention module processing.
[0041] Furthermore, in the denoising network, a spatial-spectral attention module with memory enhancement as its main component is employed, combining spatial self-attention and spectral self-attention. First, the input features pass through a 5×5 deep convolutional layer to perform conditional position encoding, resulting in... After layer normalization, spatial self-attention calculation is performed first, using a memory-enhanced proxy spatial attention mechanism.
[0042] Before performing self-attention computation, the input features It needs to be compared with the hidden state of the previous iteration phase. To perform fusion, that is, to fuse the current feature x j Compared to the previous stage of hidden state The joint features are concatenated along the channel dimension, and then processed through a memory update gate and a memory reset gate to dynamically adjust the update and reset of the hidden state. Both the memory update gate and the memory reset gate consist of a 3×3 convolutional layer and a sigmoid activation function, and their calculation process can be represented as follows:
[0043]
[0044] Where σ(·) represents the Sigmoid activation function, and These represent the results after processing by the memory update gate and the memory reset gate, respectively. "[]" indicates the splicing operation on the channel dimension.
[0045] Next, the historical hidden state is gated to be appropriately reset at the current moment, and the fusion information is further extracted through the state transition unit:
[0046]
[0047] Where ⊙ represents element-wise multiplication, t j The Tanh activation function is used to perform a non-linear transformation on the current information. "[]" indicates a concatenation operation on the channel dimension.
[0048] Furthermore, after computation in the gated memory unit, for the features A 1×1 convolutional layer and a 3×3 depthwise convolutional layer are used to calculate the query, key, and value, as shown in the following formula:
[0049]
[0050] in, This represents a convolution operation with a filter size of N×N. This represents a depthwise convolution operation with a filter size of N×N.
[0051] Then, a window of size M×M is used for Q. spa K spa V spa Perform non-overlapping partitioning to obtain
[0052] The computation process of proxy attention can be divided into two core stages: proxy aggregation and proxy broadcasting. In the proxy aggregation stage, proxy label A is used as query Q. spa The agent, with key K spa Calculate attention, extract key information, and in the value V spa Feature aggregation is performed under the mapping to achieve efficient spatial modeling. The calculation is as follows:
[0053]
[0054] Here, Softmax() represents normalizing the attention score along the last dimension. For the calculated surrogate features, d is a scaling factor to prevent gradient vanishing, and in this method it is set to be the same as the window size M.
[0055] During the proxy broadcast phase, the extracted information V A It is further passed to the original query Q spa This is to enhance feature representation and improve the effectiveness of attention computation. Specifically, the surrogate tag A serves as the key K. spa The agent, and then with V A Perform attention calculations to obtain the final attention output:
[0056]
[0057] in, This is the final result of the attention calculation.
[0058] Finally, O undergoes feature transformation via a 1×1 convolutional layer. Subsequently, to further enhance feature diversity, an additional 3×3 depthwise convolutional layer is used to transform V. spa The processed features are then added to the result of the feature transformation to obtain the optimized output features.
[0059] Furthermore, after the proxy attention calculation, window spectral self-attention calculation is performed. Within each spatial window, channels of different bands are treated as independent tokens, and self-attention calculation is performed along the spectral dimension. First, the input features... After performing layer normalization, the query, key, and value are calculated separately using 1×1 convolutions, as shown in the following formulas:
[0060]
[0061] in, This represents a convolution operation with a filter size of N×N. Then, Q... spe K spe and V spe Divide the M×M window into non-overlapping sections to obtain the shape as follows The window-level features are then used. Finally, a spectral self-attention operation is performed within the window, as shown in the following formula:
[0062]
[0063] Where d represents the scaling factor, and in this method it is set to be the same as the window size M.
[0064] Furthermore, during the K iterations of the unfolded network, linear projection and denoising are performed alternately, continuously approximating the solution space jointly determined by measurement constraints and data priors. After K iterations, the final reconstructed hyperspectral image v is obtained. K If the true hyperspectral image is denoted as The overall network loss function adopts a staged weighted loss strategy, defined as follows:
[0065]
[0066] in, Represents the losses in the final reconstruction phase. These represent the losses of the three stages preceding the final stage, and are assigned different weights to reflect the optimization contribution of multi-stage training. The loss function for each stage k is defined as follows:
[0067]
[0068] Where, ‖·‖2 is the two-dimensional Euclidean norm, used to calculate the pixel-level error between the current estimate and the reference true value.
[0069] The beneficial effects of this invention are as follows:
[0070] This invention proposes a hyperspectral snapshot compressed imaging reconstruction method combining memory and deep unfolding. It utilizes a deep unfolding inference framework to progressively approximate the optimal solution, balancing the interpretability of the physical model with the powerful expressive capabilities of deep networks, achieving high-quality hyperspectral image reconstruction. This invention proposes a degradation learning module to model the difference between the actual observation matrix and the ideal imaging model, and dynamically corrects the linear projection process within the network, improving the accuracy and convergence speed of the reconstruction results. Furthermore, this invention proposes a memory-enhanced spatial-spectral attention mechanism. By introducing gated memory units to enhance the transitivity of features between stages, and combining surrogate spatial attention and windowed spectral attention, it fully captures the spatial details, global dependencies, and key spectral features of the image, improving the model's feature utilization efficiency and generalization ability. Experimental results on multiple hyperspectral datasets demonstrate that the proposed method outperforms existing methods in terms of reconstruction accuracy and stability.
[0071] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0072] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0073] Figure 1 This is a flowchart of the method of the present invention;
[0074] Figure 2 The overall architecture of the spatial-spectral attention robust unfolding network for memory enhancement is shown in the figure, where (a) is the RGAP linear projection, (b) is the degradation learning module, and (c) is the RGAP deep unfolding network.
[0075] Figure 3 This is a diagram of the denoising network structure of the present invention;
[0076] Figure 4 This is a structural diagram of the memory-enhanced spatial-spectral attention module of the present invention;
[0077] Figure 5 This is a diagram of the memory-enhanced agent spatial attention structure of the present invention;
[0078] Figure 6 This is a structural diagram of the gated memory unit of the present invention;
[0079] Figure 7 This is a diagram of the window spectral attention structure of the present invention;
[0080] Figure 8 Visualization results of different methods on four channels of a scene on the KAIST hyperspectral image dataset, where (a) TwIST, (b) DWMT, (c) HDiff-HIR, (d) DAUHST, (e) RGAP-Net, and (f) ground truth map; Detailed Implementation
[0081] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.
[0082] Figure 1 This is a flowchart of the method of the present invention. The present invention provides a hyperspectral snapshot compressed imaging reconstruction method combining memory and depth unrolling. As shown in the figure, in the image acquisition stage, a two-dimensional compressed measurement image is acquired using an coded aperture snapshot spectral imaging system. The memory-enhanced depth unrolling network used for hyperspectral image reconstruction is as follows: Figure 2 As shown, this method, by introducing a gated memory mechanism and a spatial-spectral attention module, fully utilizes historical features and measurement priors during the inference process to gradually recover high-quality hyperspectral images. Specifically, the proposed model employs a degradation learning module to dynamically model the residual between actual observations and the ideal model at each iteration stage, and guides the network to better capture spatial details, global context, and spectral features of the image through a memory-enhanced spatial-spectral attention module, effectively improving reconstruction quality and stability. The gated memory units in the network can transmit important information across stages, enhancing the feature correlation between stages, thereby further improving reconstruction accuracy and robustness. The overall scheme fully combines measurement constraints, physical interpretability, and the feature representation capabilities of deep networks, achieving higher reconstruction quality and artifact suppression. Specifically, the technical solution of this invention includes the following:
[0083] 1. Data Acquisition: Snapshot compressed image acquisition is achieved using a coded aperture snapshot spectral imaging system. First, the spectral range is selected, with the working wavelength set to 450nm to 650nm and divided into 28 channels. Then, the coded aperture is designed, introducing different optical codes at each pixel. Next, the parameters of the coded aperture snapshot spectral imaging system are set, including but not limited to the number of dispersive elements and the dispersion shift step size. Finally, a compressed measurement image of the entire scene is captured using an optical sensor.
[0084] 2. Data Preprocessing: The hyperspectral images in the training set are preprocessed. First, the hyperspectral images are divided into two parts. The first part is divided into samples using a 256×256 non-overlapping window, and the second part is divided into samples using a 128×128 non-overlapping window, which are then randomly combined into several 256×256 samples. Next, data augmentation is performed on both parts of the samples, including random rotation and flipping. Then, using a simulated imaging system, the hyperspectral images are modulated using a pre-designed physical mask, followed by cropping, offsetting, and summing to generate corresponding two-dimensional compressed measurements. The data-augmented hyperspectral images and their corresponding two-dimensional compressed measurements together constitute training sample pairs.
[0085] 3. Initialize the training samples from the two-dimensional compressed measurement processed by the simulated imaging system. First, perform sliding extraction along the channel dimension, rearranging each column in the compressed measurement and restoring it to the corresponding position in the output hyperspectral image, thus converting the original two-dimensional data into a coarse three-dimensional hyperspectral image. Then, stitch the physical mask of the imaging system with this coarse hyperspectral image along the channel dimension, and restore the number of channels to the number of channels in the original hyperspectral image through a 1×1 convolution, thus obtaining the initial input to the network. The specific formula is shown below:
[0086]
[0087] Where X is the input feature map, M is the physical mask, and I(X) is the obtained initial input. The symbol represents a convolution operation with a filter size of N×N, and "[]" represents a concatenation operation along the channel dimension.
[0088] 4. Recovering a three-dimensional hyperspectral image from two-dimensional measurement data reduces to the following optimization problem:
[0089]
[0090] Where argmin represents the variable value that minimizes the objective function, y is the two-dimensional compressed measurement, Φ is the sensing matrix, and s is the original three-dimensional spectral data. For the predicted three-dimensional spectral data, τ is the regularization parameter, and R(s) is the regularization term.
[0091] Then, an auxiliary variable v is introduced, and variable v is updated stage by stage through alternating optimization. k With x k This allows the reconstruction results to gradually converge to the target hyperspectral image, as shown in the following formula:
[0092]
[0093] Each stage first uses the auxiliary variable x k Reconstruction estimate v from the previous stagek-1 A linear projection operation is performed to ensure that the estimation results are consistent with the physical measurement model. This projection step can be represented as a constrained least-squares optimization problem:
[0094]
[0095] 5. To find the closed-form solution to a constrained least squares optimization problem, the formula is as follows:
[0096]
[0097] in, This is used to project the current measurement residual back to a linear manifold that satisfies the measurement constraints, thereby correcting the estimation error, (y-Φv) k ) represents the current estimated value v k The difference between v and the actual measured data y k This represents the reconstruction estimate at the current stage, and through a projection step, a new auxiliary variable x is obtained that better conforms to the constraints of the measurement data. k+1 .
[0098] A degradation learning module (including a 1×1 convolutional layer and a 3×3 deep convolutional layer) is used to learn and model the degradation residual between the actual observation matrix and the ideal matrix. A 3×3 deep convolutional layer is also used to implement gradient correction, thereby optimizing the solution process of linear projection. The formula is as follows:
[0099]
[0100] in, This is the corrected observation matrix estimated by the degradation learning module in this stage. This indicates a deep convolutional layer used for gradient correction.
[0101] The next goal is to make result x k+1 To more closely approximate the true hyperspectral image distribution, the optimization objective becomes:
[0102]
[0103] This process employs a training denoising network. To mitigate potential information loss during multi-stage iterations of the deep unfolded network, this method introduces a memory mechanism into each denoising network, specifically by incorporating the hidden state h from the previous stage. k And the observation matrix information Φ, the formula is as follows:
[0104] v k+1 ,h k+1 =D k+1 (x k+1 ,h k ,Φ)
[0105] Among them, D k+1 h represents the deep denoising network used in stage k+1. k The set of multi-scale memory features generated by the k-th stage denoising network, i.e. It contains hidden states at different levels in a U-shaped network, where Φ represents measurement matrix information.
[0106] 6. The denoising network consists of a U-shaped network, comprising three stages: downsampling, intermediate layers, and upsampling. First, for x... k A 1×1 convolution is applied to adjust the channel dimensions and perform preliminary feature extraction. Next, a 3×3 convolutional layer is used to fuse the encoded aperture mask M, where M is copied N times along the channel dimensions to ensure its shape aligns with the input features, ultimately yielding the mask-embedded feature x. ME The formula is as follows:
[0107]
[0108] Among them, M r T represents the result of repeating M N times along the channel dimension. I For a full 1 tensor, This represents a convolution operation with a filter size of N×N.
[0109] The downsampling module performs a 4×4 convolution operation, as shown in the following formula:
[0110]
[0111] Where X is the input feature map, and U(X) is the feature map obtained after upsampling. This represents a convolution operation with a filter size of N×N.
[0112] The upsampling module uses a 2×2 deconvolution operation, as shown in the following formula:
[0113]
[0114] Where X is the input feature map, and D(X) is the feature map obtained after downsampling. This represents a deconvolution operation with a filter size of N×N.
[0115] After two downsampling operations and attention module processing, the feature map output by the U-shaped network is obtained after passing through the intermediate layer and then undergoing two more upsampling operations and attention module processing.
[0116] 7. In the denoising network, a spatial-spectral attention module with memory enhancement as its main component is employed, combining spatial self-attention and spectral self-attention. First, the input features pass through a 5×5 deep convolutional layer to perform conditional position encoding, resulting in... After layer normalization, spatial self-attention calculation is performed first, using a memory-enhanced proxy spatial attention mechanism.
[0117] Before performing self-attention computation, the input features It needs to be compared with the hidden state of the previous iteration phase. To perform fusion, that is, to fuse the current feature x j Compared to the previous stage of hidden state The joint features are concatenated along the channel dimension, and then processed through a memory update gate and a memory reset gate to dynamically adjust the update and reset of the hidden state. Both the memory update gate and the memory reset gate consist of a 3×3 convolutional layer and a sigmoid activation function, and their calculation process can be represented as follows:
[0118]
[0119] Where σ(·) represents the Sigmoid activation function, and These represent the results after processing by the memory update gate and the memory reset gate, respectively. "[]" indicates the splicing operation on the channel dimension.
[0120] Next, the historical hidden state is gated to be appropriately reset at the current moment, and the fusion information is further extracted through the state transition unit:
[0121]
[0122] Where ⊙ represents element-wise multiplication, t j The Tanh activation function is used to perform a non-linear transformation on the current information. "[]" indicates a concatenation operation on the channel dimension.
[0123] 8. After computation in the gated memory unit, for the features A 1×1 convolutional layer and a 3×3 depthwise convolutional layer are used to calculate the query, key, and value, as shown in the following formula:
[0124]
[0125] in, This represents a convolution operation with a filter size of N×N. This represents a depthwise convolution operation with a filter size of N×N.
[0126] Then, a window of size M×M is used for Q. spa K spa V spa Perform non-overlapping partitioning to obtain
[0127] The computation process of proxy attention can be divided into two core stages: proxy aggregation and proxy broadcasting. In the proxy aggregation stage, proxy label A is used as query Q. spa The agent, with key K spa Calculate attention, extract key information, and in the value V spa Feature aggregation is performed under the mapping to achieve efficient spatial modeling. The calculation is as follows:
[0128]
[0129] in, For the calculated surrogate features, d is a scaling factor to prevent gradient vanishing, and in this method it is set to be the same as the window size M.
[0130] During the proxy broadcast phase, the extracted information V A It is further passed to the original query Q spa This is to enhance feature representation and improve the effectiveness of attention computation. Specifically, the surrogate tag A serves as the key K. spa The agent, and then with V A Perform attention calculations to obtain the final attention output:
[0131]
[0132] in, This is the final result of the attention calculation.
[0133] Finally, O undergoes feature transformation via a 1×1 convolutional layer. Subsequently, to further enhance feature diversity, an additional 3×3 depthwise convolutional layer is used to transform V. spa The processed features are then added to the result of the feature transformation to obtain the optimized output features.
[0134] 9. After the proxy attention calculation, window spectral self-attention calculation is performed. Within each spatial window, channels of different bands are treated as independent tokens, and self-attention calculation is performed along the spectral dimension. First, the input features... After performing layer normalization, the query, key, and value are calculated separately using 1×1 convolutions, as shown in the following formulas:
[0135]
[0136] in, This represents a convolution operation with a filter size of N×N. Then, Q... spe K spe and V spe Divide the M×M window into non-overlapping sections to obtain the shape as follows The window-level features are then used. Finally, a spectral self-attention operation is performed within the window, as shown in the following formula:
[0137]
[0138] Where d represents the scaling factor, and in this method it is set to be the same as the window size M.
[0139] 10. During the K iterations of the unfolded network, linear projection and denoising are performed alternately, continuously approximating the solution space jointly determined by measurement constraints and data priors. After K iterations, the final reconstructed hyperspectral image v is obtained. K If the true hyperspectral image is denoted as The overall network loss function adopts a staged weighted loss strategy, defined as follows:
[0140]
[0141] in, Represents the losses in the final reconstruction phase. These represent the losses of the three stages preceding the final stage, and are assigned different weights to reflect the optimization contribution of multi-stage training. The loss function for each stage k is defined as follows:
[0142]
[0143] Where, ‖·‖2 is the two-dimensional Euclidean norm, used to calculate the pixel-level error between the current estimate and the reference true value.
[0144] like Figure 8 The experimental results of the hyperspectral snapshot compression imaging reconstruction method (RAGP-Net) based on memory and depth unfolding described in this invention on an open-source hyperspectral natural scene show that the method can effectively restore the detailed structure of the scene, and the reconstructed image has high quality and is highly consistent with the real scene. The reconstruction effect of this invention can be further illustrated by comparative experiments. The method of this invention was compared with other existing methods such as GAP-TV, DWMT, HDiff-HIR, BIRNAT, DAUHST, PADUT, and RDLUF on a hyperspectral natural scene dataset. Peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) were calculated respectively. A higher PSNR indicates relatively less distortion and higher hyperspectral image quality; a higher SSIM indicates that the image structure, brightness, and contrast are more similar to the real image, and the image quality is higher. Table 1 shows the average reconstruction results of ten scenes selected from the KAIST test set using different methods.
[0145] Table 1 compares RGAP-Net with various methods on the KAIST hyperspectral image dataset (average of ten scenes).
[0146]
[0147] The method of this invention achieves the best reconstruction accuracy on this dataset. It can be seen that the method proposed in this invention significantly outperforms other hyperspectral image reconstruction methods in terms of reconstruction quality, and can more finely restore the detailed structure of the image, making the reconstruction results more natural and realistic.
[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications should be covered within the scope of the claims of the present invention.
Claims
1. A hyperspectral snapshot compressed imaging reconstruction method combining memory and depth unfolding, characterized in that: The method includes the following steps: S1: Acquire a two-dimensional compressed image of the target scene; S2: Preprocess the hyperspectral image training set to generate sample data for training; S3: Input the preprocessed training samples into the memory-enhanced spatial-spectral attention robust unfolding network for training; S4: After training is complete, input the test samples into the network to reconstruct the hyperspectral image and output the result; In step S1, a two-dimensional compressed image is acquired using a coded aperture snapshot spectral imaging system, specifically including the following steps: S11, selecting the spectral range, setting the working wavelength to 450nm to 650nm, and dividing it into 28 channels; S12, designing the coded aperture, introducing different optical codes at each pixel; S13, setting the parameters of the coded aperture snapshot spectral imaging system, including but not limited to the number of dispersive elements and the dispersive shift step size; S14, capturing a compressed measurement image of the entire scene using an optical sensor. In step S2, the hyperspectral images used as the training set are preprocessed, specifically including the following steps: S21, the hyperspectral images are divided into two parts. The first part is divided into samples using a 256×256 non-overlapping window, and the second part is divided into samples using a 128×128 non-overlapping window, and randomly combined into several 256×256 samples; S22, data augmentation is performed on the two parts of samples respectively, including random rotation and flipping; S23, through a simulated imaging system, the hyperspectral images are modulated by a pre-designed physical mask, and then cropped, offset, and summed to generate corresponding two-dimensional compressed measurements. The data-augmented hyperspectral images and their corresponding two-dimensional compressed measurements together constitute training sample pairs. In step S3, the two-dimensional compressed measurement training samples processed by the simulated imaging system are initialized. First, sliding extraction is performed in the channel dimension to rearrange each column in the compressed measurement and restore it to the corresponding position in the output hyperspectral image, thereby converting the original two-dimensional data into a coarse three-dimensional hyperspectral image. Then, the physical mask of the imaging system is stitched with the coarse hyperspectral image in the channel dimension, and the number of channels is restored to the number of channels in the original hyperspectral image through a 1×1 convolution, thus obtaining the initial input of the network. The specific formula is as follows: Where X is the input feature map, M is the physical mask, and I(X) is the obtained initial input. "[]" indicates a convolution operation with a filter size of N×N, and "[]" indicates a concatenation operation on the channel dimension.
2. The hyperspectral snapshot compression imaging reconstruction method combining memory and depth unfolding according to claim 1, characterized in that: First, the recovery of a three-dimensional hyperspectral image from two-dimensional measurement data is reduced to the following optimization problem: Where argmin represents the variable value that minimizes the objective function, y is the two-dimensional compressed measurement, Φ is the sensing matrix, and s is the original three-dimensional spectral data. For the predicted three-dimensional spectral data, τ is the regularization parameter, and R(s) is the regularization term; Then, an auxiliary variable v is introduced, and variable v is updated stage by stage through alternating optimization. k With x k This allows the reconstruction results to gradually converge to the target hyperspectral image, as shown in the following formula: Each stage first uses the intermediate variable x k Reconstruction estimate v from the previous stage k-1 A linear projection operation is performed to ensure that the estimation results are consistent with the physical measurement model. This projection step can be represented as a constrained least squares optimization problem:
3. The hyperspectral snapshot compression imaging reconstruction method combining memory and depth unfolding according to claim 2, characterized in that: The formula for finding a closed-form solution to a constrained least-squares optimization problem is as follows: in, This is used to project the current measurement residual back to a linear manifold that satisfies the measurement constraints, thereby correcting the estimation error, (y-Φv) k ) represents the current estimated value v k The difference between v and the actual measured data y k This represents the reconstruction estimate at the current stage, and through a projection step, a new auxiliary variable x is obtained that better conforms to the constraints of the measurement data. k+1 ; A degradation learning module (including a 1×1 convolutional layer and a 3×3 deep convolutional layer) is used to learn and model the degradation residual between the actual observation matrix and the ideal matrix. A 3×3 deep convolutional layer is also used to implement gradient correction, thereby optimizing the solution process of linear projection. The formula is as follows: in, This is the corrected observation matrix estimated by the degradation learning module in this stage. This represents a deep convolutional layer used for gradient correction; The next goal is to make result x k+1 To more closely approximate the true hyperspectral image distribution, the optimization objective becomes: Where R(v) represents the regularization term and τ represents the adjustment term, used to balance the importance of the data consistency term and the regularization term; this process is implemented by training a denoising network. To mitigate the information loss that may occur during the multi-stage iteration of the deep unfolded network, this method introduces a memory mechanism in each denoising network, that is, introducing the hidden state h from the previous stage. k And the observation matrix information Φ, the formula is as follows: v k+1 ,h k+1 =D k+1 (x k+1 ,h k ,Φ) Among them, D k+1 h represents the deep denoising network used in stage k+1. k The set of multi-scale memory features generated by the k-th stage denoising network, i.e. It contains hidden states at different levels in a U-shaped network, where Φ represents measurement matrix information.
4. The hyperspectral snapshot compression imaging reconstruction method combining memory and depth unfolding according to claim 3, characterized in that: The denoising network consists of a U-shaped network, comprising three stages: downsampling, intermediate layers, and upsampling; firstly, x... k A 1×1 convolution is applied to adjust the channel dimensions and perform preliminary feature extraction. Next, a 3×3 convolutional layer is used to fuse the encoded aperture mask M, where M is copied N times along the channel dimensions to ensure its shape aligns with the input features, ultimately yielding the mask-embedded feature x. ME The formula is as follows: Among them, M r T represents the result of repeating M N times along the channel dimension. I For a full 1 tensor, This represents a convolution operation with a filter size of N×N; The downsampling module performs a 4×4 convolution operation, as shown in the following formula: Where X is the input feature map, and U(X) is the feature map obtained after upsampling. This represents a convolution operation with a filter size of N×N; The upsampling module uses a 2×2 deconvolution operation, as shown in the following formula: Where X is the input feature map, and D(X) is the feature map obtained after downsampling. This represents a deconvolution operation with a filter size of N×N; After two downsampling operations and attention module processing, the feature map output by the U-shaped network is obtained after passing through the intermediate layer and then undergoing two more upsampling operations and attention module processing.
5. The hyperspectral snapshot compression imaging reconstruction method combining memory and depth unfolding according to claim 4, characterized in that: In the denoising network, a spatial-spectral attention module with memory enhancement as its main component is employed, combining spatial self-attention and spectral self-attention. First, the input features pass through a 5×5 deep convolutional layer to perform conditional position encoding, resulting in... After layer normalization, spatial self-attention calculation is performed first, using a memory-enhanced proxy spatial attention mechanism. Before performing self-attention computation, the input features It needs to be compared with the hidden state of the previous iteration phase. To perform fusion, that is, to fuse the current feature x j Compared to the previous stage of hidden state The joint features are concatenated along the channel dimension. Then, they are processed through a memory update gate and a memory reset gate to dynamically adjust the update and reset of the hidden state. Both the memory update gate and the memory reset gate consist of a 3×3 convolutional layer and a sigmoid activation function, and their calculation process can be represented as follows: Where σ(·) represents the Sigmoid activation function, and These represent the results after processing by the memory update gate and the memory reset gate, respectively, and "[]" represents the splicing operation in the channel dimension; Next, the historical hidden state is gated to be appropriately reset at the current moment, and the fusion information is further extracted through the state transition unit: Where ⊙ represents element-wise multiplication, t j The Tanh activation function is used to perform a non-linear transformation on the current information, and "[]" represents the splicing operation on the channel dimension.
6. The hyperspectral snapshot compression imaging reconstruction method combining memory and depth unfolding according to claim 5, characterized in that: After computation in the gated memory unit, for the features A 1×1 convolutional layer and a 3×3 depthwise convolutional layer are used to calculate the query, key, and value, as shown in the following formula: in, This represents a convolution operation with a filter size of N×N. This represents a depthwise convolution operation with a filter size of N×N; Then, a window of size M×M is used for Q. spa K spa V spa Perform non-overlapping partitioning to obtain The computation process of proxy attention can be divided into two core stages: proxy aggregation and proxy broadcasting. In the proxy aggregation stage, proxy label A is used as query Q. spa The agent, with key K spa Calculate attention, extract key information, and in the value V spa Feature aggregation is performed under the mapping to achieve efficient spatial modeling, and its calculation is as follows: Here, Softmax() represents normalizing the attention score along the last dimension. For the calculated surrogate features, d is a scaling factor to prevent gradient vanishing, and in this method it is set to be the same as the window size M; During the proxy broadcast phase, the extracted information V A It is further passed to the original query Q spa This is to enhance feature representation and improve the effectiveness of attention computation; specifically, the surrogate tag A serves as the key K. spa The agent, and then with V A Perform attention calculations to obtain the final attention output: in, This is the final result of the attention calculation; Finally, O undergoes feature transformation via a 1×1 convolutional layer. Subsequently, to further enhance feature diversity, an additional 3×3 deep convolutional layer is used to transform V. spa The process is performed and added to the result of feature transformation to obtain the optimized output feature. After the proxy attention calculation, window spectral self-attention calculation is performed. Within each spatial window, channels of different bands are treated as independent tokens, and self-attention calculation is performed in the spectral dimension. First, the input features are... After performing layer normalization, the query, key, and value are calculated separately using 1×1 convolutions, as shown in the following formulas: in, This represents a convolution operation with a filter size of N×N; subsequently, Q... spe K spe and V spe Divide the M×M window into non-overlapping sections to obtain the shape as follows The window-level features are then analyzed; finally, a spectral self-attention operation is performed within the window, as shown in the following formula: Where d represents the scaling factor, and in this method it is set to be the same as the window size M.
7. The hyperspectral snapshot compression imaging reconstruction method combining memory and depth unfolding according to claim 6, characterized in that: During the K iterations of the unfolded network, linear projection and denoising are performed alternately, continuously approximating the solution space jointly determined by measurement constraints and data priors; after K iterations, the final reconstructed hyperspectral image v is obtained. K If the true hyperspectral image is denoted as The overall network loss function adopts a staged weighted loss strategy, defined as follows: in, Represents the losses in the final reconstruction phase. These represent the losses of the three stages preceding the final stage, and are assigned different weight coefficients to reflect the optimization contribution of multi-stage training. The loss function for each stage k is defined as follows: Where, ‖·‖2 is the two-dimensional Euclidean norm, used to calculate the pixel-level error between the current estimate and the reference true value.
Citation Information
Patent Citations
Remote sensing image space and spectral resolution recovery method and system
CN114549306A
Hyperspectral image super-resolution method based on transposed convolutional long-short term memory network
CN116563113A
Two-stage multi-scale hyperspectral snapshot compression imaging image reconstruction method
CN117974909A
Hyperspectral image compression imaging method and device based on spectral memory enhancement network
CN118379374A
Self-adaptive depth prior compressed sensing spectrum reconstruction method and system
CN119625107A