A hyperspectral snapshot compressive imaging reconstruction method fusing conditional diffusion

CN120976433BActive Publication Date: 2026-08-11CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511108677.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2026-08-11
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

然而,简单地在端到端网络中引入Transformer模型并不能充分应对噪声问题,同时在处理成像过程中的复杂的空间-光谱信息还原方面仍有较大改进空间

Benefits of technology

[0071]This invention proposes a hyperspectral snapshot compressed imaging reconstruction method (HDiff-HIR) based on conditional diffusion. It utilizes the iterative nature of the diffusion generation model to achieve high-fidelity reconstruction, effectively tapping the potential of deep generative models in hyperspectral image reconstruction tasks. The invention proposes a conditional generation module (MCGM) that integrates two-dimensional measurement information and coded aperture information, embedding the generated conditions hierarchically into different stages of the denoising network. This guides the network to progressively optimize the reconstruction results, fully utilizing measurement and prior information to improve the accuracy and stability of the reconstruction. Furthermore, this invention proposes a novel attention mechanism, Local-Global-Spectral Enhanced Multi-Head Self-Attention (LGS-MSA), which simultaneously captures local detail information, long-range dependencies, and fine spectral features, effectively improving the model's feature representation ability and reconstruction quality. Experimental results on three hyperspectral datasets demonstrate that the proposed HDiff-HIR outperforms existing state-of-the-art methods in terms of reconstruction accuracy, denoising capability, and artifact suppression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976433B_ABST
    Figure CN120976433B_ABST
Patent Text Reader

Abstract

This invention relates to a hyperspectral snapshot compressed imaging reconstruction method with fused conditional diffusion, belonging to the field of deep learning technology. The method includes the following steps: S1: acquiring a two-dimensional compressed image of the target scene; S2: preprocessing the hyperspectral image training set to generate sample data for training; S3: inputting the preprocessed training samples into a hierarchical conditional diffusion image reconstruction network for training; S4: after the network training is completed, inputting test samples into the network to reconstruct the hyperspectral image and obtain the result. This method, while reconstructing the detailed structure of hyperspectral images, possesses stronger denoising capabilities and has significant advantages over other methods in suppressing artifacts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning technology and relates to a hyperspectral snapshot compression imaging reconstruction method that incorporates conditional diffusion. Background Technology

[0002] Traditional hyperspectral imaging methods typically rely on the movement of large mechanical or optical components to acquire a complete spectral data cube. These methods suffer from limitations such as high system cost and slow imaging speed. In contrast, hyperspectral snapshot compression imaging technology compresses a three-dimensional hyperspectral image into a two-dimensional measurement that simultaneously contains spectral and spatial information. This allows for the acquisition of both spatial and spectral information of the target in a single exposure, enabling real-time scene monitoring. Furthermore, snapshot compression imaging systems are more compact, reducing the use of mechanical components, lowering system complexity and failure rate, and significantly improving imaging efficiency.

[0003] Among these snapshot compression imaging systems, the coded aperture snapshot spectral imaging system is a highly efficient technique. It combines a coded aperture with a spectroscopic element to acquire spatial-spectral information in a single exposure and reconstructs the complete hyperspectral data cube through optimized reconstruction algorithms. The core advantage of this snapshot compression imaging system lies in its optical encoding of hyperspectral data, enabling compressed sensing methods to directly encode hyperspectral information during the data acquisition phase, thereby significantly reducing sampling requirements and improving system imaging speed. Its core challenge lies in accurately reconstructing a three-dimensional hyperspectral signal from limited two-dimensional measurements, a process that is typically an ill-posed inverse problem.

[0004] Traditional hyperspectral image reconstruction algorithms rely on predefined artificial priors to regularize the reconstruction process, obtaining reconstruction results by constructing an optimization problem with prior constraints. These methods typically employ techniques such as sparse representation, low-rank constraints, or total variational regularization. However, since the design of artificial priors often depends on specific scenarios, the expressive power of traditional methods is limited, making it difficult to fully capture the complex features of hyperspectral data. In recent years, deep learning-based methods have been widely used. These methods can adaptively learn more expressive feature representations from large-scale data, with typical methods being Transformer-based end-to-end training methods. However, simply introducing a Transformer model into an end-to-end network cannot adequately address noise issues, and there is still considerable room for improvement in handling the complex spatial-spectral information restoration during the imaging process. Although some research has attempted to use diffusion models to address these problems, these methods typically rely on plug-and-play approaches, where the model training phase and the diffusion process are often independent, lacking deep coupling and hindering model generalization and end-to-end optimization. Furthermore, in order to further improve the adaptability of the diffusion model in hyperspectral snapshot compressed imaging tasks, it is still necessary to study its fusion mechanism with prior information of the imaging system (such as coded aperture) to achieve better reconstruction results. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a hyperspectral snapshot compressed imaging reconstruction method that integrates conditional diffusion. It employs an integrated framework consisting of a conditional generation module and a diffusion reconstruction network to guide the reconstruction process of hyperspectral images. Through multi-layered conditional guidance, the generation process can be effectively constrained in both spatial and spectral dimensions, thereby better integrating measurement information and imaging priors. A novel attention mechanism, Local-Global-Spectral Enhanced Multi-Head Self-Attention (LGS-MSA), is used in the denoising network. LGS-MSA is divided into two branches. One branch performs window self-attention calculation on the query, key, and value after depthwise convolution, while the other branch performs self-attention calculation on the key and value using a larger window after dilated convolution. Finally, the calculation results of these two branches are fused. This approach can simultaneously capture local details and global contextual information, effectively enhancing the spatial-spectral feature representation capability of hyperspectral data. This method not only reconstructs the detailed structure of hyperspectral images but also possesses stronger denoising capabilities and exhibits significant advantages over other methods in suppressing artifacts.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A hyperspectral snapshot compressed imaging reconstruction method with fusion conditional diffusion is disclosed, comprising the following steps: S1: acquiring a two-dimensional compressed image of the target scene; S2: preprocessing the hyperspectral image training set to generate sample data for training; S3: inputting the preprocessed training samples into a hierarchical conditional diffusion image reconstruction network for training; S4: after the network training is completed, inputting the test samples into the network to reconstruct the hyperspectral image and obtain the result.

[0008] 1. Further, in step S1, a two-dimensional compressed image is acquired using a coded aperture snapshot spectral imaging system, specifically including: S11, selecting the spectral range, setting the working wavelength to 450nm-650nm, and dividing it into 28 channels; S12, designing the coded aperture to introduce different optical codes at each pixel; S13, setting the parameters of the coded aperture snapshot spectral imaging system, including but not limited to the number of dispersive elements and the dispersive shift step size; S14, capturing a compressed measurement image of the entire scene through an optical sensor.

[0009] 2. Further, in step S2, the hyperspectral images used as the training set are preprocessed, specifically including: S21, dividing the hyperspectral images into two parts; the first part uses a 256×256 non-overlapping window to generate samples, and the second part uses a 128×128 non-overlapping window to divide the image, and randomly combining them into several 256×256 samples; S22, performing data augmentation on the two parts respectively, including random rotation and flipping; S23, using a simulated imaging system, that is, after the hyperspectral images are modulated by a pre-designed physical mask, they are sheared, shifted, and summed to generate corresponding two-dimensional compressed measurements. The data-augmented hyperspectral images and their corresponding two-dimensional compressed measurements together constitute training sample pairs.

[0010] 3. Further, step S3 includes two processes: a forward diffusion process and a backward reasoning process. The forward process involves gradually adding Gaussian noise to the original hyperspectral image X0 through a Markov chain within T steps to generate a pure Gaussian noise image X. T ~N(0,I), the formula is as follows:

[0011]

[0012] Among them, X T This represents the pure noise image at the end of the forward diffusion process, where I is the identity matrix and α 1:T ∈(0,1) is a scalar hyperparameter that determines the noise variance introduced in each iteration.

[0013] Based on reparameterization technology, X t Given X0, the distribution can be directly sampled by omitting the intermediate transformation process, thus simplifying the calculation. The formula is as follows:

[0014]

[0015] in, Therefore, the output of each diffusion step can be obtained directly, as shown in the following formula:

[0016]

[0017] Furthermore, the reverse reasoning process starts with pure Gaussian noise and gradually optimizes the sample over T iterations, as shown in the following formula:

[0018]

[0019] p(X T )=N(X T ;0,I)

[0020] p θ (X t-1 |X t ,Y,M)=N(X t-1 μ θ (X t ,Y,M,γ t ),σ 2 I)

[0021] Wherein, the conditional distribution p θ (X t-1 |X t Let Y, M represent the inference process during training. This method uses a denoising network to predict... The Denoising Diffusion Implicit Model (DDIM) is employed to accelerate sampling. Its forward diffusion process can be modeled as a non-Markovian process, and the reverse process iteratively samples X. t-1 The formula is as follows:

[0022]

[0023] Where, ∈ θ (X t (,Y,M,t) represents from X θ (X t The noise variable calculated from (Y, M, t) is given by the following formula:

[0024]

[0025] DDIM uses a subsequence of the original time step, i.e., {τ1, τ2, ..., τ...} n The core idea of ​​the method is skip sampling, which reduces the number of sampling steps from T to n without significantly reducing the reconstruction quality.

[0026] Next, the noisy image is input into the denoising network. The denoising network consists of a U-shaped network, comprising three stages: downsampling, intermediate layers, and upsampling. The downsampling module uses a 4×4 convolution operation, as shown in the following formula:

[0027]

[0028] Where X is the input feature map, and U(X) is the feature map obtained after upsampling. This represents a convolution operation with a filter size of N×N.

[0029] The upsampling module uses a 2×2 deconvolution operation, as shown in the following formula:

[0030]

[0031] Where X is the input feature map, and D(X) is the feature map obtained after downsampling. This represents a deconvolution operation with a filter size of N×N.

[0032] After two downsampling operations and attention module processing, the feature map output by the U-shaped network is obtained after passing through the intermediate layer and then undergoing two more upsampling operations and attention module processing.

[0033] Furthermore, in the denoising network, the main components are Local-Global-Spectral Enhanced Attention Block (LGSAB) and its simplified version, Local-Spectral Enhanced Attention Block (LSAB). The main components of the module are Local-Global-Spectral Enhanced Multi-Head Self-Attention (LGS-MSA) and Local-Spectral Enhanced Multi-Head Self-Attention (LS-MSA), respectively. LGS-MSA and LS-MSA first transform the input feature map into query value, key value, and attribute value through three linear layers, respectively. They then calculate the temporal embedding using sine-cosine positional encoding at each time step, and obtain the temporal token through the Swish activation function and MLP layer. This token is also processed through linear layers, and the final query value, key value, and attribute value are obtained by summing them, as shown in the following formula:

[0034]

[0035] Where X represents the input feature map, query, key, and value represent the obtained query value, key value, and attribute value, respectively, and W... q W k W v These represent the projected weights of the query value, key value, and attribute value, respectively.

[0036] For LGS-MSA, the query value, key value, and attribute value are divided into two equal parts along the channel dimension, resulting in [Q]. l ,K l Vl ] and [Q g ,K g V g The process is divided into local and global branches, which are processed separately. The local branch first uses a 3×3 depthwise convolutional layer for local feature enhancement, as shown in the following formula:

[0037]

[0038] in, This indicates a 3×3 depth convolutional layer.

[0039] Then, The partitioning is performed using non-overlapping windows (M1×M1), resulting in: Subsequently, they are divided into multiple attention heads along the channel dimension, as shown in the following formula:

[0040]

[0041] Where h represents the number of attention heads. The self-attention of each head is calculated as follows:

[0042]

[0043] Here, Softmax() represents the normalization process for the attention score, and d is the dimension of each head. These are learnable parameters that represent additional location information;

[0044] The global branch first uses a 3×3 dilated convolution to process K. g and V g The formula is as follows:

[0045]

[0046] in, This represents a 3×3 hollow convolutional layer.

[0047] Then, the query, key, and value are partitioned using a larger non-overlapping window (M2×M2) compared to the window size used in the local branch, resulting in... The subsequent operations are the same as for the local branch. Therefore, the self-attention of the global branch can be calculated as follows:

[0048]

[0049] Finally, the local and global attention calculation results are concatenated and a linear transformation is applied to obtain the final attention output:

[0050]

[0051] Where Concat[] represents concatenating the channel dimensions of the outputs of multiple attention heads, and Linear() represents a linear transformation, i.e., a fully connected layer;

[0052] For LS-MSA, the complete query, key, and value are processed directly without bisecting the channel, and then the same processing method as the local branch of LGS-MSA is used to calculate the attention.

[0053] Furthermore, in LGSAB and LSAB, a residual is first calculated by performing a 3×3 depthwise convolution on the feature map, and then added to the input feature map for conditional location encoding, as shown in the following formula:

[0054]

[0055] Where X is the input feature map, and CPE(X) is the feature map after conditional location encoding.

[0056] Then, layer normalization is applied to the position-encoded feature map, followed by LGS-MSA or LS-MSA processing to obtain the residual, which is then added to the input feature map, as shown in the following formula:

[0057] LGS(X)=X+LayerNorm(LGSMSA(X))

[0058] LS(X) = X + LayerNorm(LSMSA(X))

[0059] Where X is the input feature map, LGS() and LS() represent the feature maps after passing through the LGSAB module and LSAB module, respectively, LayerNorm() represents layer normalization, LGSMA() represents Local-Global-Spectral Enhanced Multi-Head Self-Attention (LGS-MSA), and LSMSA() represents Local-Spectral Enhanced Multi-Head Self-Attention (LS-MSA).

[0060] Finally, the result after self-attention calculation is normalized again, and then passed through a feedforward neural network. The feedforward neural network contains 1×1 convolutions, 3×3 depthwise convolutions, and 1×1 convolutions. Before each two convolutions, the GELU activation function is applied, as shown in the following formula:

[0061]

[0062] Where X is the input feature map, and F(X) represents the feature map after processing by the attention module.

[0063] Furthermore, the two-dimensional compressed measurement training samples processed by the simulated imaging system are initialized to generate conditions for diffusion generation. First, by performing sliding extraction along the channel dimension, each column in the compressed measurement is restored to a specific position in the output hyperspectral image, thereby transforming the original two-dimensional data into a coarse three-dimensional hyperspectral image with higher dimensions. Then, the physical mask of the imaging system is stitched with this hyperspectral image along the channel dimension, and then a 1×1 convolution is performed to restore the number of channels to the number of channels in the original hyperspectral image, obtaining the initial input to the network, as shown in the following formula:

[0064]

[0065] Where X is the input feature map, M is the physical mask, and I(X) is the obtained initial input. The symbol represents a convolution operation with a filter size of N×N, and "[]" represents a concatenation operation along the channel dimension.

[0066] This is then passed to different branches. Each branch consists of a 3×3 convolutional layer followed by an LSAB to generate input conditions for each layer of the denoising network (downsampling is performed using 4×4 convolutional layers if necessary). The generated input conditions are then embedded into each layer of the U-shaped network encoder, which effectively utilizes multi-scale features to generate more accurate results during the reverse inference process.

[0067] Furthermore, by setting the prediction target to the original data X0, the following loss function for network training is obtained:

[0068] Loss=||X0-X θ (X t ,Y,M,t)||1

[0069] Among them, X θ (X t Let (Y, M, t) represent the hyperspectral image predicted by the denoising network under given conditions (Y, M) and temporal embedding, and ||·||1 represent the L1 norm, i.e., the sum of absolute errors between pixels. A loss function is calculated, and the model parameters of the hyperspectral image reconstruction framework are optimized based on the cross-loss function and the backpropagation process. After training, a trained hyperspectral image reconstruction framework is obtained. The input samples are reconstructed using the trained hyperspectral image reconstruction framework, and the resulting hyperspectral image reconstruction is output.

[0070] The beneficial effects of this invention are as follows:

[0071] This invention proposes a hyperspectral snapshot compressed imaging reconstruction method (HDiff-HIR) based on conditional diffusion. It utilizes the iterative nature of the diffusion generation model to achieve high-fidelity reconstruction, effectively tapping the potential of deep generative models in hyperspectral image reconstruction tasks. The invention proposes a conditional generation module (MCGM) that integrates two-dimensional measurement information and coded aperture information, embedding the generated conditions hierarchically into different stages of the denoising network. This guides the network to progressively optimize the reconstruction results, fully utilizing measurement and prior information to improve the accuracy and stability of the reconstruction. Furthermore, this invention proposes a novel attention mechanism, Local-Global-Spectral Enhanced Multi-Head Self-Attention (LGS-MSA), which simultaneously captures local detail information, long-range dependencies, and fine spectral features, effectively improving the model's feature representation ability and reconstruction quality. Experimental results on three hyperspectral datasets demonstrate that the proposed HDiff-HIR outperforms existing state-of-the-art methods in terms of reconstruction accuracy, denoising capability, and artifact suppression.

[0072] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0073] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0074] Figure 1 This is a flowchart of the method of the present invention;

[0075] Figure 2 This is a schematic diagram of the Coding Aperture Snapshot Spectral Imaging System (CASSI).

[0076] Figure 3 This is a schematic diagram of the hyperspectral image reconstruction process based on a diffusion model.

[0077] Figure 4 The diagram shows the overall architecture of the conditional diffusion-based hyperspectral snapshot compressed imaging reconstruction method (HDiff-HIR), where (a) is the overall architecture of the hierarchical conditional denoising diffusion probability model, (b) is the mask integration conditional generation module, (c) is the local-global-spectral enhancement attention block, and (d) is the local-spectral enhancement attention block.

[0078] Figure 5 The diagram shows the structure of the attention mechanism of the present invention, wherein (a) local-global-spectral enhanced multi-head self-attention, and (b) local-spectral enhanced multi-head self-attention;

[0079] Figure 6 Visualization results of different methods on four channels of a scene on the KAIST hyperspectral image dataset, where (a) TwIST, (b) TSA-Net, (c) MST, (d) DWMT, (e) HDiff-HIR, and (f) ground truth map;

[0080] Figure 7 Visualization results of different methods on two channels of a scene on a real CASSI image dataset, where (a) TwIST, (b) TSA-Net, (c) MST, (d) DWMT, and (e) HDiff-HIR. Detailed Implementation

[0081] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.

[0082] Figure 1 This is a flowchart of the method of the present invention. The present invention provides a hyperspectral snapshot compressed imaging reconstruction method with fused conditional diffusion. As shown in the figure, in the image acquisition stage, a coded aperture snapshot spectral imaging system is used to acquire two-dimensional compressed images. The conditional diffusion model used for hyperspectral image reconstruction is as follows: Figure 4 As shown, this method guides the diffusion process to gradually recover high-quality hyperspectral images in both spatial and spectral dimensions by introducing conditional information generated from two-dimensional measurements and encoded aperture. Specifically, the proposed model employs a Conditional Generation Module (MCGM) to effectively fuse measurement information with prior knowledge and embeds hierarchically generated conditions into different stages of the denoising network, guiding the network to gradually denoise and reconstruct image details during back-inference. The Local-Global-Spectral Enhancement Attention Block (LGSAB) in the network uses a local-global-spectral enhancement multi-head self-attention mechanism, which can simultaneously capture local detail information, long-range dependencies, and spectral features, thereby further improving reconstruction quality and stability. The overall scheme fully combines the capabilities of measurement constraints and generative models, effectively improving the accuracy, robustness, and artifact suppression of reconstruction. Specifically, the technical solution of this invention includes the following:

[0083] 1. Data Acquisition: Snapshot compressed image acquisition is achieved using a coded aperture snapshot spectral imaging system. First, the spectral range is selected, with the working wavelength set to 450nm-650nm, divided into 28 channels. Then, a coded aperture is designed to introduce different optical codes at each pixel. Next, the parameters of the coded aperture snapshot spectral imaging system are set, including but not limited to the number of dispersive elements and the dispersion shift step size. Finally, a compressed measurement image of the entire scene is captured using an optical sensor.

[0084] 2. Data Preprocessing: The hyperspectral images in the training set are preprocessed. First, the hyperspectral images are divided into two parts. The first part uses a 256×256 non-overlapping window to generate samples, and the second part uses a 128×128 non-overlapping window to divide the image, which is then randomly combined into several 256×256 samples. Next, data augmentation is performed on both parts, including random rotation and flipping. Then, using a simulated imaging system, the hyperspectral images are modulated using a pre-designed physical mask, then sheared, shifted, and summed to generate corresponding two-dimensional compressed measurements. The data-augmented hyperspectral images and their corresponding two-dimensional compressed measurements together constitute training sample pairs.

[0085] 3. Construct a forward diffusion process, such as Figure 3 As shown, the forward process uses a Markov chain to progressively add Gaussian noise to the original hyperspectral image X0 within T steps, generating a pure Gaussian noise image X. T ~N(0,I), the formula is as follows:

[0086]

[0087] Among them, X T This represents the pure noise image at the end of the forward diffusion process, where I is the identity matrix and α 1:T ∈(0,1) is a scalar hyperparameter that determines the noise variance introduced in each iteration.

[0088] Based on reparameterization technology, X t Given X0, the distribution can be directly sampled by omitting the intermediate transformation process, thus simplifying the calculation. The formula is as follows:

[0089]

[0090] in, Therefore, the output of each diffusion step can be obtained directly, as shown in the following formula:

[0091]

[0092] 4. Construct a reverse reasoning process, such as... Figure 3 As shown. The reverse inference process starts with pure Gaussian noise and optimizes the sample step by step over T iterations, as shown in the following formula:

[0093]

[0094] p(X T )=N(X T ;0,I)

[0095] p θ (X t-1 |X t,Y,M)=N(X t-1 μ θ (X t ,Y,M,γ t ),σ 2 I)

[0096] Wherein, the conditional distribution p θ (X t-1 |X t Let Y, M represent the inference process during training. This method uses a denoising network to predict... The Denoising Diffusion Implicit Model (DDIM) is employed to accelerate sampling. Its forward diffusion process can be modeled as a non-Markovian process, and the reverse process iteratively samples X. t-1 The formula is as follows:

[0097]

[0098] Where, ∈ θ (X t (,Y,M,t) represents from X θ (X t The noise variable calculated from (Y, M, t) is given by the following formula:

[0099]

[0100] DDIM uses a subsequence of the original time step, i.e., {τ1, τ2, ..., τ...} n The core idea of ​​the method is skip sampling, which reduces the number of sampling steps from T to n without significantly reducing the reconstruction quality.

[0101] 5. Input the noisy image into the denoising network. The denoising network consists of a U-shaped network, including three stages: downsampling, intermediate layers, and upsampling. The downsampling module uses a 4×4 convolution operation, as shown in the following formula:

[0102]

[0103] Where X is the input feature map, and U(X) is the feature map obtained after upsampling. This represents a convolution operation with a filter size of N×N.

[0104] The upsampling module uses a 2×2 deconvolution operation, as shown in the following formula:

[0105]

[0106] Where X is the input feature map, and D(X) is the feature map obtained after downsampling. This represents a deconvolution operation with a filter size of N×N.

[0107] After two downsampling operations and attention module processing, the feature map output by the U-shaped network is obtained after passing through the intermediate layer and then undergoing two more upsampling operations and attention module processing.

[0108] 6. In the denoising network, the main components are Local-Global-Spectral Enhanced Attention Block (LGSAB) and its simplified version, Local-Spectral Enhanced Attention Block (LSAB). The main components of the module are Local-Global-Spectral Enhanced Multi-Head Self-Attention (LGS-MSA) and Local-Spectral Enhanced Multi-Head Self-Attention (LS-MSA), respectively. LGS-MSA and LS-MSA first transform the input feature map into query value, key value, and attribute value through three linear layers, respectively. They then calculate the temporal embedding using sine-cosine positional encoding at each time step, and obtain the temporal token through the Swish activation function and MLP layer. This token is also processed through linear layers, and the final query value, key value, and attribute value are obtained by summing them, as shown in the following formula:

[0109]

[0110] Where X represents the input feature map, query, key, and value represent the obtained query value, key value, and attribute value, respectively, and W... q W k W v These represent the projected weights of the query value, key value, and attribute value, respectively.

[0111] For LGS-MSA, the query value, key value, and attribute value are divided into two equal parts along the channel dimension, resulting in [Q]. l ,K l V l ] and [Q g ,K g V g The process is divided into local and global branches, which are processed separately. The local branch first uses a 3×3 depthwise convolutional layer for local feature enhancement, as shown in the following formula:

[0112]

[0113] in, This indicates a 3×3 depth convolutional layer.

[0114] Then, The partitioning is performed using non-overlapping windows (M1×M1), resulting in: Subsequently, they are divided into multiple attention heads along the channel dimension, as shown in the following formula:

[0115]

[0116] Where h represents the number of attention heads. The self-attention of each head is calculated as follows:

[0117]

[0118] Where d is the dimension of each head. These are learnable parameters that represent additional location information.

[0119] The global branch first uses a 3×3 dilated convolution to process K. g and V g The formula is as follows:

[0120]

[0121] in, This represents a 3×3 hollow convolutional layer.

[0122] Then, the query, key, and value are partitioned using a larger non-overlapping window (M2×M2) compared to the window size used in the local branch, resulting in... The subsequent operations are the same as for the local branch. Therefore, the self-attention of the global branch can be calculated as follows:

[0123]

[0124] Finally, the local and global attention calculation results are concatenated and a linear transformation is applied to obtain the final attention output:

[0125]

[0126] Where Concat[] represents concatenating the channel dimensions of the outputs of multiple attention heads, and Linear() represents a linear transformation, i.e., a fully connected layer;

[0127] For LS-MSA, the complete query, key, and value are processed directly without bisecting the channel, and then the same processing method as the local branch of LGS-MSA is used to calculate the attention.

[0128] 7. In LGSAB and LSAB, the feature map is first subjected to a 3×3 depthwise convolution to calculate the residual, which is then added to the input feature map for conditional location encoding. The formula is as follows:

[0129]

[0130] Where X is the input feature map, and CPE(X) is the feature map after conditional location encoding.

[0131] Then, layer normalization is applied to the position-encoded feature map, followed by LGS-MSA or LS-MSA processing to obtain the residual, which is then added to the input feature map, as shown in the following formula:

[0132] LGS(X)=X+LayerNorm(LGSMSA(X))

[0133] LS(X) = X + LayerNorm(LSMSA(X))

[0134] Where X is the input feature map, LGS() and LS() represent the feature maps after passing through the LGSAB module and LSAB module, respectively, LayerNorm() represents layer normalization, LGSMA() represents Local-Global-Spectral Enhanced Multi-Head Self-Attention (LGS-MSA), and LSMSA() represents Local-Spectral Enhanced Multi-Head Self-Attention (LS-MSA).

[0135] Finally, the result after self-attention computation undergoes another layer normalization operation, followed by a feedforward neural network. The feedforward neural network contains 1×1 convolutions, 3×3 depthwise convolutions, and 1×1 convolutions, with GELU activation applied before each convolution, as shown in the following formula:

[0136]

[0137] Where X is the input feature map, and F(X) represents the feature map after processing by the attention module.

[0138] 8. Initialize the training samples of the two-dimensional compressed measurement processed by the simulated imaging system to generate conditions for diffusion generation. First, by performing sliding extraction in the channel dimension, each column in the compressed measurement is restored to a specific position in the output hyperspectral image, thereby converting the original two-dimensional data into a coarse three-dimensional hyperspectral image with higher dimensions. Then, the physical mask of the imaging system is stitched with this hyperspectral image in the channel dimension, and then a 1×1 convolution is performed to restore the number of channels to the number of channels in the original hyperspectral image, obtaining the initial input of the network, as shown in the following formula:

[0139]

[0140] Where X is the input feature map, M is the physical mask, and I(X) is the obtained initial input. The symbol represents a convolution operation with a filter size of N×N, and "[]" represents a concatenation operation along the channel dimension.

[0141] This is then passed to different branches. Each branch consists of a 3×3 convolutional layer followed by an LSAB to generate input conditions for each layer of the denoising network (downsampling is performed using 4×4 convolutional layers if necessary). The generated input conditions are then embedded into each layer of the U-shaped network encoder, which effectively utilizes multi-scale features to generate more accurate results during the reverse inference process.

[0142] 9. Setting the prediction target as the original data X0, we obtain the following loss function for network training:

[0143] Loss=||X0-X θ (X t ,Y,M,t)||1

[0144] Among them, X θ (X t Let (Y, M, t) represent the hyperspectral image predicted by the denoising network under given conditions (Y, M) and temporal embedding, and ||·||1 represent the L1 norm, i.e., the sum of absolute errors between pixels. A loss function is calculated, and the model parameters of the hyperspectral image reconstruction framework are optimized based on the cross-loss function and the backpropagation process. After training, a trained hyperspectral image reconstruction framework is obtained. The input samples are reconstructed using the trained hyperspectral image reconstruction framework, and the resulting hyperspectral image reconstruction is output.

[0145] like Figure 6 The experimental results of the conditional diffusion-based hyperspectral snapshot compression imaging reconstruction method described in this invention on an open-source hyperspectral natural scene show that the scene details are well recovered, and the reconstructed image quality is high and highly consistent with the real scene. The reconstruction effect of this invention can be further illustrated by comparative experiments. The method of this invention was compared with other existing methods such as GAP-TV, DeSCI, TSA-Net, HDNet, MST, CST, and DWMT on a hyperspectral natural scene dataset. Peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) were calculated respectively. A higher PSNR indicates relatively less distortion and higher hyperspectral image quality; a higher SSIM indicates that the image structure, brightness, and contrast are more similar to the real image, resulting in higher image quality. Table 1 shows the average reconstruction results of ten scenes selected from the KAIST test set using different methods.

[0146] Table 1. Comparison of HDiff-HIR with various methods on the KAIST hyperspectral image dataset (average of ten scenes).

[0147]

[0148] The method of this invention achieves the best reconstruction accuracy on this dataset. Figure 7 The presentation showcases the visualized reconstruction results of different methods on a typical scene in the real CASSI dataset. It can be seen that the method proposed in this invention significantly outperforms other hyperspectral image reconstruction methods in terms of reconstruction quality, enabling more refined recovery of image details and exhibiting a clear advantage in suppressing artifacts, resulting in a more natural and realistic reconstruction.

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications should be covered within the scope of the claims of the present invention.

Claims

1. A hyperspectral snapshot compression imaging reconstruction method with fusion conditional diffusion, characterized in that: The method includes the following steps: S1: Acquire a two-dimensional compressed image of the target scene; S2: Preprocess the hyperspectral image training set to generate sample data for training; S3: Input the preprocessed training samples into the hierarchical conditional diffusion image reconstruction network for training; S4: After the network training is completed, the test samples are input into the network to reconstruct the hyperspectral image and obtain the results; The noisy image is input into the denoising network; the denoising network consists of a U-shaped network, including three stages: downsampling, intermediate layers, and upsampling; the downsampling module uses a 4×4 convolution operation, as shown in the following formula: Where X is the input feature map, and U(X) is the feature map obtained after upsampling. This represents a convolution operation with a filter size of 4×4; The upsampling module uses a 2×2 deconvolution operation, as shown in the following formula: Where X is the input feature map, and D(X) is the feature map obtained after downsampling. This represents a deconvolution operation with a filter size of 2×2; After two downsampling operations and attention module processing, the feature map output by the U-shaped network is obtained after passing through the intermediate layer and then two upsampling operations and attention module processing. In the denoising network, the main components are Local-Global-Spectral Enhanced Attention Block (LGSAB) and its simplified version, Local-Spectral Enhanced Attention Block (LSAB). The main components of the module are Local-Global-Spectral Enhanced Multi-Head Self-Attention (LGS-MSA) and Local-Spectral Enhanced Multi-Head Self-Attention (LS-MSA). LGS-MSA and LS-MSA first transform the input feature map into query value, key value, and attribute value through three linear layers, respectively. They then calculate the temporal embedding using sine-cosine positional encoding at each time step, and obtain the temporal token through the Swish activation function and MLP layer. This token is also processed through linear layers, and the final query value, key value, and attribute value are obtained by summing them, as shown in the following formula: in, and This represents the input feature map, where Q, K, and V represent the obtained query value, key value, and attribute value, respectively. and Representing the input features respectively and The weights projected onto the query value. and Representing the input features respectively and The weights projected onto the key values. and Representing the input features respectively and Weights projected onto the query value; For LGS-MSA, the query value, key value, and attribute value are divided into two equal parts along the channel dimension, resulting in: and The process is divided into local and global branches, which are processed separately. The local branch first uses a 3×3 deep convolutional layer to enhance local features, obtaining the enhanced query value, key value, and attribute value. The calculation formula is as follows: in, , , This represents three different 3×3 depth convolutional layers; Then, The partition is performed using non-overlapping windows, with each window measuring M1×M1. The result is... Where H, W, and C represent height, width, and number of channels, respectively; subsequently, they are divided into multiple attention heads along the channel dimension, as shown in the following formula: Where h is the number of attention heads, , , Given the query value, key value, and attribute value of the i-th head, the self-attention features of the i-th head are calculated as follows: in, () indicates that the attention score is normalized, and d is the dimension of each head. These are learnable parameters that represent additional location information; The global branch is first processed using a 3×3 dilated convolution. and The formula is as follows: in, , This represents two different 3×3 hollow convolutional layers; Then, the query value, key value, and attribute value are divided using a larger non-overlapping window than the window size corresponding to the local branch. The window size is M2×M2, resulting in... Subsequently, they are divided into h attention heads along the channel dimension, and the query value, key value, and attribute value of the i-th attention head are represented as follows: , , Therefore, the i-th head self-attention feature of the global branch can be calculated as follows: Finally, the local and global attention calculation results are concatenated and a linear transformation is applied to obtain the final attention output: Where Concat[] represents concatenating the channel dimensions of the outputs of multiple attention heads, and Linear() represents a linear transformation, i.e., a fully connected layer; For LS-MSA, the complete query value, key value, and attribute value are processed directly without channel bisection, and then the same processing method as LGS-MSA local branch is used to calculate attention.

2. The hyperspectral snapshot compression imaging reconstruction method according to claim 1, characterized in that: In step S1, a two-dimensional compressed image is acquired using a coded aperture snapshot spectral imaging system, specifically including the following steps: S11, selecting the spectral range and setting the working wavelength to 450nm-650nm, dividing it into 28 channels; S12, designing the coded aperture to introduce different optical codes at each pixel; S13, setting the parameters of the coded aperture snapshot spectral imaging system, including but not limited to the number of dispersive elements and the dispersive shift step size; S14, capturing a compressed measurement image of the entire scene using an optical sensor. In step S2, the hyperspectral images used as the training set are preprocessed, specifically including the following steps: S21, the hyperspectral images are divided into two parts. The first part is divided into samples using a 256×256 non-overlapping window, and the second part is divided into images using a 128×128 non-overlapping window, and randomly combined into several 256×256 samples; S22, data augmentation is performed on the two parts respectively, including random rotation and flipping; S23, through a simulated imaging system, the hyperspectral images are modulated by a pre-designed physical mask, then sheared, shifted, and summed to generate corresponding two-dimensional compressed measurements; the data-augmented hyperspectral images and their corresponding two-dimensional compressed measurements together constitute training sample pairs; Step S3 includes two processes: a forward diffusion process and a backward reasoning process. The forward process involves inverse reasoning of the original hyperspectral image within step T via a Markov chain. Gaussian noise is gradually added to generate a pure Gaussian noise image. The formula for calculating the joint distribution of the image after adding noise is as follows: The distribution after adding noise in step t is: in, N represents the pure noise image at the end of the forward diffusion process. The Gaussian distribution satisfied after adding noise at step t is... It is the identity matrix. It is a scalar hyperparameter that determines the noise variance introduced in each iteration; Based on reparameterization technology In the given In this case, the distribution can be directly sampled by omitting the intermediate transformation process, thus simplifying the calculation. The distribution calculation formula is as follows: in, Therefore, the output of each diffusion step can be directly obtained, as shown in the following formula: The reverse inference process starts with pure Gaussian noise and optimizes the sample step by step over T iterations, as shown in the following formula: Among them, conditional distribution This represents the reasoning process during training and learning, where N() represents the Gaussian distribution it follows. Indicates 2D compression measurement, This method uses a denoising network to predict the perception matrix. The denoising diffusion implicit model DDIM is used to accelerate sampling. Its forward diffusion process can be modeled as a non-Markovian process, and the reverse process iteratively samples. The formula is as follows: in, This represents the image predicted after denoising at step t. The noise variable calculated in the above formula is as follows: DDIM uses a subsequence of the original time step, i.e. Its core idea is skip sampling, which reduces the number of sampling steps from T to n without significantly reducing the reconstruction quality.

3. The hyperspectral snapshot compression imaging reconstruction method according to claim 1, characterized in that: In LGSAB and LSAB, a 3×3 depthwise convolution is first applied to the feature map. The residual is calculated and then added to the input feature map to perform conditional location encoding, as shown in the following formula: Where X is the input feature map, and CPE(X) is the feature map after conditional location encoding; Then, layer normalization is applied to the position-encoded feature map, followed by LGS-MSA or LS-MSA processing to obtain the residual, which is then added to the input feature map, as shown in the following formula: Where X is the input feature map, LGS() and LS() represent the feature maps after passing through the LGSAB module and LSAB module, respectively, LayerNorm() represents layer normalization, LGSMA() represents Local-Global-Spectral Enhanced Multi-Head Self-Attention (LGS-MSA), and LSMSA() represents Local-Spectral Enhanced Multi-Head Self-Attention (LS-MSA). Finally, the result after self-attention computation undergoes another layer normalization operation, followed by a feedforward neural network; the feedforward neural network contains 1×1 convolutions. 3×3 depthwise convolution 1×1 convolution Before each two convolutions, the GELU activation function is used for activation, as shown in the following formula: Where X is the input feature map, and F(X) represents the feature map after processing by the attention module.

4. The hyperspectral snapshot compression imaging reconstruction method according to claim 1, characterized in that: The training samples from the two-dimensional compressed measurements processed by the simulated imaging system are initialized to generate conditions for diffusion generation. First, by performing sliding extraction along the channel dimension, each column in the compressed measurement is restored to a specific position in the output hyperspectral image, thereby transforming the original two-dimensional data into a coarse three-dimensional hyperspectral image with higher dimensions. Then, the physical mask of the imaging system is stitched with this hyperspectral image along the channel dimension, and then a 1×1 convolution is performed to restore the number of channels to the number of channels in the original hyperspectral image, obtaining the initial input to the network, as shown in the following formula: Where X is the input feature map, M is the physical mask, and I(X) is the obtained initial input. This indicates a convolution operation with a filter size of 1×1, and "[]" indicates a concatenation operation along the channel dimension; This is then passed to different branches; each branch includes a 3×3 convolutional layer followed by LSAB to generate input conditions for each layer of the denoising network; the generated input conditions are then embedded into each layer of the U-shaped network encoder, which can effectively utilize multi-scale features to generate more accurate results during the reverse inference process.

5. The hyperspectral snapshot compression imaging reconstruction method according to claim 1, characterized in that: Set the prediction target as the original data. Thus, the following loss function for network training is obtained: in, Indicates under given conditions In the case of temporal embedding, denoising networks predict images. Indicates 2D compression measurement, Represents the perception matrix, The L1 norm is represented, which is the sum of absolute errors between pixels; the loss function is calculated, and the model parameters of the hyperspectral image reconstruction framework are optimized based on the loss function and the backpropagation process. After training, the trained hyperspectral image reconstruction framework is obtained; the input samples are reconstructed using the trained hyperspectral image reconstruction framework, and the hyperspectral image reconstruction result is output.

Citation Information

Patent Citations

  • Hyperspectral reconstruction method and system based on CASSI optical system

    CN118052897A

  • Medical image unified diffusion generation method oriented to random mode deficiency and related device

    CN120107393A