Hyperspectral snapshot compression imaging reconstruction method fusing conditional diffusion
By combining a conditional generation module and a diffusion reconstruction network with a local-global-spectral enhancement multi-head self-attention mechanism, the problems of noise processing and complex information restoration in hyperspectral imaging are solved, achieving efficient and accurate hyperspectral image reconstruction.
Patent Information
- Application Number
- CN202511108677.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Traditional hyperspectral imaging methods rely on large mechanical components, which are costly and slow. Existing deep learning methods are insufficient in noise processing and complex information reconstruction in hyperspectral snapshot compressed imaging, and lack deep coupling and fusion with prior information of the imaging system.
By employing a conditional generation module and a diffusion reconstruction network, combined with a local-global-spectral enhanced multi-head self-attention mechanism, the reconstruction process is guided by multiple layers of conditions to capture spatial and spectral information, enhance denoising capabilities, and suppress artifacts.
It achieves high-fidelity and efficient hyperspectral image reconstruction, improves reconstruction accuracy and stability, and significantly improves noise processing and artifact suppression.
Smart Images

Figure CN120976433A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of deep learning, and relates to a hyperspectral snapshot compressive imaging reconstruction method fusing conditional diffusion. BACKGROUND
[0002] Traditional hyperspectral imaging methods usually rely on the movement of large mechanical or optical components to obtain a complete spectral data cube. These methods have limitations such as high system cost and slow imaging speed. In contrast, hyperspectral snapshot compressive imaging technology can simultaneously obtain spatial and spectral information of a target in a single exposure by compressing a three-dimensional hyperspectral image into a two-dimensional measurement that contains both spectral and spatial information, thereby realizing real-time monitoring of a scene. In addition, the snapshot compressive imaging system is more compact in structure, reduces the use of mechanical components, reduces system complexity and failure rate, and significantly improves imaging efficiency.
[0003] Among these snapshot compressive imaging systems, the coded aperture snapshot spectral imaging system is a highly efficient technology. It combines coded aperture with a light splitting element to obtain spatial-spectral information in a single exposure and restores a complete hyperspectral data cube through an optimized reconstruction algorithm. The core advantage of this snapshot compressive imaging system is that it modulates hyperspectral data through optical coding, enabling compressive sensing methods to directly encode hyperspectral information during data acquisition, thereby significantly reducing sampling requirements and improving system imaging speed. The core challenge lies in accurately reconstructing three-dimensional hyperspectral signals from limited two-dimensional measurements, which is usually an ill-posed inverse problem.
[0004] Traditional hyperspectral image reconstruction algorithms rely on predefined artificial priors to regularize the reconstruction process, and obtain the reconstruction results by constructing optimization problems with prior constraints. These methods usually use techniques such as sparse representation, low-rank constraint or total variation regularization. However, due to the dependence of artificial prior design on specific scenarios, the expression ability of traditional methods is limited, and it is difficult to fully capture the complex characteristics of hyperspectral data. In recent years, deep learning-based methods have been widely applied. These methods can adaptively learn more expressive feature representations from large-scale data, and typical methods are end-to-end training methods based on Transformer. However, simply introducing a Transformer model into an end-to-end network cannot fully cope with the noise problem, and there is still much room for improvement in restoring complex spatial-spectral information during imaging. Although existing research has attempted to use diffusion models to solve these problems, they usually rely on a plug-and-play method, which often separates the model training stage from the diffusion process, lacks deep coupling, and is not conducive to model generalization and end-to-end optimization. In addition, in order to further improve the adaptability of the diffusion model in the hyperspectral snapshot compressive imaging task, it is still necessary to further study the fusion mechanism of the imaging system prior information (such as the coded aperture), so as to achieve better reconstruction results. SUMMARY
[0005] Therefore, the present application aims to provide a hyperspectral snapshot compressive imaging reconstruction method based on conditional diffusion, which adopts an integrated framework composed of a conditional generation module and a diffusion reconstruction network to guide the reconstruction process of hyperspectral images. Through the guidance of multiple conditions, the generation process can be effectively constrained in the spatial and spectral dimensions, thereby better fusing the measurement information and imaging prior. A new attention mechanism, local-global-spectral enhanced multi-head self-attention (LGS-MSA), is used in the denoising network. LGS-MSA is divided into two branches. One branch performs window self-attention calculation on query, key, and value after deep convolution processing. The other branch performs self-attention calculation with a larger window on key and value after hole convolution processing. Finally, the calculation results of the two branches are fused. In this way, local details and global context information can be captured simultaneously, and the spatial-spectral feature expression ability of hyperspectral data can be effectively enhanced. This method has stronger denoising ability while reconstructing the details and structures of hyperspectral images, and has obvious advantages in suppressing artifacts compared to other methods.
[0006] To achieve the above purpose, the present application provides the following technical solutions:
[0007] The application discloses a hyperspectral snapshot compression imaging reconstruction method based on conditional diffusion, and the method comprises the following steps: S1, acquiring a two-dimensional compressed image of a target scene; S2, preprocessing a hyperspectral image training set to generate sample data for training; S3, inputting the preprocessed training sample into a hierarchical conditional diffusion image reconstruction network for training; and S4, after the network training is completed, inputting a test sample into the network for hyperspectral image reconstruction to obtain a result.
[0008] 1. Further, in step S1, a two-dimensional compressed image is acquired by using an encoding aperture snapshot spectral imaging system, and the step specifically comprises the following steps: S11, selecting a spectral range, setting the working wavelength to 450nm-650nm, and dividing the wavelength range into 28 channels; S12, designing an encoding aperture for introducing different optical codes at each pixel; S13, setting parameters of the encoding aperture snapshot spectral imaging system, including but not limited to the number of dispersion elements, the displacement step of dispersion and the like; and S14, capturing a compressed measurement image of the entire scene by using an optical sensor.
[0009] 2. Further, in step S2, the hyperspectral image serving as the training set is preprocessed, and the step specifically comprises the following steps: S21, dividing the hyperspectral image into two parts, generating samples by using a 256x256 non-overlapping window for the first part, dividing the image by using a 128x128 non-overlapping window for the second part, and randomly combining the second part into a plurality of 256x256 samples; S22, performing data enhancement on the two parts, including random rotation and flipping; and S23, generating corresponding two-dimensional compressed measurements by simulating the imaging system, that is, modulating the hyperspectral image by using a pre-designed physical mask, and then performing shearing, shifting and adding, wherein the data-enhanced hyperspectral image and the corresponding two-dimensional compressed measurement jointly constitute a training sample pair.
[0010] 3. Further, in step S3, two processes, that is, a forward diffusion process and a reverse inference process, are included: the forward process gradually adds Gaussian noise to an original hyperspectral image X0 within T steps by using a Markov chain to generate a pure Gaussian noise image X T ~N(0,I), and the formula is as follows:
[0011]
[0012] wherein X T represents a pure noise image at the end of the forward diffusion process, I is a unit matrix, a 1:T ∈(0,1) is a scalar hyperparameter, and determines the noise variance introduced in each iteration process.
[0013] Based on the reparameterization technique, X t The distribution under the condition of X0 can be directly sampled by omitting the intermediate conversion process, thereby simplifying the calculation, and the formula is as follows:
[0014]
[0015] where, Thus, the output of each diffusion step can be obtained directly, as follows:
[0016]
[0017] Further, the reverse inference process starts from a pure Gaussian noise and optimizes the sample step by step in T iterations, as follows:
[0018]
[0019] p(X T )=N(X T ;0,I)
[0020] p θ (X t-1 |X t ,Y,M)=N(X t-1 ;μ θ (X t ,Y,M,γ t ),σ 2 I)
[0021] where the conditional distribution p θ (X t-1 |X t ,Y,M) represents the inference process of the training learning. The method uses a denoising network to predict and uses a denoising diffusion implicit model (DDIM) to accelerate sampling, and the forward diffusion process can be modeled as a non-Markov process, and the reverse process iteratively samples X t-1 , as follows:
[0022]
[0023] where ∈ θ (X t ,Y,M,t) represents a noise variable calculated from X θ (X t ,Y,M,t), as follows:
[0024]
[0025] DDIM uses a subsequence of the original time steps, i.e., {τ1,τ2,…,τ n}∈{1,...,T}, and the core idea is to jump sampling, i.e., to reduce the number of sampling steps from T to n without significantly reducing the reconstruction quality.
[0026] Further, the image after adding noise is input into the denoising network. The denoising network is composed of a U-shaped network, including three stages of down-sampling, intermediate layer and up-sampling. Among them, the down-sampling module is a 4x4 convolution operation, and the formula is as follows:
[0027]
[0028] wherein X is an input feature map, and U(X) is a feature map obtained after up-sampling, denotes a convolution operation with a filter size of N x N.
[0029] The up-sampling module is a 2x2 deconvolution operation, and the formula is as follows:
[0030]
[0031] wherein X is an input feature map, and D(X) is a feature map obtained after down-sampling, denotes a deconvolution operation with a filter size of N x N.
[0032] After two down-sampling operations and attention module processing, the intermediate layer is passed through, and then two up-sampling operations and attention module processing are performed, and then the feature map output by the U-shaped network is obtained.
[0033] Further, in the denoising network, the main components are local-global-spectral enhancement attention block (LGSAB) and its simplified version local-spectral enhancement attention block (LSAB), and the main components of the module are local-global-spectral enhancement multi-head self-attention (LGS-MSA) and local-spectral enhancement multi-head self-attention (LS-MSA). LGS-MSA and LS-MSA first convert the input feature map into query value, key value and attribute value through three linear layers, calculate the time embedding using the sine-cosine position encoding of the time step, and obtain the time token through the Swish activation function and the MLP layer. After processing by the linear layer, the final query value, key value and attribute value are obtained, and the formula is as follows:
[0034]
[0035] wherein X represents an input feature map, query, key and value represent obtained query value, key value and attribute value, W q , W k and W v represent the projection weights of the query value, the key value and the attribute value, respectively.
[0036] For LGS-MSA, the query value, the key value and the attribute value are divided into two equal parts in the channel dimension, and the result is [Q l , K l , Vl ] and [Q g ,K g ,V g ] are processed separately by local branch and global branch. The local branch first uses a 3x3 depth convolution layer to enhance local features, as follows:
[0037]
[0038] where, denotes a 3x3 depth convolution layer.
[0039] Then, are divided using non-overlapping windows (size M1xM1), and the results are Subsequently, they are divided into multiple attention heads along the channel dimension, as follows:
[0040]
[0041] where h is the number of attention heads. The self-attention of each head is calculated as follows:
[0042]
[0043] where Softmax() denotes normalization processing of attention scores, d is the dimension of each head, is a learnable parameter, and denotes additional position information.
[0044] The global branch first uses a 3x3 dilated convolution to process K g and V g , as follows:
[0045]
[0046] where, denotes a 3x3 dilated convolution layer.
[0047] Then, the query, key, and value are divided using a larger non-overlapping window (size M2xM2) than the window used by the local branch, obtaining The subsequent operations are the same as those of the local branch. Therefore, the self-attention of the global branch can be calculated as follows:
[0048]
[0049] Finally, the local and global attention calculation results are spliced, and the final attention output is obtained through linear transformation:
[0050]
[0051] where Concat[] represents channel dimension concatenation of the outputs of multiple attention heads, and Linear() represents linear transformation, i.e., a fully connected layer.
[0052] For LS-MSA, the complete query, key, and value are directly processed without channel two-splitting, and then the same processing as the local branch of LGS-MSA is adopted to calculate attention.
[0053] Further, in LGSAB and LSAB, a 3x3 deep convolution is first used to calculate a residual from a feature map, and then the residual is added to the input feature map to perform conditional position encoding, as follows:
[0054]
[0055] where X is an input feature map, and CPE(X) is a feature map after conditional position encoding.
[0056] Then, a layer normalization operation is performed on the feature map after position encoding, and then LGS-MSA or LS-MSA processing is performed to obtain a residual, which is then added to the input feature map, as follows:
[0057] LGS(X) = X + LayerNorm(LGSMSA(X))
[0058] LS(X) = X + LayerNorm(LSMSA(X))
[0059] where X is an input feature map, LGS() and LS() respectively represent a feature map after LGSAB module and LSAB module, LayerNorm() represents layer normalization, LGSMSA() represents local-global-spectral enhanced multi-head self-attention (LGS-MSA), and LSMSA() represents local-spectral enhanced multi-head self-attention (LS-MSA).
[0060] Finally, a layer normalization operation is performed again on the result after self-attention calculation, and then a feedforward neural network is performed. The feedforward neural network includes 1x1 convolution, 3x3 deep convolution, and 1x1 convolution, and a GELU activation function is used for activation between each two convolutions, as follows:
[0061]
[0062] where X is an input feature map, and F(X) represents a feature map after attention module processing.
[0063] Further, the two-dimensional compressed measurement training samples processed by the analog imaging system are initialized to generate conditions for diffusion generation; first, by sliding extraction in the channel dimension, each column in the compressed measurement is restored to a specific position of the output hyperspectral image, thereby converting the original two-dimensional data into a rough three-dimensional hyperspectral image with higher dimensions; then the physical mask of the imaging system is spliced with the hyperspectral image in the channel dimension, and then the channel number is restored to the channel number of the original hyperspectral image through 1x1 convolution to obtain the initial input of the network, as follows:
[0064]
[0065] wherein X is the input feature map, M is the physical mask, I(X) is the obtained initial input, represents a convolution operation with a filter size of NxN, and [] represents a splicing operation in the channel dimension.
[0066] Then it is passed to different branches. Each branch includes a 3x3 convolution layer followed by LSAB to generate the input condition of each layer of the denoising network (if necessary, use a 4x4 convolution layer for down-sampling). Then the generated input condition is embedded into each layer of the U-shaped network encoder, which can effectively utilize multi-scale features to generate more accurate results in the reverse reasoning process.
[0067] Further, the prediction target is set to the original data X0, and the following loss function of network training is obtained:
[0068] Loss=||X0-X θ (X t ,Y,M,t)||1
[0069] wherein X θ (X t ,Y,M,t) represents the hyperspectral image predicted by the denoising network under the given conditions (Y, M) and time embedding, and ||·||1 represents the L1 norm, i.e., the sum of absolute errors between pixels. The loss function is calculated, the model parameters of the hyperspectral image reconstruction framework are optimized according to the cross-loss function and the back propagation process, and after training, the trained hyperspectral image reconstruction framework is obtained; the input sample is reconstructed by the trained hyperspectral image reconstruction framework, and the hyperspectral image reconstruction effect diagram is output.
[0070] The beneficial effects of the present application are:
[0071] The application proposes a hyperspectral snapshot compressive imaging reconstruction method (HDiff-HIR) based on conditional diffusion, which realizes high-fidelity reconstruction by using the step-by-step iteration characteristics of the diffusion generative model, and effectively taps the potential of the deep generative model in the hyperspectral image reconstruction task. The application proposes a conditional generative module (MCGM) that fuses two-dimensional measurement information and coded aperture information, and embeds the generated conditions in different stages of the denoising network in layers, guides the network to gradually optimize the reconstruction result, fully utilizes the measurement and prior information, and improves the accuracy and stability of the reconstruction. The application proposes a new attention mechanism, local-global-spectral enhanced multi-head self-attention (LGS-MSA), which simultaneously captures local detail information, long-range dependency and fine spectral features, effectively improving the feature expression ability and reconstruction quality of the model. The experimental results on three hyperspectral data sets show that the proposed HDiff-HIR is superior to existing advanced methods in terms of reconstruction accuracy, denoising ability and artifact suppression.
[0072] Other advantages, objects, and features of the present application will be apparent to those skilled in the art from the following specification, which is to be taken in connection with the accompanying drawings illustrating the principles of the present application. The objects and other advantages of the present application will be achieved and obtained by means of the features and combinations hereinafter set forth. BRIEF DESCRIPTION OF DRAWINGS
[0073] In order to make the purposes, technical solutions and advantages of the present application clearer, the preferred detailed description of the present application will be made below in combination with the drawings, in which:
[0074] Figure 1 The flowchart of the method of the present application is shown in Figure 1.
[0075] Figure 2 The principle diagram of the coded aperture snapshot spectral imaging system (CASSI) is shown in Figure 2.
[0076] Figure 3 The schematic diagram of the hyperspectral image reconstruction process based on the diffusion model is shown in Figure 3.
[0077] Figure 4 The overall architecture diagram of the hyperspectral snapshot compressive imaging reconstruction method (HDiff-HIR) based on conditional diffusion is shown in Figure 4, wherein (a) is the overall architecture of the hierarchical conditional denoising diffusion probability model, (b) is the mask integrated conditional generative module, (c) is the local-global-spectral enhanced attention block, and (d) is the local-spectral enhanced attention block.
[0078] Figure 5 The structure diagram of the attention mechanism of the present application is shown in Figure 5, wherein (a) is the local-global-spectral enhanced multi-head self-attention, and (b) is the local-spectral enhanced multi-head self-attention.
[0079] Figure 6 Visualization results of different methods on four channels of a scene on the KAIST hyperspectral image dataset, where (a) TwIST, (b) TSA-Net, (c) MST, (d) DWMT, (e) HDiff-HIR, and (f) ground truth map;
[0080] Figure 7 Visualization results of different methods on two channels of a scene on a real CASSI image dataset, where (a) TwIST, (b) TSA-Net, (c) MST, (d) DWMT, and (e) HDiff-HIR. Detailed Implementation
[0081] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.
[0082] Figure 1 This is a flowchart of the method of the present invention. The present invention provides a hyperspectral snapshot compressed imaging reconstruction method with fused conditional diffusion. As shown in the figure, in the image acquisition stage, a coded aperture snapshot spectral imaging system is used to acquire two-dimensional compressed images. The conditional diffusion model used for hyperspectral image reconstruction is as follows: Figure 4 As shown, this method guides the diffusion process to gradually recover high-quality hyperspectral images in both spatial and spectral dimensions by introducing conditional information generated from two-dimensional measurements and encoded aperture. Specifically, the proposed model employs a Conditional Generation Module (MCGM) to effectively fuse measurement information with prior knowledge and embeds hierarchically generated conditions into different stages of the denoising network, guiding the network to gradually denoise and reconstruct image details during back-inference. The Local-Global-Spectral Enhancement Attention Block (LGSAB) in the network uses a local-global-spectral enhancement multi-head self-attention mechanism, which can simultaneously capture local detail information, long-range dependencies, and spectral features, thereby further improving reconstruction quality and stability. The overall scheme fully combines the capabilities of measurement constraints and generative models, effectively improving the accuracy, robustness, and artifact suppression of reconstruction. Specifically, the technical solution of this invention includes the following:
[0083] 1. Data Acquisition: Snapshot compressed image acquisition is achieved using a coded aperture snapshot spectral imaging system. First, the spectral range is selected, with the working wavelength set to 450nm-650nm, divided into 28 channels. Then, a coded aperture is designed to introduce different optical codes at each pixel. Next, the parameters of the coded aperture snapshot spectral imaging system are set, including but not limited to the number of dispersive elements and the dispersion shift step size. Finally, a compressed measurement image of the entire scene is captured using an optical sensor.
[0084] 2. Data preprocessing: For the hyperspectral image in the training set, first, divide the hyperspectral image into two parts, the first part uses a non-overlapping window of 256x256 to generate samples, and the second part uses a non-overlapping window of 128x128 to divide the image and randomly combine it into several 256x256 samples. Then, data augmentation is performed on the two parts, including random rotation and flipping. Then, simulate the imaging system, that is, after the hyperspectral image is modulated by the pre-designed physical mask, perform shear shift and addition to generate the corresponding two-dimensional compressed measurement. The hyperspectral image after data augmentation and its corresponding two-dimensional compressed measurement jointly constitute a training sample pair.
[0085] 3. Construct a forward diffusion process, as shown in Figure 3 . The forward process gradually adds Gaussian noise to the original hyperspectral image X0 within T steps through Markov chain to generate a pure Gaussian noise image X T ~ N(0, I), as follows:
[0086]
[0087] where X T represents the pure noise image at the end of the forward diffusion process, I is the unit matrix, a 1:T ∈(0, 1) is a scalar hyperparameter that determines the noise variance introduced in each iteration process.
[0088] Based on the reparameterization technique, the distribution of X t under the given X0 can be directly sampled by omitting the intermediate conversion process, thereby simplifying the calculation, as follows:
[0089]
[0090] where, Therefore, the output of each diffusion step can be directly obtained, as follows:
[0091]
[0092] 4. Construct a backward inference process, as shown in Figure 3 . The backward inference process starts from a pure Gaussian noise and gradually optimizes the sample in T iteration steps, as follows:
[0093]
[0094] p(X T ) = N(X T ; 0, I)
[0095] p θ (X t-1 | X tY, M) = N(X t-1 ; μ θ (X t , Y, M, γ t ), σ 2 I)
[0096] where the conditional distribution p θ (X t-1 |X t , Y, M) represents the inference process of the trained model. The method uses a denoising network to predict and uses a denoising diffusion implicit model (DDIM) to accelerate sampling, and the forward diffusion process can be modeled as a non-Markov process, and the reverse process iteratively samples X t-1 , the formula is as follows:
[0097]
[0098] where ∈ θ (X t , Y, M, t) represents a noise variable calculated from X θ (X t , Y, M, t), the formula is as follows:
[0099]
[0100] DDIM uses a subsequence of the original time steps, i.e. {τ1, τ2, …, τ n} ∈ {1, …, T}, the core idea is to jump sampling, that is, under the premise of not significantly reducing the reconstruction quality, the number of sampling steps is reduced from T to n.
[0101] 5. The noisy image is input into the denoising network. The denoising network is composed of a U-shaped network, including down-sampling, intermediate layer, and up-sampling three stages. Among them, the down-sampling module is 4x4 convolution operation, the formula is as follows:
[0102]
[0103] where X is the input feature map, U(X) is the feature map obtained after up-sampling, represents the convolution operation with filter size N x N.
[0104] The up-sampling module is a 2x2 deconvolution operation, the formula is as follows:
[0105]
[0106] where X is the input feature map, D(X) is the feature map obtained after down-sampling, represents the deconvolution operation with filter size N x N.
[0107] After two downsampling operations and attention module processing, through the intermediate layer, and then two upsampling operations and attention module processing, the feature map output by the U-shaped network is obtained.
[0108] 6. In the denoising network, the main components are local-global-spectral enhancement attention block (LGSAB) and its simplified version local-spectral enhancement attention block (LSAB), and the main components of the module are local-global-spectral enhancement multi-head self-attention (LGS-MSA) and local-spectral enhancement multi-head self-attention (LS-MSA). LGS-MSA and LS-MSA first convert the input feature map into query value, key value, and attribute value through three linear layers, calculate the time embedding using the sine-cosine position encoding of the time step, and obtain the time token through the Swish activation function and the MLP layer. After processing through the linear layer, the final query value, key value, and attribute value are obtained, and the formula is as follows:
[0109]
[0110] Where X represents the input feature map, query, key, and value represent the obtained query value, key value, and attribute value, W q , W k , and W v represent the projection weights of the query value, key value, and attribute value, respectively.
[0111] For LGS-MSA, the query value, key value, and attribute value are divided into two equal parts in the channel dimension, resulting in [Q l , K l , and V l ] and [Q g , K g , and V g ], which are divided into local branches and global branches for processing. The local branch first uses a 3x3 deep convolution layer for local feature enhancement, and the formula is as follows:
[0112]
[0113] Where, represents a 3x3 deep convolution layer.
[0114] Then, use non-overlapping windows (size M1xM1) to divide the results into Subsequently, they are divided into multiple attention heads along the channel dimension, and the formula is as follows:
[0115]
[0116] where h is the number of attention heads. The self-attention of each head is calculated as follows:
[0117]
[0118] where d is the dimension of each head, is a learnable parameter, representing additional positional information.
[0119] While the global branch first processes K g and V g using a 3x3 dilated convolution, the formula is as follows:
[0120]
[0121] where, represents a 3x3 dilated convolution layer.
[0122] Then, the query, key, and value are divided using a window size larger than the window size used by the local branch (M2xM2), obtaining The subsequent operations are the same as the local branch. Therefore, the self-attention of the global branch can be calculated as follows:
[0123]
[0124] Finally, the local and global attention calculation results are spliced, and the final attention output is obtained through linear transformation:
[0125]
[0126] where Concat[] represents channel dimension splicing of multiple attention head outputs, and Linear() represents linear transformation, i.e., a fully connected layer;
[0127] For LS-MSA, the complete query, key, and value are directly processed without channel bisection, and then the same processing method as the LGS-MSA local branch is used to calculate the attention.
[0128] 7. In LGSAB and LSAB, first, a 3x3 deep convolution is used to calculate the residual, and then it is added to the input feature map to perform conditional position encoding, the formula is as follows:
[0129]
[0130] where X is the input feature map, and CPE(X) is the feature map after conditional position encoding.
[0131] Then, the layer normalization operation is performed on the position-coded feature map, and then the LGS-MSA or LS-MSA processing is performed to obtain the residual, and then the residual is added to the input feature map, and the formula is as follows:
[0132] LGS(X) = X + LayerNorm(LGSMSA(X))
[0133] LS(X) = X + LayerNorm(LSMSA(X))
[0134] Wherein, X is the input feature map, LGS(), LS() respectively represent the feature map after LGSAB module and LSAB module, LayerNorm() represents layer normalization, LGSMSA() represents local-global-spectrum enhanced multi-head self-attention (LGS-MSA), LSMSA() represents local-spectrum enhanced multi-head self-attention (LS-MSA).
[0135] Finally, the layer normalization operation is performed again on the result after the self-attention calculation, and then a feedforward neural network is performed. The feedforward neural network includes 1x1 convolution, 3x3 deep convolution, 1x1 convolution, wherein GELU is used for activation before each two convolutions, and the formula is as follows:
[0136]
[0137] Wherein, X is the input feature map, F(X) represents the feature map processed by the attention module.
[0138] 8. The two-dimensional compressed measurement training sample processed by the simulation imaging system is initialized to generate conditions for diffusion generation; first, by sliding extraction in the channel dimension, each column in the compressed measurement is restored to a specific position of the output hyperspectral image, so that the original two-dimensional data is converted into a rough three-dimensional hyperspectral image with higher dimension; then the physical mask of the imaging system is spliced with the hyperspectral image in the channel dimension, and then the channel number is restored to the channel number of the original hyperspectral image through 1x1 convolution to obtain the initial input of the network, and the formula is as follows:
[0139]
[0140] Wherein, X is the input feature map, M is the physical mask, I(X) is the obtained initial input, Indicates the convolution operation with filter size N x N, and “[]” represents the splicing operation in the channel dimension.
[0141] Then it is passed to different branches. Each branch includes a 3x3 convolutional layer followed by LSAB to generate the input condition for each layer of the denoising network (if necessary, down-sampling is performed using a 4x4 convolutional layer). The generated input condition is then embedded into each layer of the U-shaped network encoder, which can effectively utilize multi-scale features to generate more accurate results in the reverse inference process.
[0142] 9. The prediction target is set as the original data X0, so as to obtain the loss function of the network training:
[0143] Loss=||X0-X θ (X t ,Y,M,t)||1
[0144] Wherein, X θ (X t ,Y,M,t) represents the hyperspectral image predicted by the denoising network under the given condition (Y, M) and time embedding, and ||·||1 represents the L1 norm, that is, the sum of the absolute errors between pixels. The loss function is calculated, the model parameters of the hyperspectral image reconstruction framework are optimized according to the cross loss function and the back propagation process, and after the training is completed, the trained hyperspectral image reconstruction framework is obtained; the input sample is reconstructed by the trained hyperspectral image reconstruction framework, and the hyperspectral image reconstruction effect diagram is output.
[0145] As Figure 6 is the experimental result of the hyperspectral snapshot compression imaging reconstruction method based on conditional diffusion according to the application on an open source hyperspectral natural scene. It can be seen that the details of the scene are well restored, and the reconstructed image has high quality and is highly consistent with the real scene. The reconstruction effect of the application can be further illustrated by a comparative experiment. The application method and other existing methods GAP-TV, DeSCI, TSA-Net, HDNet, MST, CST, DWMT are compared on the hyperspectral natural scene data set, and the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) are calculated, respectively. The greater the peak signal-to-noise ratio (PSNR), the smaller the distortion, and the higher the quality of the hyperspectral image. The greater the structural similarity (SSIM), the more similar the structure, brightness and contrast of the image to the real image, and the higher the image quality. Table 1 shows the average reconstruction results of ten scenes selected from the KAIST test set by different methods:
[0146] Table 1 Comparison of HDiff-HIR and various methods on KAIST hyperspectral image dataset (average value of ten scenes)
[0147]
[0148] The method of the present application achieves the best reconstruction accuracy on the data set. Figure 7 The visualization reconstruction results of different methods on a typical scene in the real CASSI data set are shown. It can be seen that the method proposed in the present application is obviously superior to other hyperspectral image reconstruction methods in terms of reconstruction quality, can more finely recover the detailed structure of the image, and also has obvious advantages in suppressing artifacts, and the reconstruction effect is more natural and real.
[0149] Finally, it should be explained that the above examples are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified without departing from the purpose and scope of the technical solutions, and all should be covered in the scope of the claims of the present application.
Claims
1. A hyperspectral snapshot compression imaging reconstruction method with fusion conditional diffusion, characterized in that: The method includes the following steps: S1: Acquire a two-dimensional compressed image of the target scene; S2: Preprocess the hyperspectral image training set to generate sample data for training; S3: Input the preprocessed training samples into the hierarchical conditional diffusion image reconstruction network for training; S4: After the network training is completed, the test samples are input into the network to reconstruct the hyperspectral image and obtain the results; In step S1, a two-dimensional compressed image is acquired using a coded aperture snapshot spectral imaging system, specifically including the following steps: S11, selecting the spectral range and setting the working wavelength to 450nm-650nm, dividing it into 28 channels; S12, designing the coded aperture to introduce different optical codes at each pixel; S13, setting the parameters of the coded aperture snapshot spectral imaging system, including but not limited to the number of dispersive elements and the dispersive shift step size; S14, capturing a compressed measurement image of the entire scene using an optical sensor. In step S2, the hyperspectral images used as the training set are preprocessed, specifically including the following steps: S21, the hyperspectral images are divided into two parts. The first part is divided into samples using a 256×256 non-overlapping window, and the second part is divided into images using a 128×128 non-overlapping window, and randomly combined into several 256×256 samples; S22, data augmentation is performed on the two parts respectively, including random rotation and flipping; S23, through a simulated imaging system, the hyperspectral images are modulated by a pre-designed physical mask, then sheared, shifted, and summed to generate corresponding two-dimensional compressed measurements; the data-augmented hyperspectral images and their corresponding two-dimensional compressed measurements together constitute training sample pairs; Step S3 includes two processes: a forward diffusion process and a backward reasoning process. The forward process involves gradually adding Gaussian noise to the original hyperspectral image X0 within T steps using a Markov chain to generate a pure Gaussian noise image X. T ~N(0,I), the formula is as follows: Among them, X T This represents the pure noise image at the end of the forward diffusion process, where I is the identity matrix and α... 1:T ∈(0,1) is a scalar hyperparameter that determines the noise variance introduced in each iteration; Based on reparameterization technology, X t Given X0, the distribution can be directly sampled by omitting the intermediate transformation process, thus simplifying the calculation. The formula is as follows: in, Therefore, the output of each diffusion step can be obtained directly, as shown in the following formula: The reverse inference process starts with pure Gaussian noise and optimizes the sample step by step over T iterations, as shown in the following formula: p(X T )=N(X T ;0,I) p θ (X t-1 |X t ,Y,M)=N(X t-1 ;m θ (X t ,Y,M,c t ),s 2 I) Wherein, the conditional distribution p θ (X t-1 |X t (Y, M) represents the inference process during training; this method uses a denoising network for prediction. The Denoising Diffusion Implicit Model (DDIM) is employed to accelerate sampling. Its forward diffusion process can be modeled as a non-Markovian process, and the reverse process iteratively samples X. t-1 The formula is as follows: Where, ∈ θ (X t (,Y,M,t) represents from X θ (X t The noise variable calculated from (Y, M, t) is given by the following formula: DDIM uses a subsequence of the original time step, i.e., {τ1, τ2, ..., τ...} n The core idea of the method is skip sampling, which reduces the number of sampling steps from T to n without significantly reducing the reconstruction quality.
2. The hyperspectral snapshot compression imaging reconstruction method according to claim 1, characterized in that: The noisy image is input into the denoising network; the denoising network consists of a U-shaped network, including three stages: downsampling, intermediate layers, and upsampling; the downsampling module uses a 4×4 convolution operation, as shown in the following formula: Where X is the input feature map, and U(X) is the feature map obtained after upsampling. This represents a convolution operation with a filter size of N×N; The upsampling module uses a 2×2 deconvolution operation, as shown in the following formula: Where X is the input feature map, and D(X) is the feature map obtained after downsampling. This represents a deconvolution operation with a filter size of N×N; After two downsampling operations and attention module processing, the feature map output by the U-shaped network is obtained after passing through the intermediate layer and then undergoing two more upsampling operations and attention module processing.
3. The hyperspectral snapshot compression imaging reconstruction method according to claim 2, characterized in that: In the denoising network, the main components are Local-Global-Spectral Enhanced Attention Block (LGSAB) and its simplified version, Local-Spectral Enhanced Attention Block (LSAB). The main components of the module are Local-Global-Spectral Enhanced Multi-Head Self-Attention (LGS-MSA) and Local-Spectral Enhanced Multi-Head Self-Attention (LS-MSA). LGS-MSA and LS-MSA first transform the input feature map into query value, key value, and attribute value through three linear layers, respectively. They then calculate the temporal embedding using sine-cosine positional encoding at each time step, and obtain the temporal token through the Swish activation function and MLP layer. This token is also processed through linear layers, and the final query value, key value, and attribute value are obtained by summing them, as shown in the following formula: Where X represents the input feature map, query, key, and value represent the obtained query value, key value, and attribute value, respectively, and W... q W k W v These represent the projected weights of the query value, key value, and attribute value, respectively. For LGS-MSA, the query value, key value, and attribute value are divided into two equal parts along the channel dimension, resulting in [Q]. l ,K l V l ] and [Q g ,K g V g The process is divided into local and global branches, which are processed separately. The local branch first uses a 3×3 depth convolutional layer for local feature enhancement, as shown in the following formula: in, This represents a 3×3 depth convolutional layer; Then, The partitioning is performed using non-overlapping windows (M1×M1), resulting in: Subsequently, they are divided into multiple attention heads along the channel dimension, as shown in the following formula: Where h is the number of attention heads, and the self-attention of each head is calculated as follows: Here, Softmax() represents the normalization process for the attention score, and d is the dimension of each head. These are learnable parameters that represent additional location information; The global branch first uses a 3×3 dilated convolution to process K. g and V g The formula is as follows: in, This represents a 3×3 hollow convolutional layer; Then, the query, key, and value are partitioned using a larger non-overlapping window (M2×M2) compared to the window size used in the local branch, resulting in... The subsequent operations are the same as for the local branch; therefore, the self-attention of the global branch can be calculated as follows: Finally, the local and global attention calculation results are concatenated and a linear transformation is applied to obtain the final attention output: Where Concat[] represents concatenating the channel dimensions of the outputs of multiple attention heads, and Linear() represents a linear transformation, i.e., a fully connected layer; For LS-MSA, the complete query, key, and value are processed directly without bisecting the channel, and then the same processing method as the local branch of LGS-MSA is used to calculate the attention.
4. The hyperspectral snapshot compression imaging reconstruction method according to claim 3, characterized in that: In LGSAB and LSAB, the feature map is first subjected to a 3×3 depthwise convolution to calculate the residual, which is then added to the input feature map to perform conditional location encoding, as shown in the following formula: Where X is the input feature map, and CPE(X) is the feature map after conditional location encoding; Then, layer normalization is applied to the position-encoded feature map, followed by LGS-MSA or LS-MSA processing to obtain the residual, which is then added to the input feature map, as shown in the following formula: LGS(X)=X+LayerNorm(LGSMSA(X)) LS(X) = X + LayerNorm(LSMSA(X)) Where X is the input feature map, LGS() and LS() represent the feature maps after passing through the LGSAB module and LSAB module, respectively, LayerNorm() represents layer normalization, LGSMA() represents Local-Global-Spectral Enhanced Multi-Head Self-Attention (LGS-MSA), and LSMSA() represents Local-Spectral Enhanced Multi-Head Self-Attention (LS-MSA). Finally, the result after self-attention calculation is normalized again, and then passed through a feedforward neural network. The feedforward neural network contains 1×1 convolutions, 3×3 depthwise convolutions, and 1×1 convolutions. Before each two convolutions, the GELU activation function is used for activation, as shown in the following formula: Where X is the input feature map, and F(X) represents the feature map after processing by the attention module.
5. The hyperspectral snapshot compression imaging reconstruction method according to claim 4, characterized in that: The training samples from the two-dimensional compressed measurements processed by the simulated imaging system are initialized to generate conditions for diffusion generation. First, by performing sliding extraction along the channel dimension, each column in the compressed measurement is restored to a specific position in the output hyperspectral image, thereby transforming the original two-dimensional data into a coarse three-dimensional hyperspectral image with higher dimensions. Then, the physical mask of the imaging system is stitched with this hyperspectral image along the channel dimension, and then a 1×1 convolution is performed to restore the number of channels to the number of channels in the original hyperspectral image, obtaining the initial input to the network, as shown in the following formula: Where X is the input feature map, M is the physical mask, and I(X) is the obtained initial input. This indicates a convolution operation with a filter size of N×N, and "[]" indicates a concatenation operation along the channel dimension; This is then passed to different branches; each branch includes a 3×3 convolutional layer followed by LSAB to generate input conditions for each layer of the denoising network (using 4×4 convolutional layers for downsampling if necessary); the generated input conditions are then embedded into each layer of the U-shaped network encoder, which can effectively utilize multi-scale features to generate more accurate results during the reverse inference process.
6. The hyperspectral snapshot compression imaging reconstruction method according to claim 5, characterized in that: Setting the prediction target to the original data X0, we obtain the following loss function for network training: Loss=||X0-X θ (X t ,Y,M,t)||1 Among them, X θ (X t Let (Y, M, t) represent the hyperspectral image predicted by the denoising network under given conditions (Y, M) and temporal embedding, and ||·||1 represent the L1 norm, i.e., the sum of absolute errors between pixels. The loss function is calculated, and the model parameters of the hyperspectral image reconstruction framework are optimized based on the cross-loss function and the backpropagation process. After training, a trained hyperspectral image reconstruction framework is obtained. The input samples are reconstructed using the trained hyperspectral image reconstruction framework, and the resulting hyperspectral image reconstruction image is output.
Citation Information
Patent Citations
Hyperspectral reconstruction method and system based on CASSI optical system
CN118052897A
Method for accelerating image inpainting based on diffusion model
CN118608427A
Medical image unified diffusion generation method oriented to random mode deficiency and related device
CN120107393A
Cited By
Multi-spectral imaging satellite data space-time fusion method, medium, equipment and product
CN121259643A
Hyperspectral fusion imaging method and system of double-branch diffusion model framework
CN121707841A