Hyperspectral image reconstruction method and device based on high compression ratio snapshot compression and readable storage medium thereof
Through the lightweight space-spectral state space model, the GHSB, BPSB and GFFN modules are used to solve the high accuracy and low complexity problems of hyperspectral image reconstruction under high compression ratio, and efficient space-spectral information integration and image quality improvement are achieved.
Patent Information
- Application Number
- CN202510483102.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-22
AI Technical Summary
The existing hyperspectral image reconstruction algorithm is difficult to take into account high precision and lightweight in high compression ratio scenarios, and has high computational complexity, so it is impossible to effectively integrate space-spectral information.
Using a lightweight space-spectral state space model, long-range space dependence is captured through multi-directional scanning of the grouped hybrid space module (GHSB), combined with the bidirectional patch spectral module (BPSB), and optimized feature representation using a gated feedforward network (GFFN).
Realizing high-precision and low-complexity hyperspectral image reconstruction under high compression ratio improves reconstruction quality, reduces computing resource requirements, and is suitable for resource-constrained scenarios.
Smart Images

Figure CN120355808A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hyperspectral image processing and computational imaging, and particularly to a hyperspectral image reconstruction method, apparatus, and readable storage medium based on high compression ratio snapshot compression. Background Art
[0002] Hyperspectral imaging technology captures information in hundreds of narrow spectral bands to form a three-dimensional data cube containing rich spatial-spectral information, which has important application value in fields such as environmental monitoring, biomedicine, and geological exploration. Traditional hyperspectral cameras use a point-by-point or line-by-line scanning method, resulting in long imaging times and difficulty in capturing dynamic scenes. Therefore, the coded aperture snapshot spectral imaging system (CASSI) emerged. CASSI achieves compressed measurements of multiple spectral bands through a single exposure, modulates spectral information onto a two-dimensional detector using a physical mask and a dispersive element, significantly improving the imaging speed and reducing the data transmission volume. However, recovering the original hyperspectral image from the compressed measurements relies on an efficient reconstruction algorithm.
[0003] Existing reconstruction algorithms are mainly based on convolutional neural networks (CNNs) or Transformer architectures: CNNs are good at capturing local spatial features but difficult to model long-range spatial dependencies; although Transformer can handle global dependencies, its computational complexity grows quadratically with the input scale, resulting in a large number of model parameters and low computational efficiency. In addition, existing research mostly focuses on low compression ratio (such as 28 bands) scenarios, and there is a lack of lightweight models that can efficiently integrate spatial-spectral information for hyperspectral image reconstruction at high compression ratios (such as 83 bands), resulting in insufficient reconstruction accuracy and high consumption of computational resources.
[0004] Therefore, there is an urgent need for an end-to-end reconstruction method based on a lightweight spatial-spectral aggregation block (SSAB) to solve the problems existing in the prior art. Summary of the Invention
[0005] Embodiments of the present invention provide a hyperspectral image reconstruction method, apparatus, and readable storage medium based on high compression ratio snapshot compression, aiming at the problems existing in the current technology that there is a contradiction between the CASSI reconstruction algorithm in long-range spatial dependency modeling and computational efficiency, and it is difficult to balance the reconstruction accuracy and lightweight requirements at high compression ratios (such as 83 bands).
[0006] The core technology of the present invention mainly proposes a lightweight spatio-spectral state space model, which captures long-range spatial dependencies through multi-directional scanning of the grouped hybrid spatial module (GHSB), and combines the bidirectional patch spectral module (BPSB) to efficiently aggregate spectral context, achieving high-precision and low-complexity reconstruction at a high compression ratio (83 bands).
[0007] In a first aspect, the present invention provides a hyperspectral image reconstruction method based on high-compression ratio snapshot compression, and the method includes the following steps: S1. Obtain two-dimensional compressed measurement values collected by a CASSI system; S2. Perform a dispersion inverse operation on the two-dimensional compressed measurement values to generate an initial three-dimensional hyperspectral signal; S3. Concatenate the initial hyperspectral signal with a physical mask, and generate an initial feature map through a convolution operation; S4. Input the initial feature map into a reconstruction model based on a U-shaped network, and perform multi-level feature extraction and reconstruction through an encoder-bottleneck-decoder structure, where the core unit of the reconstruction model is a spatio-spectral aggregation block; The spatio-spectral aggregation block sequentially includes: Grouped hybrid spatial module: Group the input features along the channel dimension, and capture long-range spatial dependencies through multi-directional scanning and inter-group residual connections; Bidirectional patch spectral module: Divide the features into small blocks and concatenate them along the spectral dimension, and capture spectral context information through bidirectional scanning; Gated feed-forward network: Integrate spatial and spectral features through a gating mechanism to optimize the output representation; S5. Add the residual features output by the decoder to the initial input to generate a reconstructed hyperspectral image.
[0008] Further, in step S4, the implementation of the encoder-bottleneck-decoder structure includes: The encoder gradually compresses the spatial dimension and expands the channels through N1 consecutive spatio-spectral aggregation blocks and downsampling operations; The bottleneck part extracts high-level feature representations through N3 spatio-spectral aggregation blocks; The decoder restores the spatial details through upsampling and skip connections, and the skip connections are concatenated with the features in the corresponding stage of the encoder.
[0009] Further, in step S4, the specific operation of the grouped hybrid spatial module is: Divide the input data into G groups along the channel dimension, and each group of data is subjected to multi-directional scanning through a Spatial-SSM module, and the scanning directions include at least two of left to right, right to left, up to down, and down to up; Except for the first group, the input of each group fuses the output of the Spatial-SSM module of the previous group to form an inter-group residual connection; Concatenate the outputs of all groups and integrate the information through 1×1 convolution.
[0010] Furthermore, in step S4, the Spatial-SSM module unfolds the input features into a one-dimensional sequence and models the spatial pixel correlation through the state transition equation. The state transition equation is:
[0011] where h t is the hidden state at the current spatial position, used to store and transmit spatial context information; h t-1 is the hidden state at the previous spatial position, reflecting the dependency between spatial positions; x t is the input feature at the current spatial position; is calculated by the learnable parameters ΔA, ΔB through and is used to adjust the contribution weights of the hidden state h t-1 and the input x t to the current hidden state h t , introducing a non-linear transformation to capture complex spatial relationships; C, D are learnable parameters used to map the hidden state h t and the input x t to the output y t .
[0012] Furthermore, in step S4, the processing steps of the bidirectional patch spectral module include: Divide the input features into blocks of size M×M, concatenate the flattened pixels of each block along the spectral dimension to form a sequence representation in the spectral dimension; perform a bidirectional scan on the spectral sequence to capture the context dependencies between spectral bands through the Spectral-SSM module. The bidirectional scan includes a forward scan from low frequency to high frequency and a backward scan from high frequency to low frequency.
[0013] Furthermore, in step S4, the processing steps of the gated feed-forward network include: Expand the channel dimension through 1×1 convolution and divide the features into two groups along the channels; The first group captures spatial context through 3×3 convolution and introduces non-linearity through the GELU function, and the second group is multiplied by the features of the first group through a gating mechanism; Integrate the two groups of features through 1×1 convolution to output the optimized spatial-spectral fusion features.
[0014] Furthermore, for the hyperspectral image corresponding to the high compression ratio scenario, the number of spectral bands ≥ 83 and the compression ratio ≥ 30:1.
[0015] In a second aspect, the present invention provides a hyperspectral image reconstruction apparatus based on high compression ratio snapshot compression, comprising: A dispersion inverse operation module that performs a dispersion inverse operation on the two-dimensional compressed measurement values collected by the CASSI system to generate an initial three-dimensional hyperspectral signal; A mask stitching module that stitches the initial hyperspectral signal with a physical mask and generates an initial feature map through a convolution operation; A reconstruction module that inputs the initial feature map into a reconstruction model based on a U-shaped network, performs multi-level feature extraction and reconstruction through an encoder-bottleneck-decoder structure, wherein the core unit of the reconstruction model is a spatial-spectral aggregation block; An output module that adds the residual features output by the decoder to the initial input to generate a reconstructed hyperspectral image; The spatial-spectral aggregation block sequentially includes: A grouped hybrid spatial module: groups the input features along the channel dimension, and captures long-range spatial dependencies through multi-directional scanning and inter-group residual connections; A bidirectional patch spectral module: divides the features into small blocks and stitches them along the spectral dimension, and captures spectral context information through bidirectional scanning; A gated feed-forward network: integrates spatial and spectral features through a gating mechanism to optimize the output representation.
[0016] In a third aspect, the present invention provides an electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the above-mentioned hyperspectral image reconstruction method based on high compression ratio snapshot compression.
[0017] In a fourth aspect, the present invention provides a readable storage medium, in which a computer program is stored, and the computer program includes program codes for controlling a process to execute the process, and the process includes the hyperspectral image reconstruction method based on high compression ratio snapshot compression as described above.
[0018] The main contributions and innovations of the present invention are as follows: 1. Reduce the computational complexity and achieve lightweight: Through the grouping process of the grouped hybrid spatial module (GHSB) and the block segmentation mechanism of the bidirectional patch spectral module (BPSB), the problem that the computational complexity of the traditional Transformer increases quadratically with the input scale (O(N 2 )) is avoided. For example, GHSB reduces the computational complexity of the spatial dimension from O(C 2 HW) to O((C / G) 2 HW) (G is the number of groups), significantly reducing the computational amount and parameter scale in high compression ratio scenarios, improving the model efficiency, and meeting the application requirements of resource-constrained scenarios.
[0019] 2. Enhance the ability to model long-range dependencies: GHSB captures long-range spatial dependencies through multi-directional scanning (such as left→right, up→down, etc.), breaking through the limitation of CNN which is only good at local features; BPSB captures context dependencies between spectral bands through bidirectional scanning (forward and backward), which is more comprehensive than unidirectional scanning. The combination of the two achieves efficient modeling of long-range dependencies in both spatial and spectral dimensions, superior to the existing methods of single local modeling by CNN or high-complexity global modeling by Transformer, and significantly improving the reconstruction accuracy.
[0020] 3. Excellent adaptability to high compression ratio scenarios: Existing technologies mostly focus on low compression ratio (such as 28 bands) scenarios, and the present invention is optimized for high compression ratio (such as the number of bands ≥ 83). Through the collaboration of GHSB, BPSB and gated feed-forward network (GFFN), it can still efficiently integrate spatial-spectral information under high compression ratio, avoiding information loss and degradation of reconstruction quality caused by too high compression ratio, and filling the gap of high-precision reconstruction in high compression ratio scenarios.
[0021] 4. Spatial-spectral joint optimization to improve image quality: The modeling of spatial dependencies by GHSB can correct spatial artifacts (such as edge blurring) in the reconstructed image, and the capture of spectral context by BPSB can suppress spectral outliers (such as noisy bands). The combination of the two realizes regularization in both spatial and spectral dimensions. Compared with the existing situation where spatial and spectral modeling are separated or inefficient, the present invention makes the reconstructed image more accurate in terms of spatial structure and spectral features, comprehensively improving the reconstruction quality of hyperspectral images.
[0022] Details of one or more embodiments of the present invention are set forth in the following drawings and description to make other features, objects, and advantages of the present invention more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is the flow of a hyperspectral image reconstruction method based on high compression ratio snapshot compression according to an embodiment of the present invention; Figure 2 is the structural diagram of a spatial-spectral aggregation block (SSAB) according to an embodiment of the present invention; Figure 3 is the structural diagram of a gated feed-forward network (GFFN) according to an embodiment of the present invention; Figure 4 is the schematic hardware structure diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0025] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0026] The current CASSI reconstruction algorithm faces a triangular contradiction of "computational complexity - long-range modeling - high compression ratio support": neither CNN nor Transformer can balance efficient computation and high-precision reconstruction; in the scenario of high compression ratio (such as 83 bands), the reconstruction quality of existing algorithms drops significantly, resulting in spatial blurring or spectral distortion.
[0027] Based on this, the present invention is based on a lightweight spatio-spectral state space model to solve the problems existing in the prior art.
[0028] Embodiment 1 The present invention aims to propose a hyperspectral image reconstruction method based on high compression ratio snapshot compression. Specifically, referring to Figure 1 , the method includes the following steps: S1. Obtain two-dimensional compressed measurement values collected by the CASSI system; In this embodiment, the CASSI system uses a physical mask to modulate hyperspectral signals of multiple wavelengths and projects them onto a detection plane through a dispersive element to output two-dimensional compressed measurement values.
[0029] S2. Perform an inverse dispersion operation (such as a mathematical inverse transform based on a known dispersion model) on the two-dimensional compressed measurement values to generate an initial three-dimensional hyperspectral signal; In this embodiment, the CASSI system disperses optical signals of different wavelengths along the spatial dimension (such as the row / column direction of the detector) through a dispersion element (such as a prism or a grating), so that the signal of each pixel contains the mixed information of multiple spectral bands. For example, the light with a wavelength of λ1 will be offset by d pixels on the detector after dispersion, forming spatial aliasing with the optical signal with a wavelength of λ2. This operation compresses the three-dimensional hyperspectral data (spatial dimension H×W + spectral dimension C) into two-dimensional measurements (H×W), realizing the hybrid modulation of "spectral-spatial" information. The inverse dispersion operation is the reverse recovery of the dispersion process of the CASSI system, that is, according to the known dispersion parameters (such as the mapping relationship between wavelength and spatial offset), the two-dimensional mixed measurement is restored to the initial three-dimensional hyperspectral signal (H×W×C). The specific steps include: Demix the mixed signal of each pixel point to separate the contributions of each band; According to the spatial offset corresponding to the wavelength, remap the separated signal to the corresponding position in the three-dimensional data cube.
[0030] The purpose is to restore the three-dimensional data structure (spatial dimension + spectral dimension) of the hyperspectral image, providing a basic input for subsequent feature extraction and spatial-spectral joint modeling.
[0031] S3. Stitch the initial hyperspectral signal with the physical mask and generate the initial feature map through a convolution operation; In this embodiment, the initial signal is stitched with the physical mask, and channel fusion is performed through a 1×1 convolution, and then an initial feature map is generated using a 3×3 convolution. For example, the hyperspectral image H undergoes a "shift" operation (inverse dispersion), is stitched with the mask M in the channel (⊕), and the channels are fused through a 1×1 convolution (Conv1x1) to generate an H×W×C feature map.
[0032] S4. Input the initial feature map into a reconstruction model based on a U-shaped network, and perform multi-level feature extraction and reconstruction through an encoder-bottleneck-decoder structure, where the core unit of the reconstruction model is the spatial-spectral aggregation block (SSAB); The spatial-spectral aggregation block includes in sequence: Grouped hybrid spatial module (GHSB): Group the input features along the channel dimension, and capture long-range spatial dependencies through multi-directional scanning and inter-group residual connections; Bidirectional patch spectral module (BPSB): Divide the features into small blocks and stitch them along the spectral dimension, and capture spectral context information through bidirectional scanning; Gated feed-forward network (GFFN): Integrate spatial and spectral features through a gating mechanism to optimize the output representation; In this embodiment, the Encoder starts with N1 consecutive SSAB blocks, followed by a 4×4 convolution (stride = 2) for spatial downsampling and channel doubling (as ). This sequence is then repeated through N2 SSAB blocks and another identical downsampling operation. In the Bottleneck, the compressed features go through N3 SSAB transforms to capture high-level semantic features. Corresponding to the encoder structure, the Decoder uses a transposed convolution with a stride of 2×2 for upsampling instead of the downsampling operation, and connects the corresponding encoder and decoder stages through skip connections (via 1×1 convolution) to preserve spatial details. The final reconstruction result is obtained by applying a 3×3 convolution to the decoder output to estimate the residual component, i.e., finally generating the reconstructed image Z through "Mapping" and residual connection (+).
[0033] In this embodiment, the structure of the Spatial-Spectral Aggregation Block (SSAB) is as Figure 2 shown, which integrates a Grouped Hybrid Spatial Module (GHSB), a Bidirectional Patch Spectral Module (BPSB), a Gated Feed-Forward Network (GFFN), and multiple layer normalization modules. The feature transfer between modules is optimized through layer normalization and residual connection to achieve joint modeling of spatial-spectral dependencies.
[0034] Among them, the GHSB focuses on modeling spatial information. To reduce the computational complexity, the present invention groups the input data along the channel dimension, processes each group of data separately, and then combines the outputs to obtain the final result. To compensate for the reduced inter-group information exchange caused by grouping, the present invention introduces an inter-group residual connection to enhance the interaction of inter-group information. As Figure 4 shown, the input data is divided into four groups along the channel dimension. Each group of data is processed by the Spatial-SSM, which unfolds the input into a one-dimensional sequence and scans it in different directions (from left to right, from right to left, from top to bottom, from bottom to top) to capture spatial pixel correlations. It should be noted that, except for the first group, the output of each group is connected to the input of the next group to promote the flow of information between groups. The outputs of all Spatial-SSMs are then connected and integrated through a 1×1 convolution to obtain the output.
[0035] For example, the input processing of the Grouped Hybrid Spatial Module (GHSB): The input features first go through a 3×3 depthwise separable convolution (DWConv3x3) and a 1×1 convolution (Conv1x1), and then are divided into 4 groups (X1 to X4) along the channel dimension, with the number of channels in each group being C / 4.
[0036] Multi-directional scanning: Each group of features passes through the Spatial-SSM module in different directions (e.g., Direction1 is 1→2→3→4, Direction2 is 4→3→2→1, etc.). Figure 4 The right block diagram in
[0037] shows the calculation logic of Spatial-SSM. By the state transition equation: t where h t-1 is the hidden state at the current spatial position, used to store and transmit spatial context information; h t-1 is the hidden state at the previous spatial position, reflecting the dependence relationship between spatial positions; x t is the input feature at the current spatial position; is calculated by the learnable parameters ΔA, ΔB through and is used to adjust the contribution weights of the hidden state h t-1 and the input x t to the current hidden state h t . Nonlinear transformation is introduced to capture complex spatial relationships; C, D are learnable parameters used to map the hidden state h t and the input x t to the output y t .
[0038] Output integration: The processed 4 groups of features (Y1 to Y4) are integrated through 1×1 convolution to output the final features.
[0039] Among them, BPSB aims to enhance the model's ability to learn from spectral features, while reducing the dependence on spatial information and lowering the computational complexity. The present invention divides the features into smaller blocks and connects these blocks along the spectral dimension. This enables the present invention to analyze the responses of each pixel in different bands in more detail and pay more attention to spectral features. In addition, bidirectional scanning can thoroughly capture the context relationship between spectral bands. In this way, the present invention reduces the complexity introduced by the spatial structure, enabling the model to pay more attention to interpreting spectral information rather than the spatial layout, thereby improving the computational efficiency and the accuracy of spectral feature extraction. Specifically, for the input, the present invention first applies a 1×1 convolution, followed by a 3×3 depth convolution to extract local context. Then, the present invention rearranges it into blocks and connects these blocks along the spectral dimension, and the flattened block pixels are used as channel representations. To fully capture the context information between spectral bands, the present invention uses Spectral-SSM for bidirectional scanning, overcoming the limitations of unidirectional scanning in capturing spectral dependence relationships.
[0040] For example, the feature partitioning and expansion of the bidirectional patch spectral module (BPSB): The input features first go through a 3×3 depthwise separable convolution and a 1×1 convolution, then the features are partitioned into P×P blocks, stitched together along the spectrum, and flattened to obtain a sequence representation of
[0041] Bidirectional scanning: Forward (ForwardRoute) and backward (BackwardRoute) scans are performed through Spectral-SSM to capture the context dependencies between spectral bands. The calculation method is similar to that of Spatial-SSM ( ).
[0042] Feature processing: After scanning, a 1×1 convolution is used to complete the spectral-spatial feature integration.
[0043] Among them, GFFN further optimizes the integrated spatial-spectral features to enhance the overall feature representation. As Figure 3 shown, the gated feed-forward network (GFFN) of the present invention first expands the channel dimension through a 1×1 convolution to enhance cross-band interaction. The feature map is then divided into multiple groups along the channel dimension. One of the groups captures the spatial context through a 3×3 convolution and introduces non-linearity through the GELU function \cite{gelu}. Another group is multiplied with it through a gating mechanism to further optimize the information flow. Finally, these features are integrated through a 1×1 convolution to generate an improved output that fuses spatial and spectral information.
[0044] For example, the spatial context capture of the gated feed-forward network (GFFN): After the input goes through a 1×1 convolution, non-linearity is introduced through a 3×3 depthwise separable convolution (DWConv3x3) and the GELU activation function to capture the spatial context. Gating fusion: The features are screened and fused through a gating mechanism (×), and finally, an optimized spatial-spectral fusion feature is output through a 1×1 convolution.
[0045] These three modules model the hyperspectral features from the perspectives of space, spectrum, and their fusion respectively, and together constitute the core components of the hyperspectral image reconstruction model, realizing the efficient capture and feature optimization of the spatial-spectral dependencies of hyperspectral images.
[0046] S5. Add the residual features output by the decoder to the initial input to generate the reconstructed hyperspectral image.
[0047] In this embodiment, these residual components are combined with the initial input through element-wise addition to generate the reconstructed hyperspectral image.
[0048] Embodiment 2 Based on the same concept, the present invention also proposes a hyperspectral image reconstruction device based on high compression ratio snapshot compression, including: The dispersion inverse operation module performs a dispersion inverse operation on the two-dimensional compressed measurement values collected by the CASSI system to generate an initial three-dimensional hyperspectral signal; The mask stitching module stitches the initial hyperspectral signal with a physical mask and generates an initial feature map through a convolution operation; The reconstruction module inputs the initial feature map into a reconstruction model based on a U-shaped network, performs multi-level feature extraction and reconstruction through an encoder-bottleneck-decoder structure, where the core unit of the reconstruction model is a spatio-spectral aggregation block; The spatio-spectral aggregation block sequentially includes: The grouped hybrid spatial module: groups the input features along the channel dimension, and captures long-range spatial dependencies through multi-directional scanning and inter-group residual connections; The bidirectional patch spectral module: divides the features into small blocks and stitches them along the spectral dimension, and captures spectral context information through bidirectional scanning; The gated feed-forward network: integrates spatial and spectral features through a gating mechanism to optimize the output representation; The output module adds the residual features output by the decoder to the initial input to generate a reconstructed hyperspectral image.
[0049] The following are the explanations of the technical terms appearing in the present invention, which are accurately defined in combination with the content of the technical solution: I. Core technical concepts 1. Hyperspectral Image A three-dimensional data cube composed of hundreds of narrow spectral bands (spatial dimensions H×W×C, where H and W respectively represent the sizes of the spatial dimensions, and C represents the number of channels in the spectral dimension). Each pixel contains continuous spectral information and can be used for material composition analysis (such as environmental monitoring, biomedicine). Due to the large amount of hyperspectral image data, traditional transmission and processing methods often face efficiency bottlenecks. The emergence of compression reconstruction technology provides the possibility for efficient processing of hyperspectral images. By reducing the data volume through compression technology and then using reconstruction algorithms to restore the original information, not only the processing efficiency is improved, but also the storage and transmission costs are reduced, making the efficient utilization of hyperspectral images a reality.
[0050] 2. Coded Aperture Snapshot Spectral Imaging (CASSI) An imaging system that captures multi-spectral band information through a single exposure. It uses a physical mask (coded aperture) and a dispersion element (such as a prism) to modulate optical signals of different wavelengths onto a two-dimensional detector to generate compressed measurement values (two-dimensional mixed signals), solving the problems of slow scanning speed and only being able to scan static objects in traditional hyperspectral cameras.
[0051] 3. Spatial-Spectral Aggregation Block (SSAB) The core building unit of the present invention integrates a Grouped Hybrid Spatial Block (GHSB), a Bi-directional Patch Spectral Block (BPSB), and a Gated Feed-Forward Network (GFFN). Through layer normalization and residual connections, it jointly models the spatial dependencies and spectral context relationships of hyperspectral images to achieve lightweight and efficient reconstruction.
[0052] II. Core Modules and Algorithms 4. Grouped Hybrid Spatial Block (GHSB) Function: Model long-range spatial dependencies and reduce computational complexity.
[0053] Mechanism: The input is grouped along the channel dimension (e.g., into 4 groups). Each group captures the spatial pixel correlations through a Spatial-SSM module with multi-directional scans (left → right, right → left, etc.). Information interaction between groups is enhanced through residual connections to avoid information isolation caused by grouping.
[0054] 5. Bi-directional Patch Spectral Block (BPSB) Function: Capture the context dependencies between spectral bands and focus on spectral feature analysis.
[0055] Mechanism: The features are divided into M×M small blocks. The pixels are concatenated and flattened along the spectral dimension, and the spectral correlations are modeled through a Bi-directional Spectral-SSM module (forward + backward) to reduce the interference of the spatial structure on spectral analysis.
[0056] 6. Gated Feed-Forward Network (GFFN) Function: Optimize the spatial-spectral fusion features and enhance the cross-band non-linear interaction.
[0057] Mechanism: The channels are expanded through 1×1 convolutions. After grouped processing, one group captures the spatial context (3×3 convolution + GELU), and the other group is screened and fused through a gating mechanism, and finally the optimized features are output.
[0058] 7. State Space Model (SSM) Spatial-SSM: Used for the spatial dimension, the features are unfolded into a one-dimensional sequence, and the spatial pixel dependencies are recursively modeled through the state transition equation Generated by exponentiating learnable parameters to capture long-range spatial correlations. Spectral-SSM: For the spectral dimension, similar to Spatial-SSM in principle, it models the dependencies between spectral bands through bidirectional scanning.
[0059] III. Key Operations and Technologies 8. Dispersion and Reverse Dispersion Dispersion: In the CASSI system, the dispersion element disperses optical signals of different wavelengths along the spatial dimension (e.g., the light of wavelength λ is offset by d pixels), enabling the two-dimensional detector to record the mixed spectral signal.
[0060] Reverse Dispersion: The reverse process, which restores the two-dimensional mixed measurements to the three-dimensional hyperspectral signal (H×W×C) according to the dispersion parameters, de-aliases and restores the spatial positions of each band, providing the initial input for reconstruction.
[0061] 9. DownSample and UpSample DownSample: In the encoder, the spatial resolution is reduced (e.g., H×W→H / 2×W / 2) through a 4×4 convolution (stride = 2), while doubling the number of channels to extract high-level features.
[0062] UpSample: In the decoder, the spatial resolution is restored through a 2×2 transposed convolution, and the skip connections are combined to retain details.
[0063] 10. Skip Connection and Residual Connection Skip Connection: Feature maps are directly transmitted (after 1×1 convolution) between the corresponding stages of the encoder and the decoder to avoid information loss in the deep network and retain spatial details.
[0064] Residual Connection: Inside the module or between groups (such as GHSB), the input and output are fused through element-wise addition to enhance information flow and alleviate the vanishing gradient problem.
[0065] IV. Performance-Related Terms 11. High Compression Ratio Refers to the dimensional compression ratio of the two-dimensional measurements output by the CASSI system to the original three-dimensional hyperspectral data (e.g., an 83-band hyperspectral image is compressed into two-dimensional measurements). This invention is optimized for scenarios with a compression ratio ≥ 30:1 and the number of bands ≥ 83, solving the problem of low reconstruction accuracy in the prior art under limited computing resources.
[0066] 12. Long-Range Dependency Spatial long-range dependence: The correlation of pixels with a relatively large distance in the image (such as the global consistency of edges and textures) is difficult to model by traditional CNNs. The present invention effectively captures it through multi-directional scanning of GHSB.
[0067] Spectral long-range dependence: The context relationship between different spectral bands (such as the correlation between band λ1 and λ n ), and BPSB realizes two-way modeling through two-way scanning.
[0068] V. Basic technical terms 13. Convolutional Neural Network (CNN) A neural network that extracts local spatial features through convolutional operations. In the present invention, it is used as a comparative technology, and its defect is the lack of long-range dependence modeling ability.
[0069] 14. Transformer architecture A network architecture that models global dependence through self-attention mechanism. The defect is that the computational complexity grows quadratically with the input scale (O(N 2 )). The present invention avoids this problem through lightweight design.
[0070] 15. Layer Normalization Normalize the input of the neural network layer to stabilize the training process. In the present invention, it is used inside the SSAB module to improve the feature stability.
[0071] The above term explanations are strictly based on the technical solutions of the present invention to clarify the functions, mechanisms, and technical advantages of each module, and ensure that those skilled in the art can accurately understand the core innovation points and technical details of the invention.
[0072] Embodiment III This embodiment also provides an electronic device. Refer to Figure 4 , including a memory 404 and a processor 402. A computer program is stored in the memory 404, and the processor 402 is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0073] Specifically, the above-mentioned processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0074] Among them, the memory 404 may include a mass storage 404 for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 404 may include removable or non-removable (or fixed) media. Where appropriate, the memory 404 may be internal or external to the data processing device. In a particular embodiment, the memory 404 is non-volatile memory. In a particular embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended date out dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0075] The memory 404 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.
[0076] By reading and executing the computer program instructions stored in the memory 404, the processor 402 implements any one of the hyperspectral image reconstruction methods based on high compression ratio snapshot compression in the above embodiments.
[0077] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.
[0078] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above network may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0079] The input / output device 408 is used to input or output information.
[0080] Embodiment 4 This embodiment also provides a readable storage medium, in which a computer program is stored. The computer program includes program codes for controlling a process to execute the process, and the process includes the hyperspectral image reconstruction method based on high compression ratio snapshot compression according to Embodiment 1.
[0081] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.
[0082] Generally, various embodiments can be implemented in hardware or special circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented by firmware or software executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representations, it should be understood that, as a non-limiting example, the blocks, devices, systems, technologies, or methods described herein can be implemented in hardware, software, firmware, special circuits or logic, general hardware or a controller, or other computing devices, or some combination thereof.
[0083] Embodiments of the present invention can be implemented by computer software that is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components that are configured to perform the embodiments when the program runs. The one or more computer-executable components can be at least one software code or a part thereof. Additionally, in this regard, it should be noted that any box in the logical flow, as Figure 1 shown, can represent a program step, or interconnected logic circuits, boxes, and functions, or a combination of program steps and logic circuits, boxes, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media are non-transitory media.
[0084] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0085] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A hyperspectral image reconstruction method based on high compression ratio snapshot compression, characterized in that It includes the following steps: S1. Obtain the two-dimensional compressed measurement values collected by the CASSI system; S2. Perform a dispersion inverse operation on the two-dimensional compressed measurement values to generate an initial three-dimensional hyperspectral signal; S3. Stitch the initial hyperspectral signal with a physical mask and generate an initial feature map through a convolution operation; S4. Input the initial feature map into a reconstruction model based on a U-shaped network, and perform multi-level feature extraction and reconstruction through an encoder-bottleneck-decoder structure, where the core unit of the reconstruction model is a spatial-spectral aggregation block; The spatial-spectral aggregation block sequentially includes: Grouped hybrid spatial module: Group the input features along the channel dimension, and capture long-range spatial dependencies through multi-directional scanning and inter-group residual connections; Bidirectional patch spectral module: Divide the features into small blocks and stitch them along the spectral dimension, and capture spectral context information through bidirectional scanning; Gated feed-forward network: Integrate spatial and spectral features through a gating mechanism to optimize the output representation; S5. Add the residual features output by the decoder to the initial input to generate a reconstructed hyperspectral image.
2. The hyperspectral image reconstruction method based on high compression ratio snapshot compression according to claim 1, characterized in that, In step S4, the implementation of the encoder-bottleneck-decoder structure includes: The encoder gradually compresses the spatial dimension and expands the channels through N1 consecutive spatial-spectral aggregation blocks and downsampling operations; The bottleneck part extracts high-level feature representations through N3 spatial-spectral aggregation blocks; The decoder restores the spatial details through upsampling and skip connections, and stitches the features of the skip connections with the corresponding stages of the encoder.
3. A hyperspectral image reconstruction method based on high compression ratio snapshot compression according to claim 1, characterized in that, In step S4, the specific operation of the grouped hybrid spatial module is: Divide the input data into G groups along the channel dimension, and each group of data is scanned in multiple directions through the Spatial-SSM module, and the scanning directions include at least two of left to right, right to left, up to down, and down to up; Except for the first group, the input of each group fuses the output of the previous group's Spatial-SSM module to form an inter-group residual connection; Stitch the outputs of all groups and integrate the information through a 1×1 convolution.
4. The hyperspectral image reconstruction method based on high compression ratio snapshot compression according to claim 3, wherein, In step S4, the Spatial-SSM module unfolds the input features into a one-dimensional sequence and models the spatial pixel correlation through a state transition equation, and the state transition equation is: ; Among them, h t is the hidden state of the current spatial position, used to store and transmit spatial context information; h t-1 is the hidden state of the previous spatial position, reflecting the dependence relationship between spatial positions; x t is the input feature of the current spatial position; is calculated by learnable parameters ΔA, ΔB through and is used to adjust the hidden state h t-1 and the input x t to the contribution weights of the current hidden state h t , introducing a non-linear transformation to capture complex spatial relationships; C, D are learnable parameters used to map the hidden state h t and the input x t to the output y t .
5. The hyperspectral image reconstruction method based on high compression ratio snapshot compression according to claim 1, characterized in that, In step S4, the processing steps of the bidirectional patch spectral module include: Divide the input features into blocks of size M×M, stitch the flattened pixels of each block along the spectral dimension to form a sequence representation in the spectral dimension; perform bidirectional scanning on the spectral sequence, and capture the context dependencies between spectral bands through the Spectral-SSM module, and the bidirectional scanning includes a forward scan from low frequency to high frequency and a backward scan from high frequency to low frequency.
6. The hyperspectral image reconstruction method based on high compression ratio snapshot compression according to claim 1, characterized in that, In step S4, the processing steps of the gated feed-forward network include: Expand the channel dimension through a 1×1 convolution and divide the features into two groups along the channel; The first group captures spatial context through a 3×3 convolution and introduces non-linearity through the GELU function, and the second group multiplies the features of the first group through a gating mechanism; Integrate the features of the two groups through a 1×1 convolution and output the optimized spatial-spectral fusion features.
7. A hyperspectral image reconstruction method based on high compression ratio snapshot compression according to any one of claims 1-6, characterized in that, For the hyperspectral image corresponding to the high compression ratio scenario, the number of bands ≥ 83 and the compression ratio ≥ 30:
1.
8. A hyperspectral image reconstruction device based on high compression ratio snapshot compression, characterized in that, It includes: The dispersion inverse operation module performs a dispersion inverse operation on the two-dimensional compressed measurement values collected by the CASSI system to generate an initial three-dimensional hyperspectral signal; The mask stitching module stitches the initial hyperspectral signal with the physical mask and generates an initial feature map through a convolution operation; The reconstruction module inputs the initial feature map into a reconstruction model based on a U-shaped network, performs multi-level feature extraction and reconstruction through an encoder-bottleneck-decoder structure, and the core unit of the reconstruction model is the spatial-spectral aggregation block; The output module adds the residual features output by the decoder to the initial input to generate a reconstructed hyperspectral image; The spatial-spectral aggregation block sequentially includes: The grouped hybrid spatial module: groups the input features along the channel dimension, and captures long-range spatial dependencies through multi-directional scanning and inter-group residual connections; The bidirectional patch spectral module: divides the features into small blocks and stitches them along the spectral dimension, and captures spectral context information through bidirectional scanning; The gated feed-forward network: integrates spatial and spectral features through a gating mechanism to optimize the output representation.
9. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the hyperspectral image reconstruction method based on high-compression-ratio snapshot compression according to any one of claims 1 to 7.
10. A readable storage medium, characterized in that, A computer program is stored in the readable storage medium, and the computer program includes program codes for controlling a process to execute the process, and the process includes the hyperspectral image reconstruction method based on high-compression-ratio snapshot compression according to any one of claims 1 to 7.
Citation Information
Cited By
Hyperspectral image reconstruction method and application thereof
CN122473004A
Hyperspectral image reconstruction method and application thereof
CN122473004B