Hyperspectral image reconstruction method and system combining multi-scale feature extraction and mask gating mechanism

Through multi-scale feature extraction and mask gated mechanism hyperspectral image reconstruction network MSMGNet, the problem of high computing complexity and insufficient reconstruction quality in the prior art is solved, and efficient and lightweight hyperspectral image reconstruction is achieved, which is suitable for resource-constrained equipment deployment.

CN120070615APending Publication Date: 2025-05-30HANGZHOU DIANZI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510086786.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing hyperspectral image reconstruction algorithm has high computational complexity, is difficult to deploy on resource-constrained devices, and the reconstruction quality is insufficient.

Method used

The hyperspectral image reconstruction network MSMGNet, which combines multi-scale feature extraction and mask gating mechanism, is used to realize efficient hyperspectral image reconstruction through the hyperspectral feature extraction module, mask gating module, mask feature extraction module and fusion output module.

Benefits of technology

Effectively capture local details and global information of hyperspectral images, improve reconstruction accuracy, suppress noise interference, is suitable for edge device deployment, and can effectively preserve image texture and edge details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070615A_ABST
    Figure CN120070615A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral image reconstruction method and system combining multi-scale feature extraction and a mask gating mechanism. By obtaining a two-dimensional measurement image and a mask image, a hyperspectral image reconstruction network combining multi-scale feature extraction and a mask gating mechanism is constructed, and efficient reconstruction of a hyperspectral image is realized. Through the multi-scale feature extraction module, local details and global information of a hyperspectral image can be effectively captured, and the reconstruction precision is improved; a mask gating mechanism is introduced through the mask gating module, so that a dynamic mask can be generated, feature fusion is dynamically adjusted, noise interference is suppressed, the image quality is further improved, and redundant features are effectively suppressed; a lightweight design strategy is adopted, the model parameter quantity is remarkably reduced, the method is suitable for edge device deployment, efficient reconstruction of a real scene hyperspectral image is achieved, and image texture and edge details can be effectively reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hyperspectral image processing, and particularly relates to a hyperspectral image reconstruction method and system combining multi-scale feature extraction and mask gating mechanism. Background Art

[0002] Hyperspectral images (HSIs) can capture rich spectral and spatial information within a wide spectral range, and thus have wide applications in fields such as remote sensing, medical imaging, and agricultural monitoring. However, traditional hyperspectral imaging methods (such as point scanning and line scanning) are inefficient and difficult to meet the requirements of real-time applications. In recent years, coded aperture snapshot spectral imaging (CASSI) technology has greatly improved the imaging speed and system compactness by compressing three-dimensional spectral data into two-dimensional measurements. However, existing CASSI-based signal reconstruction algorithms usually have high computational complexity and rely on high-performance computing resources, which limits their applications in resource-constrained devices (such as drones and mobile devices).

[0003] Most of the existing deep learning-based methods, although achieving remarkable results in hyperspectral image reconstruction, usually require a large amount of computing resources and are difficult to be deployed on edge devices. Among them, Transformer-based methods perform excellently in global feature modeling, but have high requirements for computing performance and memory. In contrast, CNN-based algorithms, with their efficient local receptive field operations, weight sharing mechanism, and mature hardware optimization, are suitable for applications in resource-constrained environments due to their low complexity, low energy consumption, and excellent performance on small-scale datasets.

[0004] Therefore, there is a need for a lightweight and efficient hyperspectral image reconstruction method to balance accuracy and computational efficiency. Summary of the Invention

[0005] The purpose of the present invention is to solve the problems of high computational complexity, insufficient reconstruction quality, and difficulty in being deployed on edge devices existing in the prior art, and to provide a hyperspectral image reconstruction method and system combining multi-scale feature extraction and mask gating mechanism.

[0006] In a first aspect, the present invention provides a hyperspectral image reconstruction method combining multi-scale feature extraction and mask gating mechanism, the method comprising: Step S1, generating a two-dimensional measurement image through physical modulation according to a three-dimensional hyperspectral real image and a mask image; constructing a data set with the two-dimensional measurement image and using the three-dimensional hyperspectral real image as a label; and then dividing the data set into a training data set and a test data set according to a ratio; Step S2, constructing a hyperspectral image reconstruction network MSMGNet combining multi-scale feature extraction and mask gating mechanism, and training and testing it using the training set and the test set respectively; Step S3: Use the hyperspectral image reconstruction network MSMGNet that combines multi-scale feature extraction and mask gating mechanism to efficiently reconstruct the hyperspectral image; Among them, the hyperspectral image reconstruction network MSMGNet that combines multi-scale feature extraction and mask gating mechanism includes a hyperspectral feature extraction module, a mask gating module, a mask feature extraction module, and a fusion output module; The mask feature extraction module extracts mask region features from the mask image; The mask gating module further extracts mask region weight coefficients from the mask region features; The hyperspectral feature extraction module extracts hyperspectral region features of different scales from the two-dimensional measurement image; The fusion output module fuses the mask region weight coefficients and the hyperspectral region features to obtain a hyperspectral reconstructed image.

[0007] In a second aspect, the present invention provides a hyperspectral image reconstruction system, including: A data acquisition module that acquires a two-dimensional measurement image and a mask image; A reconstruction module that inputs the two-dimensional measurement image and the mask image into the hyperspectral image reconstruction network MSMGNet that combines multi-scale feature extraction and mask gating mechanism to obtain a hyperspectral image reconstruction result.

[0008] Compared with the prior art, the present invention has the following beneficial effects: Through the multi-scale feature extraction module, the present invention can effectively capture the local details and global information of the hyperspectral image, improving the reconstruction accuracy; By introducing the mask gating mechanism through the mask gating module, the present invention can generate a dynamic mask, thereby dynamically adjusting feature fusion, suppressing noise interference, further improving the image quality, and effectively suppressing redundant features; The present invention adopts a lightweight design strategy, significantly reducing the number of model parameters and being suitable for deployment on edge devices; The present invention realizes the efficient reconstruction of hyperspectral images in real scenes and can effectively retain image texture and edge details. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0010] Figure 1 It is a schematic diagram of the module flow of the present invention.

[0011] Figure 2 It is a schematic diagram of the structure of the hyperspectral image reconstruction network MSMGNet.

[0012] Figure 3 It is a schematic diagram of the structure of the multi-scale feature extraction layer.

[0013] Figure 4 It is a comparison diagram of simulation scenes.

[0014] Figure 5 It is a comparison diagram of real scenes.

[0015] Figure 6 In (a) - Figure 6 In (b) respectively show the results of reconstructing two spectral curves in the simulation scene. Specific implementation manner

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0017] This embodiment provides a hyperspectral image reconstruction method combining multi-scale feature extraction and mask gating mechanism. Refer to Figure 1 The method includes: Step S1: Generate a two-dimensional measurement image through physical modulation according to the three-dimensional hyperspectral real image and the mask image; construct a data set with the two-dimensional measurement image and use the hyperspectral real image as the label; then divide the data set into a training data set and a test data set according to a ratio; Step S2: Construct a hyperspectral image reconstruction network MSMGNet combining multi-scale feature extraction and mask gating mechanism, and train and test it respectively using the training set and the test set.

[0018] Step S3: Use the trained and tested hyperspectral image reconstruction network MSMGNet combining multi-scale feature extraction and mask gating mechanism to achieve efficient reconstruction of hyperspectral images.

[0019] In one implementation manner, in step S1, the two-dimensional measurement image is obtained by modulating the hyperspectral real image with the mask image, passing it through a prism, shearing and offsetting the two-dimensional images of different bands along the y-axis, and finally synthesizing a two-dimensional measurement image; the mask image is a physical mask with a specific pattern; In one implementation manner, refer to Figure 2The hyperspectral image reconstruction network MSMGNet that combines multi-scale feature extraction and mask gating mechanism adopts an encoder-decoder architecture in the style of Unet, integrating a hyperspectral feature extraction module, a mask gating module (Mask Gating Module), a mask feature extraction module (Mask Filter Module), and a fusion output module. The network structure is as follows:

[0020] The encoder is responsible for extracting preliminary features from the input compressed hyperspectral image data; The decoder combines multi-scale features and dynamic masks to generate the final hyperspectral image reconstruction result.

[0021] Exemplarily, the mask feature extraction module extracts mask region features from the mask image; it includes a 3x3 convolutional layer, a PReLU activation function, and a 3x3 convolutional layer connected in sequence. The specific implementation process:

[0022] First, the mask image passes through a convolutional layer with C convolutional kernels and a kernel size of 3x3, then through a PReLU non-linear activation function, and finally through a convolutional layer with C convolutional kernels and a kernel size of 3x3 to output the mask region features.

[0023] The mask gating module further extracts mask region weight coefficients from the mask region features to selectively enhance or suppress the input features; specifically, the input features pass through a convolutional layer and a non-linear activation function to generate a dynamic mask, and then multiply with the input features to complete the weighting process. It includes a 1x1 convolutional layer, a PReLU non-linear activation function, a 3x3 convolutional layer, and a Sigmoid non-linear activation function connected in sequence. The specific implementation process:

[0024] The mask region features (with a size of C*W*H) output by the mask feature extraction module pass through a convolutional layer with C convolutional kernels and a kernel size of 1×1, then followed by a PReLU non-linear activation function, and then through a convolutional layer with n convolutional kernels and a kernel size of 3×3, followed by a Sigmoid non-linear activation function, and then output the mask region weight coefficients.

[0025] The hyperspectral feature extraction module extracts hyperspectral region features of different scales from the two-dimensional measurement image; refer to Figure 3 It includes a channel attention layer, a multi-scale feature extraction layer (MixStage), and a fusion splicing layer; specifically: The two-dimensional measurement image is respectively fed into the channel attention layer and the multi-scale feature extraction layer. The multi-scale feature extraction layer extracts features of different scales, and the multi-scale information is combined through the fusion splicing layer. At the same time, residual connections are introduced to ensure the smooth flow of information and avoid gradient disappearance.

[0026] The multi-scale feature extraction layer adopts three resolution branches (1×, 2×, 4×) to obtain information at different levels, including a first semantic feature extraction layer, a second semantic feature extraction layer, and a third semantic feature extraction layer; specifically: The first semantic feature extraction layer performs initial semantic feature (i.e., detail feature) extraction (1× resolution branch) on the two-dimensional measurement image (where H is the height, W is the width, and C is the number of channels): Among them, represents the activation function, is a 3×3 depthwise separable convolution; The second semantic feature extraction layer performs middle-level semantic feature extraction (2× downsampling branch) on the initial semantic feature : Among them, represents the 2× downsampling operation, is a 3x3 depthwise separable convolution for capturing middle-scale information.

[0027] The third semantic feature extraction layer performs deep semantic feature (i.e., global feature) extraction (4× downsampling branch) on the middle-level semantic feature : Among them, represents the 2× downsampling operation, is a 3×3 depthwise separable convolution for capturing middle-scale information.

[0028] The multi-scale feature extraction module is designed as follows: The two-dimensional measurement image with size C×W×H is respectively sent into the channel attention layer, the first semantic feature extraction layer, the second semantic feature extraction layer, and the third semantic feature extraction layer; then the outputs of the above four channels are concatenated, and at this time the image size is 4C×W×H. The concatenated image passes through a convolution layer with the number of convolution kernels being C and the convolution kernel size being 1×1, and then is followed by a PReLU non-linear activation layer. Finally, the obtained output is added to the original two-dimensional measurement image to output the two-dimensional measurement image after extracting semantic features, with size C×W×H. The two-dimensional measurement image after extracting semantic features is multiplied by the masked image passing through the masked gated convolution module, and the obtained result is then combined with the original input two-dimensional measurement image to obtain the output.

[0029] The channel attention layer obtains local saliency and global information through max pooling and average pooling operations respectively, and weights the importance of each channel, specifically as follows: For the input feature , through global max pooling and average pooling operations, generate channel global information: Among them, and respectively represent average pooling and max pooling operations on the spatial dimension. Generate the attention weights for each channel by passing the pooled features through a shared fully connected layer and an activation function:

[0030] Where: represent two fully connected layers, is the activation function, is the Sigmoid activation function, is the finally generated channel attention weight.

[0031] Apply the channel weight to the original input feature , and generate the output feature with enhanced channels: The channel attention layer is designed as follows: The channel attention layer is used to learn the features of different channels, including the max-pooling branch and the global pooling branch. The max-pooling branch includes a 1×1 max-pooling layer, a 1×1 convolutional layer with the number of channels being C / 2, a ReLU non-linear activation function, a 1×1 convolutional layer with the number of channels being C, and a Sigmoid non-linear activation function connected in sequence. The global pooling branch includes a 1×1 global pooling layer, a 1×1 convolutional layer with the number of channels being C / 2, a ReLU non-linear activation function, a 1×1 convolutional layer with the number of channels being C, and a Sigmoid non-linear activation function connected in sequence. Finally, the outputs of these two branches are concatenated and then multiplied by the original input. That is, the input two-dimensional measurement image with the size of C*W*H is sent into the adaptive max-pooling layer and the adaptive average pooling layer respectively. Then, both of these two outputs pass through a convolutional layer with the number of convolutional kernels being C / 2 and the size of the convolutional kernel being 1×1, and then followed by a PReLU non-linear activation function. Then, it passes through a convolutional layer with the number of convolutional kernels being C and the size of the convolutional kernel being 1×1, and then followed by a Sigmoid non-linear activation function. Finally, the outputs of these two are concatenated together. At this time, the size of the output image is 2C×W×H. This output passes through a convolutional layer with the number of convolutional kernels being C and the size of the convolutional kernel being 1×1, and then followed by a Sigmoid non-linear activation function. At this time, the output is C×W×H. This output is multiplied by the original input two-dimensional measurement image to obtain the final output with the size of C×W×H.

[0032] The outputs of the channel attention layer, the first semantic feature extraction layer, the second semantic feature extraction layer, and the third semantic feature extraction layer are sent into the fusion and concatenation layer for concatenation, convolution operation, and PReLU non-linear activation to obtain semantic features; finally, the semantic features are added to the original two-dimensional measurement image to output the hyperspectral region features.

[0033] The fusion output module fuses the mask region weight coefficient and the hyperspectral region features to obtain the hyperspectral reconstructed image. Specifically: the hyperspectral region features are multiplied by the mask region weight coefficient passing through the mask gated convolution module, then the multiplication result and the original input two-dimensional measurement image are weighted and concatenated, and finally, through a 1×1 convolutional layer, the reconstructed hyperspectral image is output.

[0034] Experiments and Result Analysis A. Experimental Setup Experimental Setup and Datasets: We used both simulated and real hyperspectral image (HSI) datasets to evaluate the proposed method. In the simulated scenario, the publicly available datasets KAIST (2704×3376×31) and CAVE (512×512×31) were used. Consistent with the setup of TSA-Net, the CAVE dataset was used for training, and 10 selected scenes in the KAIST dataset were used for test data. To prepare the training data, the CAVE dataset was cropped into small patches of size 256×256×31 by sliding windows, and data augmentation techniques such as flipping and rotation were applied to improve the diversity of training. The experiments focused on analyzing 28 spectral bands, consistent with the configuration of TSA-Net. In the real-scenario experiment, the compressed measurement data with a spatial resolution of 660×714 collected and configured by TSA-Net were used to evaluate the performance of the method in actual HSI reconstruction tasks. Our method was implemented in Pytorch. The ADAM optimizer and a cosine annealing learning rate schedule with 300 epochs were used to train the model. The initial learning rate was set to 4×10 −4 , and the batch size was set to 2. All experiments were conducted on a single RTX2080Ti GPU.

[0035] Evaluation Metrics: To evaluate the performance of HSI reconstruction, we adopted two metrics: Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM). Meanwhile, to evaluate the complexity of the model, the number of parameters Params(M) in millions and the Floating Point Operations (FLOPs) were used. The lower these two quantities are, the simpler the model structure is.

[0036] We compared the proposed method with other state-of-the-art open-source methods, including three model-based methods (GAP-TV, DeSCI, TWIST) and six CNN-based methods ( -Net, TSA-Net, DGSMP, ADMM-Net, GAP-Net, HD-Net).

[0037] B. Results Comparison Simulated Scenario Experiments: Figure 4Shows the comparison of the simulated HSI reconstruction results of the second scenario in the KAIST dataset. Three out of 28 spectral channels were selected (487.0 nm, 558.5 nm, and 614.5 nm). It can be observed from the figure that traditional methods such as TWIST, GAP-TV, and DeSCI show significant blurring in visual effects, especially in complex scenarios (such as 558.5 nm and 614.5 nm), with serious loss of edge texture information and poor color fidelity. In contrast, MSMGNet shows clearer object edges and higher color fidelity in the reconstruction results, especially in the restoration of high-frequency details and structured regions (such as the edges of small objects), significantly outperforming traditional methods.

[0038] Real-scene experiments: Figure 5 Shows the comparison results of MSMGNet and other methods on a real dataset. At wavelengths of 558.5 nm, 594.5 nm, and 614.5 nm, MSMGNet is significantly superior to other methods in visual performance, especially in terms of detail preservation and noise suppression. In contrast, traditional methods such as TWIST and GAP-TV show weaker detail performance and more obvious noise at higher wavelengths. Although -Net and TSA-Net improve the reconstruction quality, their high computational requirements remain a significant limitation. And MSMGNet achieves the best balance between visual quality and quantization results with lower computational complexity.

[0039] Table 1 presents the quantitative analysis of the simulation results for 10 scenarios (the higher the PSNR and SSIM, the better)

[0040] As shown in the above table, compared with traditional compressive sensing reconstruction methods (such as TWIST, GAP-TV, DeSCI) and deep learning-based methods (such as -Net and TSA-Net), our proposed MSMGNet shows significant advantages in multiple metrics.

[0041] The average PSNR of MSMGNet reaches 34.53 dB, which is better than most of the comparison methods. At the same time, our method significantly reduces the number of parameters while maintaining comparable performance with HDNet, making it more suitable for deployment on edge devices. It is worth noting that in complex scenarios such as S4, MSMGNet obtains the highest score, demonstrating its excellent performance in detail restoration and structure preservation. The SSIM metric also confirms the ability of MSMGNet to retain texture and structure, with an average SSIM of 0.940, comparable to HDNet but significantly higher than -Net and TSA-Net. In contrast, the PSNR and SSIM scores of traditional methods (such as TWIST and GAP-TV) are significantly lower than those of MSMGNet, indicating that traditional methods have limited performance in dealing with highly complex scenarios.

[0042] The efficiency of MSMGNet is reflected in its lightweight design, which contains 0.87M parameters and 35.81 GFLOPs, significantly lower than other CNN-based methods, such as DGMSP (3.76M parameters, 660.65G FLOPs). This reduction in computational complexity makes MSMGNet very suitable for resource-constrained environments, such as edge computing and mobile devices, while still achieving good performance. Its efficient design ensures faster inference speed and lower resource requirements, making it an ideal solution for practical hyperspectral image reconstruction tasks.

[0043] Table 2 shows the ablation experiments on simulated data, obtaining the average PSNR and SSIM of 10 simulation scenarios

[0044] As shown in the above table, ablation experiments were conducted on a publicly available simulated HSI dataset to study the effects of each module. The baseline model is the hyperspectral feature extraction module in Figure 2 , using the initialized HSI as the input. Using only the baseline model, the reconstructed result is 34.04 dB. In the table, model (2) represents adding the mask feature extraction module with the mask M as the input to the baseline model. Model (3) represents adding both the mask feature extraction module and the mask gating module. When the mask gating module is introduced, it can be observed that the result is improved by 0.32 dB compared to the baseline model. This is because the introduction of gated convolutional features enables the model to reconstruct based on the masked HSI, thus improving the quality of the reconstructed HSI.

[0045] Comparison of spectral curves in simulation scenarios: Figure 6 In (a)- Figure 6Figure (b) shows the results of spectral curve reconstruction in the simulation scenario, and two 10×10 regions of interest are selected for analysis. The results show that our method has comparable spectral consistency with other methods and a high similarity with the true spectrum.

[0046] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A hyperspectral image reconstruction method combining multi-scale feature extraction and mask gating mechanism, characterized in that: The method comprises: Step S1, generating a two-dimensional measurement image through physical modulation according to the three-dimensional hyperspectral real image and the mask image; constructing a data set using the two-dimensional measurement image, and using the three-dimensional hyperspectral real image as a label; and then dividing the data set into a training data set and a test data set according to a ratio; Step S2, constructing a hyperspectral image reconstruction network MSMGNet combining multi-scale feature extraction and mask gating mechanism, and training and testing it using a training set and a test set respectively; Step S3, using the hyperspectral image reconstruction network MSMGNet that combines the trained and tested multi-scale feature extraction and mask gating mechanism to achieve efficient reconstruction of the hyperspectral image; The hyperspectral image reconstruction network MSMGNet combining multi-scale feature extraction and mask gating mechanism includes a hyperspectral feature extraction module, a mask gating module, a mask feature extraction module, and a fusion output module; A mask feature extraction module extracts mask area features from a mask image; The mask gating module further extracts the mask area weight coefficient based on the mask area features; Hyperspectral feature extraction module, extracting hyperspectral regional features of different scales from two-dimensional measurement images; The fusion output module fuses the mask region weight coefficient and the hyperspectral region features to obtain the hyperspectral reconstructed image.

2. The method according to claim 1, characterized in that: In step S1, the two-dimensional measurement image is a three-dimensional hyperspectral real image modulated by a mask image, and then passed through a prism to make the two-dimensional images of different bands shear and shift along the y-axis, and finally synthesize a two-dimensional measurement image; the mask image is a physical mask with a specific pattern.

3. The method according to claim 1, characterized in that: In the hyperspectral image reconstruction network MSMGNet that combines multi-scale feature extraction with a mask gating mechanism, the mask feature extraction module includes a 3×3 convolutional layer, a PReLU activation function, and a 3×3 convolutional layer that are connected in sequence.

4. The method according to claim 3, characterized in that: The specific implementation process of the mask feature extraction module is as follows: First, the mask image passes through a convolution layer with C convolution kernels and a size of 3×3, then passes through a PReLU nonlinear activation function, and finally passes through a convolution layer with C convolution kernels and a size of 3×3 to output the mask area features.

5. The method according to claim 1, characterized in that: In the hyperspectral image reconstruction network MSMGNet that combines multi-scale feature extraction with a mask gating mechanism, the mask gating module includes a 1×1 convolutional layer, a PReLU nonlinear activation function, a 3×3 convolutional layer, and a Sigmoid nonlinear activation function connected in sequence.

6. The method according to claim 5, characterized in that: The specific implementation process of the mask gating module is as follows: The mask area features output by the mask feature extraction module are passed through a convolution layer with C convolution kernels and a convolution kernel size of 1×1, and then connected to a PReLU nonlinear activation function, and then passed through a convolution layer with n convolution kernels and a convolution kernel size of 3×3, connected to a Sigmoid nonlinear activation function, and then the mask area weight coefficient is output.

7. The method according to claim 1, characterized in that: In the hyperspectral image reconstruction network MSMGNet that combines multi-scale feature extraction with mask gating mechanism, the hyperspectral feature extraction module includes a channel attention layer, a multi-scale feature extraction layer, and a fusion splicing layer; specifically: The two-dimensional measurement image is sent to the channel attention layer and the multi-scale feature extraction layer respectively. The multi-scale feature extraction layer is used to extract features of different scales, and the multi-scale information is combined through the fusion splicing layer, and the residual connection is introduced.

8. The method according to claim 7, characterized in that: The multi-scale feature extraction layer includes a first semantic feature extraction layer, a second semantic feature extraction layer, and a third semantic feature extraction layer; Specifically: The first semantic feature extraction layer performs initial semantic feature extraction on the two-dimensional measurement image, the second semantic feature extraction layer performs intermediate semantic feature extraction on the initial semantic features, and the third semantic feature extraction layer performs deep semantic feature extraction on the intermediate semantic features. The outputs of the channel attention layer, the first semantic feature extraction layer, the second semantic feature extraction layer, and the third semantic feature extraction layer are sent to the fusion splicing layer for splicing, convolution operation, and PReLU nonlinear activation to obtain semantic features; finally, the semantic features are added to the original two-dimensional measurement image to output the hyperspectral region features.

9. The method according to claim 1, characterized in that: In the hyperspectral image reconstruction network MSMGNet that combines multi-scale feature extraction with mask gating mechanism, the fusion output module multiplies the hyperspectral region features with the mask region weight coefficients passed through the mask gated convolution module, and then weightedly concatenates the multiplication result and the original input two-dimensional measurement image, and finally passes through a 1×1 convolution layer to output the reconstructed hyperspectral image.

10. A hyperspectral image reconstruction system for implementing the method according to any one of claims 1 to 9, characterized in that: include: A data acquisition module, which acquires a two-dimensional measurement image and a mask image; The reconstruction module inputs the two-dimensional measurement image and the mask image into the hyperspectral image reconstruction network MSMGNet that combines the trained and tested multi-scale feature extraction and mask gating mechanism to obtain the hyperspectral image reconstruction result.

Citation Information

Cited By

  • Single-exposure compression imaging reconstruction method and system fused with pre-training language model

    CN122312380A

  • Hyperspectral snapshot compression imaging reconstruction method based on degradation learning and spectral diffusion

    CN122435095A