Spectral super-resolution reconstruction method and system based on spatial spectrum Mama network
Through the spectral super-resolution reconstruction method based on the null spectrum Mamba network, the scanning strategy and fusion mechanism are used to extract the null spectrum information, which solves the problem of difficult to balance global information extraction and computing efficiency in the prior art, and realizes efficient spectral super-resolution reconstruction.
Patent Information
- Application Number
- CN202510174719.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-17
AI Technical Summary
Existing spectral super-resolution reconstruction methods based on CNN and Transformer are difficult to balance between global information extraction and computational efficiency.
A spectral super-resolution reconstruction method based on the null spectrum Mamba network is proposed, and the null spectrum information is efficiently extracted through scanning strategies and fusion mechanisms, including shallow feature extraction, deep feature extraction of null spectrum and spectral reconstruction modules.
It realizes powerful global feature extraction capabilities in hyperspectral images, while maintaining a linear increase in computational complexity, solving the problem of balance between performance and efficiency of existing methods.
Smart Images

Figure CN120163709A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of artificial intelligence and hyperspectral image super-resolution reconstruction, and specifically relates to a spectral super-resolution reconstruction method and system based on an air-spectral Mamba network. Background Art
[0002] Hyperspectral images are a special type of image data that cover dozens to hundreds of bands within a continuous spectral range, can provide richer spectral information than RGB images, and have been widely applied in fields such as smart agriculture, geological exploration, and environmental monitoring. Spectral super-resolution reconstruction, as an important technology for reconstructing hyperspectral images from RGB images, can effectively overcome problems such as high time cost and large resource consumption existing in the process of traditional hyperspectral image acquisition.
[0003] In recent years, with the rapid development of deep learning theory, researchers have proposed a series of spectral super-resolution reconstruction methods based on deep learning. Most of these methods are based on CNN and Transformer and have achieved relatively high reconstruction accuracy. For example, Li et al. proposed a progressive spatial information-guided deep aggregation convolutional neural network, including multiple dense residual channel affinity learning blocks and a spatial guidance propagation module as the backbone, in the literature "Li J, Du S, Song R, et al. Progressive spatial information-guided deep aggregation convolutional network for hyperspectral spectral super-resolution[J]. IEEE Transactions on Neural Networks and Learning Systems, 2023." He et al. combined Transformer with ResNet and proposed a ResNet-based dense spectral transformer to achieve spectral super-resolution reconstruction of multi-spectral remote sensing images in "He J, Yuan Q, Li J, et al. DsTer: A dense spectral transformer for remote sensing spectral super-resolution[J]. International Journal of Applied Earth Observation and Geoinformation, 2022, 109: 102773.", which can meet the needs of three-dimensional data processing of remote sensing images and learning remote relationships.
[0004] However, there are still some urgent problems to be solved in existing methods: CNN-based methods are limited by the local receptive field of the convolution kernel and cannot effectively capture global scene information; and although Transformer-based methods can model long sequence dependencies, the attention mechanism they introduce will cause the computational complexity to grow quadratically, making it difficult to achieve an ideal balance between model performance and computational efficiency. Summary of the invention
[0005] The present invention discloses a spectral super-resolution reconstruction method and system based on a spatial-spectrum Mamba network, aiming to effectively solve the key problem of the spectral super-resolution reconstruction methods based on CNN and Transformer in the prior art that it is difficult to balance global information extraction and computational efficiency.
[0006] The present invention discloses a spatial spectrum Mamba network, which mainly realizes efficient extraction of spatial spectrum information through scanning strategy and fusion mechanism. The network can be simply divided into three steps: shallow feature extraction, spatial spectrum deep feature extraction, and spectrum reconstruction, as follows:
[0007] 1. Shallow feature extraction: First, a shallow feature extraction module is built, which is responsible for extracting preliminary features from the hyperspectral image. A single 3×3 convolution can be used here to obtain low-frequency information such as color and texture of the input image.
[0008] 2. Spatial spectrum deep feature extraction: The extraction of deep spatial spectrum features is completed through multiple spatial spectrum Mamba modules, which are specifically divided into two parts: spatial spectrum scanning module and spatial spectrum fusion module.
[0009] A spatial spectrum scanning module is constructed. The spatial features focus on the spatial distribution and arrangement of pixels in the image, covering geometric structure information such as edges, textures and shapes; while the spectral features focus on the reflection or radiation intensity of each pixel at different wavelengths. The two have different emphases on information type and processing methods. Therefore, spatial features and spectral features are extracted through two parallel branches respectively. Different scanning strategies are designed for the two features.
[0010] Construct a spatial-spectral fusion module to re-weight the features of the spatial and spectral branches. In terms of spatial interaction, a 1×1 convolution with an activation function is used to extract the input features. Here, a spatial feature map containing only a single channel is generated; in terms of spectral interaction, a spectral feature map containing only the channel dimension is generated, so an adaptive average pooling layer is added on the original basis, and the spatial-spectral feature map is finally multiplied with the input features and re-joined.
[0011] 3. Spectral reconstruction: A spectral reconstruction module is constructed to achieve spectral super-resolution. Here, a 3×3 convolution is used to recombine the extracted feature information. At the same time, the feature map will be restored to the original resolution, and finally the hyperspectral image we need will be generated.
[0012] The present invention provides a spectral super-resolution reconstruction method based on a spatial spectrum Mamba network, comprising the following steps:
[0013] Step 1, extracting shallow texture features of the image;
[0014] Step 2, constructing a spatial-spectrum Mamba module, including a spatial scanning submodule for mining deeper spatial feature information of hyperspectral images, a spectral scanning submodule for extracting spectral features, and a spatial-spectrum fusion submodule for realizing adaptive interaction and fusion of spatial and spectral features;
[0015] Step 3, reconstruct the image spectrally with super-resolution by combining the spatial spectrum features;
[0016] Step 4, training and optimizing the overall network formed by steps 1 to 3;
[0017] Step 5: Use the trained model to achieve spectral super-resolution reconstruction of the image.
[0018] Furthermore, in step 1, the shallow texture features of the image are extracted using a shallow feature extraction module and mapped to the C dimension at the same time; the shallow feature extraction module is a single 3×3 convolution F 3×3 (·) is used to obtain the shallow texture features of the input image and extract the shallow features F shallow The process is expressed as:
[0019] F shallow ∈R H×W×C =F 3×3 (x) (1)
[0020] Where x∈R H×W×3 It refers to the input RGB image, H×W is the resolution of the image, and 3 represents the number of channels of the image.
[0021] Furthermore, the specific processing process of the spatial scanning submodule is as follows:
[0022] The shallow texture feature map is expanded, and then the first layer normalization, the first linear layer processing, the deep convolution and SiLU activation are performed in sequence, and then the spatial scanning Mamba is entered, followed by the second layer normalization and the element multiplication operation with the output of the first layer normalization, and then the second linear layer processing and the addition operation with the output of the first linear layer are performed to obtain the final spatial feature;
[0023] The entire extracted spatial feature Fx ∈R H×W×C The process is expressed as:
[0024] F x = SSMamba1(Unfold(F shallow )) (2)
[0025] where Unfold(·) represents the unfolding operation that divides the feature map into local regions and serializes them, and SSMamba1(·) represents the spatial scanning sub-module used to extract spatial features. F shallow represents the shallow feature, H×W is the resolution of the image, and C represents the number of channels of the image.
[0026] Furthermore, the processing process of the spatial scanning Mamba is as follows: First, scan the first-band image patch of the input feature map, then scan the second image patch in row-major and column-major order respectively, and finish scanning all image patches; then scan the second band of the input feature map in the same way until all bands are scanned.
[0027] Furthermore, the specific processing process of the spectral scanning sub-module is as follows:
[0028] Perform the unfolding operation on the shallow texture feature map, then perform the first-layer normalization, the first linear layer processing, the depth convolution, and the SiLU activation in sequence, then enter the spectral scanning Mamba, then perform the second-layer normalization and perform the element-wise multiplication operation with the output of the first-layer normalization, and then perform the second linear layer processing and perform the addition operation with the output of the first linear layer to obtain the final spectral feature;
[0029] Extract the spectral feature F y ∈R H×W×C The process is expressed as:
[0030] F y = SSMamba2(Unfold(F shallow )) (3)
[0031] where Unfold(·) represents the unfolding operation that divides the feature map into local regions and serializes them, and SSMamba2(·) represents the spectral scanning sub-module used to extract spectral features. F shallow represents the shallow feature, H×W is the resolution of the image, and C represents the number of channels of the image.
[0032] Furthermore, the processing process of the spectral scanning Mamba is as follows: Perform the full-band scanning on the first image patch of the input feature map, and then scan the full band of the second image patch starting from the row and column respectively until the entire input feature map is scanned.
[0033] Furthermore, the empty-spectrum fusion sub-module includes two operations: spatial interaction X I and spectral interaction Y I Given two input features F x , F y ∈R H×W×C , spatial interaction graph X map ∈R H×W×1 and spectral interaction graph Y map ∈R 1×1×C , their calculations are as follows:
[0034] X Map ∈R H×W×1 = sigmoid(σ(F 1×1 (F x ))) (4)
[0035] Y Map ∈R 1×1×C = sigmoid(σ(F 1×1 (Pool(F y )))) (5)
[0036] where F 1×1 (·) represents a 1×1 convolution, Pool represents an adaptive average pooling operation, σ(·) represents the GELU method, sigmoid represents an activation function. Spectral interaction and spatial interaction are respectively applied to the spatial feature F x and the spectral feature F y . The two interaction operations are expressed as:
[0037] X I = X map ⊙F y (6)
[0038] Y I = Y map ⊙F x (7)
[0039] where ⊙ represents element-wise multiplication. The interacted features are then concatenated and fused to obtain the final deep empty-spectrum feature F deep ∈R H×W×C :
[0040] F deep ∈R H×W×C = X I + Y I (8)
[0041] where H×W is the resolution of the image, and C represents the number of channels of the image.
[0042] Further, in step 3, a spectral super-resolution reconstruction is performed on the image using a spectral reconstruction module; the spectral reconstruction module is a 3×3 convolution F without an activation function 3×3 (·), which multiplies with the finally obtained deep spatio-spectral feature F deep to obtain a hyperspectral image F with N dimensions out ∈R H×W×N :
[0043] F out ∈R H×W×N = F 3×3 (·)(F deep ) (9)
[0044] where H×W is the resolution of the image, and the input of F 3×3 (·) is the deep spatio-spectral feature F deep .
[0045] Further, during the training process of step 4, a cosine annealing strategy is adopted to dynamically adjust the learning rate, Adam is selected as the optimizer, and the mean absolute error is used as the loss function for training;
[0046] Step 5 also includes using 5 common evaluation metrics, including root mean square error, spectral angle mapper, peak signal-to-noise ratio, structural similarity index, and relative global dimension, to evaluate the final reconstruction effect.
[0047] The present invention also provides a spectral super-resolution reconstruction system based on a spatio-spectral Mamba network, including:
[0048] a processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a spectral super-resolution reconstruction method based on a spatio-spectral Mamba network as described in the above technical solution.
[0049] The advantages of the present invention are as follows:
[0050] 1. The spatio-spectral Mamba network disclosed in the present invention makes full use of the characteristics of the state space model architecture, can exhibit strong global feature extraction ability in processing hyperspectral images, and at the same time, due to its internal structure design, the network maintains a linear growth in computational complexity.
[0051] 2. The band information of hyperspectral images is essentially a natural information. The present invention realizes the collaborative modeling of the spatial dimension and the spectral dimension, as well as the adaptive interactive fusion of spatio-spectral information, by designing a specific scanning mechanism and fusion strategy.
[0052] 3. The present invention utilizes multiple public detection datasets to conduct an overall evaluation and analysis of the experimental results, ensuring the effectiveness of the spatial-spectral Mamba network. At the same time, the reconstructed image is compared with the real image to ensure good results for the reconstructed hyperspectral image. Description of the Drawings
[0053] Figure 1 is the overall technical roadmap of the present invention;
[0054] Figure 2 is the entire flowchart of the implementation process of the present invention;
[0055] Figure 3 is the spatial-spectral Mamba module diagram in the network model of the present invention;
[0056] Figure 4 is the overall process diagram of the present invention in the spatial scanning stage;
[0057] Figure 5 is the overall process diagram of the present invention in the spectral scanning stage;
[0058] Figure 6 is the overall process diagram of the present invention in the spatial-spectral fusion stage. Detailed Embodiment
[0059] The technical solution of the present invention will be further described below in conjunction with the drawings and embodiments. Referring to Figure 1 , a spectral super-resolution reconstruction method based on a spatial-spectral Mamba network provided by an embodiment of the present invention is implemented as follows:
[0060] Step 1, extract the shallow texture features of the image.
[0061] Let \(x\in R\) H×W×3 represent the input RGB image, where \(H\times W\) is the resolution of the image, and 3 represents the number of channels of the image. First, construct a shallow feature extraction module responsible for extracting the initial feature \(F\) shallow \(\in R\) H×W×C . Here, mainly use a single \(3\times3\) convolution \(F\) 3×3 (·) to obtain the low-frequency information such as the color and texture of the input image. At the same time, after the convolution operation, the dimension of the shallow feature is expanded to \(C\) channels. The process of extracting the shallow feature \(F\) shallow can be expressed as:
[0062] \(F\) shallow \(\in R\) H×W×C \(=F\) 3×3 (x) (1)
[0063] Step 2, mine the deep spatial-spectral features of the image.
[0064] The spatial-spectral Mamba module is asFigure 3 As shown, it is composed of two sub - modules: spatial scanning and spectral scanning. After the shallow feature map enters this module, it will be rearranged and then enter the spatial scanning sub - module and spectral scanning sub - module respectively to extract spatial - spectral features. In the two sub - modules, the input features will first go through the first layer of normalization operation, and then be divided into two paths. One branch will go through the first linear layer processing, depth convolution and SiLU activation, then enter the spatial scanning Mamba, then go through the second layer of normalization and perform an element - wise multiplication operation with the output of the first layer of normalization, and then go through the second linear layer processing and perform an addition operation with the output of the first linear layer to obtain the final spatial features; in the other branch, it first goes through the first linear layer processing, depth convolution and SiLU activation, then enters the spectral scanning Mamba, then goes through the second layer of normalization and performs an element - wise multiplication operation with the output of the first layer of normalization, and then goes through the second linear layer processing and performs an addition operation with the output of the first linear layer to obtain the final spectral features. Then the spatial features and spectral features will interact in the spatial - spectral fusion sub - module to fuse their respective important features.
[0065] (2a) Construct a spatial scanning sub - module to mine deeper information of hyperspectral images. Existing scanning techniques mainly focus on the spatial dimension of images, often ignoring the rich band information in the channel dimension. Moreover, the channel dimension of hyperspectral images contains a large amount of spectral information closely related to spatial information, and these spatial and spectral information are crucial for a comprehensive understanding of hyperspectral images.
[0066] The process of spatial scanning Mamba is as Figure 4 shown. First, extract the features on the first - band image patch of the input feature map, then extract the features of the second image patch in row - first and column - first manners respectively until all image patches are processed. Then, scan and extract features on the second band of the feature map in the same way until all bands are scanned. In this way, not only can spatial features be captured on each band, but also through the continuous processing of all bands, the spatial and spectral information of hyperspectral images can be comprehensively understood and utilized. The entire process of extracting spatial feature F x ∈R H×W×C can be expressed as:
[0067] F x =SSMamba1(Unfold(F shallow )) (2)
[0068] where Unfold(·) represents the unfolding operation, which divides the feature map into local regions and serializes them, and SSMamba1(.) represents the proposed spatial scanning sub - module for extracting spatial features.
[0069] (2b) Construct a spectral scanning sub - module, asFigure 5 As shown, different from spatial scanning, spectral scanning Mamba performs full-band scanning on the first image patch of the input feature map, and then scans the full band of the second image patch starting from rows and columns respectively until the entire input feature map is scanned. Spectral feature F is extracted y ∈R H×W×C The process is similar to spatial scanning:
[0070] F y = SSMamba2(Unfold(F shallow )) (3)
[0071] where Unfold(·) represents the unfolding operation, which divides the feature map into local regions and serializes them, and SSMamba2(·) represents the proposed spectral scanning sub-module for extracting spectral features.
[0072] (2c) Construct an empty-spectrum fusion sub-module for realizing the adaptive interaction and fusion of important features between spatial features and spectral features. The spatial scanning block focuses on capturing spatial features, while the spectral scanning block focuses on extracting spectral features. Simply adding these two parts of features cannot effectively couple spatial and spectral features. Specifically, this module includes spatial interaction X I and spectral interaction Y I in two steps, as Figure 6 shown. Given two input features F x , F y ∈R H×W×C , the spatial interaction graph X map ∈R H×W×1 and the spectral interaction graph Y map ∈R 1×1×C are calculated as:
[0073] X Map ∈R H×W×1 = sigmoid(σ(F 1×1 (F x ))) (4)
[0074] Y Map ∈R 1×1×C = sigmoid(σ(F 1×1 (Pool(F y )))) (5)
[0075] where F 1×1 (·) represents a 1×1 convolution, Pool represents the adaptive average pooling operation, σ(·) represents the GELU method, and sigmoid represents the activation function. Spectral interaction and spatial interaction are respectively applied to the spatial feature F x and the spectral feature F y The two interaction operations can be expressed as:
[0076] X I = X map ⊙ F y (6)
[0077] Y I = Y map ⊙ F x (7)
[0078] Where ⊙ represents element-wise multiplication, and the interacted features are then concatenated and fused to obtain the final deep feature F deep ∈ R H ×W×C :
[0079] F deep ∈ R H×W×C = X I + Y I (8)
[0080] Step 3: Implement spectral super-resolution reconstruction of the image
[0081] Construct a spectral reconstruction module for spectral super-resolution reconstruction of the image. The deep feature F after multiple feature extractions and fusions deep is used as the input. This module is designed as a 3×3 convolution F without an activation function 3×3 (·), and the input of F 3×3 (·) is the deep feature F deep , and multiplying it with the finally obtained deep feature gives a hyperspectral image F with N dimensions out ∈ R H×W×N :
[0082] F out ∈ R H×W×N = F 3×3 (·)(F deep ) (9)
[0083] Step 4: Network training and parameter optimization
[0084] In the training stage, the initial learning rate of the experiment is set to 0.0005. The cosine annealing strategy is used to dynamically adjust the learning rate. Adam is selected as the optimizer, and the mean absolute error is used as the loss function for training. During the training process, RGB and hyperspectral image samples with a size of 32×32 are loaded from the original dataset as the input. The hyperspectral image samples represent the real samples and are used to calculate the loss with the spectral reconstruction results. The batch size is set to 8, and a total of 100 epochs are trained
[0085] Step 5: Experimental testing
[0086] In the testing stage, the RGB image that requires spectral super-resolution reconstruction is input into the shallow feature extraction module to extract the low-frequency information of the image. Then, it enters multiple spatio-spectral deep feature extraction modules to mine the deep information of the image. Here, it will pass through the parallel spatial scanning and spectral scanning sub-modules. Then, the information from the two branches is adaptively interacted and important information is fused through the spatio-spectral fusion sub-module. Finally, the stitched feature map is input into the spectral reconstruction module to super-resolve the required hyperspectral image, and it is compared and evaluated with the real image to ensure that the reconstructed hyperspectral image has good results.
[0087] The following uses experimental data to illustrate the technical effects of the present invention.
[0088] 1. Experimental conditions: The present invention is carried out on a workstation with a 13th Gen i9-13900K CPU, 64G of memory, an RTX 4090 GPU, and an Ubuntu 22.04.2 LTS operating system.
[0089] 2. Datasets: Two hyperspectral datasets, CAVE and Harvard, are used. The CAVE dataset contains 32 images with a size of 512×512. It is a commonly used hyperspectral dataset. All hyperspectral images contain 31 bands, covering the spectral range of 400 - 700 nm. The Harvard dataset contains 50 indoor and outdoor hyperspectral images, with a spatial resolution of 1024×1392 and 31 spectral bands. These bands cover the spectral range from 420 - 720 nm. The images were captured under daylight illumination conditions using a commercial hyperspectral camera (Nuance FX).
[0090] 3. Evaluation metrics: To evaluate the spatio-spectral Mamba network of the present invention, 5 commonly used evaluation metrics are adopted, including Root Mean Square Error (RMSE), Spectral Angle Mapper (SAM), Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Relative Global Error in Synthesis (ERGAS). Tables 1 and 2 respectively show the quantitative results of the spectral super-resolution algorithm on the CAVE and Harvard datasets. It can be seen that the method of the present invention reaches the optimal level in all metrics.
[0091] 4. Comparison methods: The comparison methods adopted CanNet (CanY B, Timofte R. An efficient CNN for spectral reconstruction from RGB images[J]. arXiv preprint arXiv:1804.04647, 2018.), DenseUnet (Galliani S, Lanaras C, Marmanis D, et al. Learned spectral super-resolution[J]. arXiv preprint arXiv:1703.09470, 2017.), FMNet (Zhang L, Lang Z, Wang P, et al. Pixel-aware deep function-mixture network for spectral super-resolution[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2020, 34(07):12821-12828.), HSCNN+ (Shi Z, Chen C, Xiong Z, et al. Hscnn+: Advanced cnn-based hyperspectral recover.y from rgb images[C] / / Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. 2018:939-947.), HSRnet (He J, Li J, Yuan Q, et al. Spectral response function-guided deep optimization-driven network for spectral super-resolution[J]. IEEE Transactions on Neural Networks and Learning Systems, 2021, 33(9):4213-4227.), sRCNN (Gewali U B, Monteiro S T, Saber E. Spectra1 super-resolution with optimized bands[J]. Remote Sensing, 2019, 11(14):1648.) (Chen w, Zheng X, Lu X. Semisupervised spectral degradation constrained network for spectral super-resolution[J]. IEEE Geoscience and Remote Sensing Letters, 2021, 19: 1-5.) is used to verify the effectiveness of the method of the present invention.
[0092] Table 1 Quantitative results of different methods on the CAVE dataset
[0093]
[0094] Table 2 Quantitative results of different methods on the Harvard dataset
[0095]
[0096] On the other hand, an embodiment of the present invention further provides a spectral super-resolution reconstruction system based on an empty-spectrum Mamba network, including:
[0097] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a spectral super-resolution reconstruction method based on an empty-spectrum Mamba network as described in the above technical solution.
[0098] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A spectral super-resolution reconstruction method based on a spatial spectrum Mamba network, characterized in that: The steps include: Step 1, extracting shallow texture features of the image; Step 2, constructing a spatial-spectrum Mamba module, including a spatial scanning submodule for mining deeper spatial feature information of hyperspectral images, a spectral scanning submodule for extracting spectral features, and a spatial-spectrum fusion submodule for realizing adaptive interaction and fusion of spatial and spectral features; Step 3, reconstruct the image spectrally with super-resolution by combining the spatial spectrum features; Step 4, training and optimizing the empty spectrum Mamba network constructed in step 1 to step 3; Step 5: Use the trained model to achieve spectral super-resolution reconstruction of the image.
2. The spectral super-resolution reconstruction method based on the spatial spectrum Mamba network according to claim 1, characterized in that: In step 1, the shallow texture features of the image are extracted using the shallow feature extraction module and mapped to the C dimension at the same time; the shallow feature extraction module is a single 3×3 convolution F 3×3 (·) is used to obtain the shallow texture features of the input image and extract the shallow features F shallow The process is expressed as: F shallow ∈R H×W×C =F 3×3 (x) (1) Where x∈R H×W×3 It refers to the input RGB image, H×W is the resolution of the image, and 3 represents the number of channels of the image.
3. The spectral super-resolution reconstruction method based on the spatial spectrum Mamba network according to claim 1, characterized in that: The specific processing process of the spatial scanning submodule is as follows: The shallow texture feature map is expanded, and then the first layer normalization, the first linear layer processing, the deep convolution and SiLU activation are performed in sequence, and then the spatial scanning Mamba is entered, followed by the second layer normalization and the element multiplication operation with the output of the first layer normalization, and then the second linear layer processing and the addition operation with the output of the first linear layer are performed to obtain the final spatial feature; The entire extracted spatial feature F x ∈R H×W×C The process is expressed as: F x =SSMamba1(Unfold(F shallow )) (2) Unfold(·) represents the expansion operation, which divides the feature map into local regions and serializes them. SSMamba1(·) represents the spatial scanning submodule, which is used to extract spatial features. shallow represents shallow features, H×W is the resolution of the image, and C represents the number of channels of the image.
4. The spectral super-resolution reconstruction method based on the spatial spectrum Mamba network as claimed in claim 3, characterized in that: The processing process of spatial scanning Mamba is as follows: first scan the first band image block of the input feature map, then scan the second image block in row-first and column-first manner respectively, and all image blocks are executed; then scan the second band of the input feature map in the same way until all bands are scanned.
5. The spectral super-resolution reconstruction method based on the spatial spectrum Mamba network according to claim 1, characterized in that: The specific processing process of the spectrum scanning submodule is as follows: The shallow texture feature map is expanded, and then the first layer normalization, the first linear layer processing, deep convolution and SiLU activation are performed in sequence, and then the spectrum scanning Mamba is entered, followed by the second layer normalization and element-wise multiplication with the output of the first layer normalization, and then the second linear layer processing and the addition operation with the output of the first linear layer are performed to obtain the final spectral feature; Extract spectral feature F y ∈R H×W×C The process is expressed as: F y =SSMamba2(Unfold(F shallow )) (3) Unfold(·) represents the expansion operation, which divides the feature map into local regions and serializes them. SSMamba2(·) represents the spectrum scanning submodule, which is used to extract spectral features. shallow represents shallow features, H×W is the resolution of the image, and C represents the number of channels of the image.
6. The spectral super-resolution reconstruction method based on the spatial spectrum Mamba network according to claim 5, characterized in that: The processing process of spectral scanning Mamba is: perform full-band scanning on the first image block of the input feature map, and then scan the full-band of the second image block from the rows and columns respectively until the entire input feature map is scanned.
7. The spectral super-resolution reconstruction method based on the spatial spectrum Mamba network according to claim 1, characterized in that: The spatial-spectral fusion submodule includes spatial interaction X I and spectral interaction Y I Two-step operation, given two input features F x , F y ∈R H ×W×C , spatial interaction graph X map ∈R H×W×1 and spectral interaction graph Y map ∈R 1×1×C The calculation is: X Map ∈R H×W×1 =sigmoid(σ(F 1×1 (F x ))) (4) Y Map ∈R 1×1×C sigmoid(σ(F 1×1 (Pool(F y )))) (5) Among them, F 1×1 (·) represents a 1×1 convolution, Pool represents an adaptive average pooling operation, σ(·) represents the GELU method, sigmoid represents the activation function, and the spatial feature F x and spectral feature F y Apply spectral interaction and spatial interaction respectively, and the two interaction operations are expressed as: X I =X map ⊙F y (6) AND I =And map ⊙F x (7) Among them, ⊙ represents element-wise multiplication. After the interactive features are concatenated and fused, the final deep spatial spectrum feature F is obtained. deep ∈R H ×W×C : F deep ∈R H×W×C =X I +Y I (8) Among them, H×W is the resolution of the image, and C represents the number of channels of the image.
8. The spectral super-resolution reconstruction method based on the spatial spectrum Mamba network according to claim 1, characterized in that: In step 3, the spectral reconstruction module is used to perform spectral super-resolution reconstruction on the image; the spectral reconstruction module is a 3×3 convolution F without an activation function. 3×3 (·), and the deep spatial spectrum feature F deep Multiplying them together gives us a hyperspectral image F with N dimensions. out ∈R H×W×N : F out ∈R H×W×N =F 3×3 (·)(F deep ) (9) Where H×W is the resolution of the image, F 3×3 The input of (·) is the deep spatial spectrum feature F deep .
9. The spectral super-resolution reconstruction method based on the spatial spectrum Mamba network according to claim 1, characterized in that: In the training process of step 4, the cosine annealing strategy is used to dynamically adjust the learning rate, Adam is selected as the optimizer, and the mean absolute error is used as the loss function of the training; Step 5 also includes the use of five commonly used evaluation indicators, including root mean square error, spectral angle mapper, peak signal-to-noise ratio, structural similarity index and relative global dimension to evaluate the final reconstruction effect.
10. A spectral super-resolution reconstruction system based on a spatial spectrum Mamba network, characterized in that: include: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a spectral super-resolution reconstruction method based on a spatial spectrum Mamba network as described in any one of claims 1 to 9.