Spectral super-resolution reconstruction method based on bidirectional attention mechanism
Through the spectral super-resolution reconstruction method based on the bidirectional attention mechanism, the problems of high cost of hyperspectral image acquisition equipment and unstable data are solved, high-precision spectral reconstruction and spatial detail retention are achieved, and high-fidelity hyperspectral images are generated.
Patent Information
- Application Number
- CN202510553240.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing hyperspectral image acquisition equipment is costly and slow in acquisition speed, and the data acquisition process is susceptible to ambient light and noise, resulting in data instability. Traditional spectral super-resolution methods are difficult to balance spectral reconstruction and spatial detail retention, and lack an effective spectral-space interaction mechanism.
The spectral super-resolution reconstruction method based on the bidirectional attention mechanism is adopted. By establishing a spectral super-resolution data set, RGB images and hyperspectral images are preprocessed, shallow features are extracted and global information modeled, and the bidirectional spectral attention branches and independent spatial information extraction branches are used to perform multi-scale features interactive fusion, and combined with mixed loss functions are optimized.
It has achieved a significant improvement in spectral recovery accuracy, ensured full retention of image space information, improved the training stability and generalization ability of the network, and generated high-fidelity hyperspectral images.
Smart Images

Figure CN120471768A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and more specifically, to a spectral super-resolution reconstruction method based on a bidirectional attention mechanism. Background Art
[0002] Hyperspectral imagery (HSI), with its characteristic of recording continuous, narrowband spectral information at each pixel, is of great significance in applications such as material identification, target detection, environmental monitoring, agricultural remote sensing, and geological exploration. Hyperspectral imagery allows for detailed distinction of ground objects, enabling accurate classification and quantitative analysis. However, due to the technical and hardware cost limitations of hyperspectral sensors, their data acquisition process is often subject to the following constraints: Traditional hyperspectral imaging systems (such as push-broom or point-by-point scanning systems) require high-cost equipment and have slow acquisition speeds, making them unsuitable for rapid imaging in dynamic scenes; hyperspectral imaging equipment is sensitive to factors such as ambient lighting and sensor calibration, resulting in the actual collected data often containing noise and instability; and due to the high acquisition cost and time costs, hyperspectral data faces problems such as insufficient data volume or untimely updates when applied on a large scale.
[0003] To address the many difficulties of directly acquiring hyperspectral imagery, researchers have proposed spectral super-resolution (SSR) methods, which reconstruct high-dimensional hyperspectral images from low-dimensional RGB images. SSR technology mainly includes two categories of methods. Traditional SSR methods rely on mathematical tools such as sparse representation, dictionary learning, and spectral basis function reconstruction. For example, sparse dictionary-based methods represent hyperspectral data by constructing a dictionary of training samples, but such methods are often limited in representation and generalization performance. Methods based on maximum a posteriori probability and spectral component deviation functions also have difficulty capturing complex spectral features due to their simple models, resulting in poor restoration results. With the widespread application of deep neural networks, an increasing number of studies have adopted new architectures such as convolutional neural networks (CNNs) and Transformers to achieve SSR. Deep learning methods learn the mapping relationship between RGB and HSI through training on large amounts of data, achieving more accurate restoration results than traditional SSR. However, existing frameworks often struggle to balance spectral reconstruction and spatial detail preservation, and lack effective mechanisms to model bidirectional spectral-spatial interactions. These limitations lead to poor performance in situations where high-fidelity reconstruction of spectral correlations and fine-grained spatial structure is required. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of the existing technology and propose a spectral super-resolution reconstruction method based on a bidirectional attention mechanism.
[0005] First, a spectral super-resolution reconstruction method based on a bidirectional attention mechanism is provided, comprising:
[0006] Step 1: Establish a spectral super-resolution dataset, which includes corresponding hyperspectral images and RGB images;
[0007] Step 2: Preprocess the RGB image and the hyperspectral image, and extract the shallow features of the RGB image; then perform global information modeling based on the shallow features of the RGB image, extract the overall context information of the image, and generate global features;
[0008] Step 3: Invert the spectral dimension of the extracted features to obtain forward and reverse features and perform self-attention calculation on the spectral dimension to obtain the spectral information feature representation of the RGB image;
[0009] Step 4: Obtain spatial information feature representation of the RGB image;
[0010] Step 5: Perform multi-scale feature interactive fusion on the spectral information feature representation and the spatial information feature representation to obtain fusion features;
[0011] Step 6: Obtain a reconstructed hyperspectral image based on the fusion features.
[0012] Preferably, step 7 is to perform optimization through a hybrid loss function.
[0013] Preferably, in step 2, the preprocessing of the RGB image and the hyperspectral image includes:
[0014] Crop the RGB image to a fixed size and normalize the RGB image;
[0015] The corresponding hyperspectral images are normalized according to the number of channels in the dataset;
[0016] Random flipping, rotation, and random cropping methods are used to perform data enhancement on RGB images and hyperspectral images.
[0017] Preferably, step 3 includes:
[0018] Step 3.1: Perform spectral self-attention calculation on the forward token sequence;
[0019] Step 3.2: Arrange the token sequence in reverse order on the spectral dimension to obtain the reverse token sequence, and then perform self-attention calculation on the spectral dimension;
[0020] Step 3.3: Based on the self-attention calculation results, perform splicing, linear transformation, and position encoding to form the spectral information feature representation of the RGB image.
[0021] Preferably, step 4 includes:
[0022] Step 4.1: extract local features from the RGB image to obtain local detail information; the local detail information includes image edges and textures;
[0023] Step 4.2: Fuse features at all scales through U-Net-style downsampling and upsampling operations.
[0024] Step 4.3: Implement feature dimensionality reduction and information complementation through cross-layer splicing operations and convolution operations to form a complete spatial information feature representation.
[0025] Preferably, step 6 includes:
[0026] Step 6.1, performing global information modeling based on the fusion features to capture the overall spectral trend of the image;
[0027] Step 6.2: Adaptively weight the channel and spatial information to highlight key features and suppress redundant information.
[0028] Step 6.3: Calculate the hyperspectral image after spectral super-resolution and add it element-by-element with the shallow features to obtain the reconstructed hyperspectral image.
[0029] In a second aspect, a spectral super-resolution reconstruction system based on a bidirectional attention mechanism is provided, which is used to perform any of the methods described in the first aspect, including:
[0030] An establishment module is used to establish a spectral super-resolution dataset, wherein the spectral super-resolution dataset includes corresponding hyperspectral images and RGB images;
[0031] A preprocessing module is used to preprocess RGB images and hyperspectral images and extract shallow features of the RGB images; then perform global information modeling based on the shallow features of the RGB images, extract the overall context information of the images, and generate global features;
[0032] The calculation module is used to invert the spectral dimension of the extracted features, obtain the forward and reverse features, and perform self-attention calculation on the spectral dimension to obtain the spectral information feature representation of the RGB image;
[0033] A first acquisition module is used to obtain the spatial information feature representation of the RGB image;
[0034] A fusion module, configured to perform multi-scale feature interactive fusion on the spectral information feature representation and the spatial information feature representation to obtain a fusion feature;
[0035] The second acquisition module is used to acquire the reconstructed hyperspectral image according to the fusion feature.
[0036] According to a third aspect, a computer storage medium is provided, wherein a computer program is stored in the computer storage medium; when the computer program is executed on a computer, the computer executes any one of the methods described in the first aspect.
[0037] In a fourth aspect, an electronic device is provided, including:
[0038] Memory, used to store computer programs;
[0039] A processor is used to execute the computer program to implement any method as described in the first aspect.
[0040] The beneficial effects of the present invention are:
[0041] 1. The spectral super-resolution framework constructed in this paper significantly improves spectral restoration accuracy by adopting a bidirectional spectral attention mechanism and an independent spatial information extraction branch. This mechanism enables the network to capture spectral dependencies from both positive and negative directions, overcoming the information loss problem caused by traditional unidirectional modeling, thereby achieving more accurate spectral reconstruction. Simultaneously, the independently designed MSAB branch utilizes the Transformer architecture and multi-scale interaction mechanism to efficiently extract and restore spatial details such as edges and textures in the image, ensuring the full preservation of the image's spatial information.
[0042] 2. This invention also employs a multi-scale feature fusion strategy to achieve the coordinated optimization of global and local information, resulting in a high level of overall structure and local detail in the reconstructed hyperspectral image. Furthermore, by introducing a hybrid loss function and an end-to-end optimization strategy, the hyperspectral reconstruction error is significantly reduced while preserving the fidelity of the image's spatial information, thereby significantly improving the overall training stability and generalization capability of the network. These technical advantages collectively constitute a significant innovation in the field of spectral super-resolution. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 Flowchart of the spectral super-resolution reconstruction method based on the bidirectional attention mechanism and the flowchart of each module;
[0044] Figure 2 A comparison chart of experimental results on the ICVL dataset;
[0045] Figure 3 This is a comparison chart of experimental results on the DFC2018 Houston dataset;
[0046] Figure 4 A comparison chart of experimental results on the TG1HRSSC dataset. DETAILED DESCRIPTION
[0047] The present invention will be further described below with reference to the following examples. The following examples are provided only to facilitate understanding of the present invention. It should be noted that, without departing from the principles of the present invention, it is possible for a person skilled in the art to make various modifications to the present invention, and such improvements and modifications fall within the scope of the claims of the present invention.
[0048] Example 1:
[0049] To achieve high-fidelity spectral super-resolution of remote sensing images, this application provides a spectral super-resolution reconstruction method based on a bidirectional attention mechanism. First, RGB and hyperspectral images are preprocessed and data augmented. Second, shallow feature extraction is performed on the RGB images. Next, the spectral-spatial fusion interaction module (SSFIM) achieves coordinated optimization of spectral and spatial information. Finally, a designed reconstruction module adaptively fuses bidirectional spectral information to produce the final hyperspectral image.
[0050] Specifically, such as Figure 1 As shown, the method includes:
[0051] Step 1: Establish a spectral super-resolution dataset, which includes corresponding hyperspectral images and RGB images.
[0052] Step 2: Preprocess the RGB image and the hyperspectral image, and extract the shallow features of the RGB image; then perform global information modeling based on the shallow features of the RGB image, extract the overall context information of the image, and generate global features.
[0053] Specifically, preprocessing involves cropping the input RGB images to a fixed size (e.g., 256×256) and normalizing them to scale pixel values to the range [0, 1]. Furthermore, the corresponding hyperspectral images (HSI) are normalized based on the number of channels specified for each dataset (e.g., 31, 50, or 54) to ensure consistent scale across the data. To prevent overfitting during training, data augmentation is performed on the RGB images and HSI using random flipping, rotation, and random cropping. This data augmentation strategy not only expands the training set but also improves the model's robustness to noise and environmental variations.
[0054] Extracting the shallow features of the RGB image includes: the pre-processed RGB image first enters the shallow feature extraction module, such as Figure 1 (a), The main task of this module is to map low-dimensional RGB data to high-dimensional feature space, laying the foundation for the spectral and spatial information extraction of subsequent modules.
[0055] In addition, a convolution operation with a kernel of 1×1 is applied to the input RGB image X to map the original 3 channels to a higher-dimensional feature representation:
[0056] F s =CELU(Conv 1×1 (X)
[0057] Among them, Conv 1×1 It represents a convolution operation with a convolution kernel of 1×1, and GELU represents the GELU activation function.
[0058] The Vision Mamba module is used to perform global information modeling on the shallow features output by the 1×1 convolution, extract the overall context information of the image, and generate the global feature F Vim :
[0059] F Vim =VisionMamba(F s )
[0060] This module provides rich context for subsequent spectral and spatial branches by capturing the global correlation of the image.
[0061] Step 3: Perform self-attention calculation on the spectral dimension of the forward and reverse token sequences to obtain the spectral information feature representation of the RGB image.
[0062] In step 3, the extracted features are spectrally inverted to obtain forward and reverse features (i.e., forward and reverse token sequences). In order to fully capture the spectral dependency information hidden in the RGB image, this embodiment constructs a bidirectional spectral attention branch, such as Figure 1 (b) contains the forward branch (MSEB) and the reverse branch (RMSEB). The specific implementation steps are as follows:
[0063] Step 3.1: The forward branch (MSEB) directly performs spectral self-attention calculation on the token sequence.
[0064] In step 3.2, the reverse branch (RMSEB) first arranges the token sequence in reverse order in the spectral dimension and then performs the same self-attention calculation.
[0065] The specific implementation steps of step 3.1 and step 3.2 are as follows:
[0066] Q=T·W Q ,K=T·W K ,V=T·W V
[0067] Among them, W Q 、W K and W Vis a trainable parameter matrix and T is a token sequence.
[0068] Step 3.3: Based on the self-attention calculation results, perform splicing, linear transformation, and position encoding to form the spectral information feature representation of the RGB image.
[0069] Specifically, in each attention head, by calculating the attention weight matrix:
[0070]
[0071]
[0072]
[0073] And for the value V i Perform weighted summation to generate the output of each head, and then form the final spectral attention output S-MSA (X token ).
[0074] In step 3, the two branches share weights to ensure that the positive and negative features complement each other, making up for the shortcomings of unidirectional modeling. Through U-Net-style downsampling and upsampling operations, features at different scales are fused.
[0075] Step 4: Obtain the spatial information feature representation of the RGB image.
[0076] Step 4 includes:
[0077] Step 4.1: Extract local features from the RGB image to obtain local detail information.
[0078] Specifically, in order to ensure that the spatial details of the image are not ignored during the spectral super-resolution process, this embodiment designs an independent spatial information extraction branch MSAB, which uses multiple 3×3 convolutional layers, ReLU activation functions and maximum pooling layers to extract local features of RGB images and obtain local detail information such as image edges and textures.
[0079] Step 4.2: Fuse features of different scales through U-Net-style downsampling and upsampling operations.
[0080] In step 4.3, cross-layer concatenation (Cat operation) and 1×1 convolution further achieve feature dimensionality reduction and information complementation, forming a complete spatial information feature representation. The specific process is similar to step 3 and will not be described here.
[0081] Step 5: Perform multi-scale feature interactive fusion on the spectral information feature representation and the spatial information feature representation to obtain fusion features.
[0082] Specifically, step 5 includes:
[0083] In step 5.1, in the U-Net structure, the features from MSEB, RMSEB, and MSAB are downsampled layer by layer to extract low-level details, and then high-resolution information is restored through upsampling to ensure that both global structure and local details are expressed.
[0084] In step 5.2, at the corresponding downsampling and upsampling stages, 1×1 convolution and concatenation (Cat operation) are used to achieve fusion of features of different scales. The specific fusion process can be referred to the following formula:
[0085]
[0086]
[0087]
[0088]
[0089] Among them, SEB represents the spectral attention module, D represents downsampling of the feature map, U represents upsampling of the feature map, Cat represents the concatenation operation of the input features in the channel dimension, Conv 1×1 Indicates a convolution operation with a convolution kernel of 1×1. and Represents the feature map generated after the nth SEB and SAB; This is the feature map of the 5th SEB after the inverted spectral attention module branch; and is the intermediate feature; F SSFIM It is the feature map output by SSFIM.
[0090] Step 6: Based on the fusion features, the final hyperspectral image is generated after further processing in the reconstruction module.
[0091] Step 6 includes:
[0092] Step 6.1: Input the fused features into the reconstruction module. First, use the Vision Mamba module to perform global information modeling to capture the overall spectral trend of the image.
[0093] Step 6.2: Introduce the Convolutional Block Attention Module (CBAM) module to adaptively weight channel and spatial information, highlight key features, and suppress redundant information.
[0094] Step 6.3: Calculate the hyperspectral image after spectral super-resolution and add it element-by-element with the shallow features to obtain the reconstructed hyperspectral image.
[0095] Specifically, the hyperspectral image output by the reconstruction module is recorded as
[0096]
[0097] in, Represents a hyperspectral image after spectral super-resolution.
[0098] To further enhance the details, the shallow features F s and Perform element-by-element addition to obtain the final hyperspectral image Y:
[0099]
[0100] Example 2:
[0101] Based on Example 1, Example 2 of the present application provides a more specific spectral super-resolution reconstruction method based on a bidirectional attention mechanism, including:
[0102] Step 1: Establish a spectral super-resolution dataset, which includes corresponding hyperspectral images and RGB images.
[0103] Specifically, the basic data sets used in Example 2 of the present application are four existing publicly available data sets: the Salinas data set is an image of the Salinas Valley in California, USA. The corrected image contains 204 channels, covering the spectral range of 0.4-2.5μm. The ICVL data set consists of 201 images, each of which contains 31 spectral channels. We randomly selected 184 images as a training set and 6 images as a test set. Before the experiment, each image was cropped to (31,256,256). The DFC2018 Houston data set contains 50 spectral channels. In this experiment, 92 images were randomly selected as a training set and 6 images were used as a test set. The TG1HRSSC data set consists of three types of data: panchromatic images, visible and near-infrared images of 54 effective bands, and short-wave infrared images of 52 effective bands. In this experiment, we selected visible and near-infrared images, randomly selected 40 images as a training set, and 6 images as a test set. As an embodiment, the present invention uses computer software Pycharm and Pytorch framework based on NVIDIA GeForce RTX 3090 GPU to implement automatic operation process.
[0104] Step 2: Preprocess the RGB image and the hyperspectral image, and extract the shallow features of the RGB image; then perform global information modeling based on the shallow features of the RGB image, extract the overall context information of the image, and generate global features.
[0105] For example, the input RGB image is cropped to a fixed size (e.g., 256×256) and normalized, scaling pixel values to the range [0, 1]. Simultaneously, the corresponding hyperspectral image (HSI) is normalized based on the specific number of channels in each dataset (e.g., 31, 50, or 54) to ensure consistent scale across the data. To prevent overfitting due to insufficient sample size during training, data augmentation is performed on the RGB image and HSI using random flipping, rotation, and random cropping. This data augmentation strategy not only expands the training set but also improves the model's robustness to noise and environmental variations. The preprocessed data serves as both the RGB input and the HSI annotation data during training, providing a high-quality data foundation for subsequent network training and feature extraction. Furthermore, a 1×1 convolution operation is applied to the input RGB image X, mapping the original three channels to a higher-dimensional feature representation. The Vision Mamba module is used to perform global information modeling on the shallow features output by the 1×1 convolution, extracting the overall context of the image and generating global features.
[0106] Step 3: A bidirectional spectral attention branch was constructed. The forward branch (MSEB) directly processes the token sequence in the spectral dimension; the reverse branch (RMSEB) first reverses the spectral order of the token sequence before performing the same self-attention calculation. The two branches share weights to ensure that the positive and negative features complement each other, compensating for the shortcomings of unidirectional modeling. Features at different scales are fused through U-Net-style downsampling and upsampling operations.
[0107] Step 4: To ensure that spatial details of the image are not overlooked during spectral super-resolution, this embodiment designs an independent spatial information extraction branch (MSAB). This branch uses multiple convolutional layers and ReLU activation functions to extract local features from the RGB image, obtaining local details such as image edges and textures. Features at various scales are fused through U-Net-style downsampling and upsampling operations. Cross-layer concatenation (Cat operation) and 1×1 convolution further achieve feature dimensionality reduction and information complementarity, forming a complete spatial information feature representation.
[0108] Step 5: The spatial information extracted by MSAB interacts with the spectral features extracted by MSEB and RMSEB to achieve coordinated recovery of spectral and spatial information.
[0109] Step 6: Obtain a reconstructed hyperspectral image based on the fusion features.
[0110] Step 7: Optimize using a hybrid loss function.
[0111] During training, the model generates high-fidelity hyperspectral images through the combined constraints of multiple loss functions. The reconstruction loss is used to ensure that the generated hyperspectral and RGB images are highly consistent with the original hyperspectral and RGB images in terms of overall content, ensuring that the reconstructed images maintain realism and detail fidelity in the visible light range.
[0112] Specifically, to ensure that the network achieves optimal performance in the spectral and spatial reconstruction tasks, this embodiment adopts the following training and optimization strategies:
[0113]
[0114] L Finall =αL HSI +βL RGB
[0115] Among them, L HSI and L RGB They are the losses for reconstructing HSI and reconstructing RGB images, and both are constrained by L1 loss. α and β represent adjustable parameters. L Finally Represents the total loss function.
[0116] In addition, Example 2 of the present application experimentally verified the above method:
[0117] exist Figure 2 、 Figure 3 and Figure 4 The experimental results on the ICVL, DFC2018 Houston and TG1HRSSC datasets are shown in the first and third rows, respectively, where the pseudo-color images generated by the HSI are shown, and the second and fourth rows show the error maps between the images generated by the network and the labels. In the error maps, bluer colors indicate smaller differences from the true situation. The spectral super-resolution images of the proposed method are compared with those of six state-of-the-art methods. Previous methods have limitations in recovering spectral and spatial details, resulting in artifacts or spots, and are unable to accurately reconstruct image edges. In addition, their overall restoration quality is insufficient, lacking global consistency and continuity of space and spectrum. In contrast, the method proposed in the present invention overcomes these shortcomings and achieves high-quality super-resolution with consistent spectrum and spatial continuity, benefiting from more accurate spectral priors and effective supplementation of spatial information.
[0118] It should be noted that the parts in this embodiment that are the same or similar to those in Example 1 can be referenced to each other and will not be described in detail in this application.
[0119] Example 3:
[0120] Based on Example 3, Example 3 of the present application provides a spectral super-resolution reconstruction system based on a bidirectional attention mechanism, including:
[0121] An establishment module is used to establish a spectral super-resolution dataset, wherein the spectral super-resolution dataset includes corresponding hyperspectral images and RGB images;
[0122] A preprocessing module is used to preprocess RGB images and hyperspectral images and extract shallow features of the RGB images; then perform global information modeling based on the shallow features of the RGB images, extract the overall context information of the images, and generate global features;
[0123] The calculation module is used to perform self-attention calculations in the spectral dimension on the forward and reverse token sequences to obtain the spectral information feature representation of the RGB image;
[0124] A first acquisition module is used to obtain the spatial information feature representation of the RGB image;
[0125] A fusion module, configured to perform multi-scale feature interactive fusion on the spectral information feature representation and the spatial information feature representation to obtain a fusion feature;
[0126] The second acquisition module is used to acquire the reconstructed hyperspectral image according to the fusion feature.
[0127] Specifically, the system provided in this embodiment is a system corresponding to the method provided in Example 2. Therefore, the parts in this embodiment that are the same or similar to those in Example 2 can be referenced to each other and will not be repeated in this application.
Claims
1. A spectral super-resolution reconstruction method based on a bidirectional attention mechanism, characterized in that: include: Step 1: Establish a spectral super-resolution dataset, which includes corresponding hyperspectral images and RGB images; Step 2: Preprocess the RGB image and the hyperspectral image, and extract the shallow features of the RGB image; then perform global information modeling based on the shallow features of the RGB image, extract the overall context information of the image, and generate global features; Step 3: Invert the spectral dimension of the extracted features to obtain forward and reverse features and perform self-attention calculation on the spectral dimension to obtain the spectral information feature representation of the RGB image; Step 4: Obtain spatial information feature representation of the RGB image; Step 5: Perform multi-scale feature interactive fusion on the spectral information feature representation and the spatial information feature representation to obtain fusion features; Step 6: Obtain a reconstructed hyperspectral image based on the fusion features.
2. The spectral super-resolution reconstruction method based on the bidirectional attention mechanism according to claim 1 is characterized in that Also includes: Step 7: Optimize using a hybrid loss function.
3. The spectral super-resolution reconstruction method based on the bidirectional attention mechanism according to claim 1 or 2, characterized in that In step 2, the RGB image and the hyperspectral image are preprocessed, including: Crop the RGB image to a fixed size and normalize the RGB image; The corresponding hyperspectral images are normalized according to the number of channels in the dataset; Random flipping, rotation, and random cropping methods are used to perform data enhancement on RGB images and hyperspectral images.
4. The spectral super-resolution reconstruction method based on the bidirectional attention mechanism according to claim 3 is characterized in that Step 3 includes: Step 3.1: Perform spectral self-attention calculation on the forward token sequence; Step 3.2: Arrange the token sequence in reverse order on the spectral dimension to obtain the reverse token sequence, and then perform self-attention calculation on the spectral dimension; Step 3.3: Based on the self-attention calculation results, perform splicing, linear transformation, and position encoding to form the spectral information feature representation of the RGB image.
5. The spectral super-resolution reconstruction method based on the bidirectional attention mechanism according to claim 4 is characterized in that Step 4 includes: Step 4.1: extract local features from the RGB image to obtain local detail information; the local detail information includes image edges and textures; Step 4.2: Fuse features at all scales through U-Net-style downsampling and upsampling operations. Step 4.3: Implement feature dimensionality reduction and information complementation through cross-layer splicing operations and convolution operations to form a complete spatial information feature representation.
6. The spectral super-resolution reconstruction method based on the bidirectional attention mechanism according to claim 5 is characterized in that Step 6 includes: Step 6.1, performing global information modeling based on the fusion features to capture the overall spectral trend of the image; Step 6.2: Adaptively weight the channel and spatial information to highlight key features and suppress redundant information. Step 6.3: Calculate the hyperspectral image after spectral super-resolution and add it element-by-element with the shallow features to obtain the reconstructed hyperspectral image.
7. A spectral super-resolution reconstruction system based on a bidirectional attention mechanism, characterized in that: Used to perform the method according to any one of claims 1 to 6, comprising: An establishment module is used to establish a spectral super-resolution dataset, wherein the spectral super-resolution dataset includes corresponding hyperspectral images and RGB images; A preprocessing module is used to preprocess RGB images and hyperspectral images and extract shallow features of the RGB images; then perform global information modeling based on the shallow features of the RGB images, extract the overall context information of the images, and generate global features; The calculation module inverts the spectral dimension of the extracted features to obtain the forward and reverse features and performs self-attention calculation on the spectral dimension to obtain the spectral information feature representation of the RGB image; A first acquisition module is used to obtain the spatial information feature representation of the RGB image; A fusion module, configured to perform multi-scale feature interactive fusion on the spectral information feature representation and the spatial information feature representation to obtain a fusion feature; The second acquisition module is used to acquire the reconstructed hyperspectral image according to the fusion feature.
8. A computer storage medium, characterized in that The computer storage medium stores a computer program; when the computer program is run on a computer, the computer executes the method according to any one of claims 1 to 6.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Video super-resolution reconstruction method and system based on multi-scale local self-attention
CN115082308A
Hyperspectral image super-resolution reconstruction method based on multi-scale space-spectrum feature learning
CN115272078A
Hyperspectral image reconstruction method based on dual-path fusion
CN116612004A
Hyperspectral image super-resolution method and system based on natural image prior
CN117252757A
Depth spectrum super-resolution method based on spectrum and texture attention fusion
CN117437123A