Hyperspectral reconstruction method based on a single rgb image

By using a hyperspectral reconstruction method based on a single RGB image and employing a cone-shaped multi-scale feature extraction module and a pixel self-attention module, the problem of low image quality in existing methods is solved, and higher-precision hyperspectral image generation is achieved.

CN119399021BActive Publication Date: 2025-11-28UNIV OF SHANGHAI FOR SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310933882.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-27
Publication Date
2025-11-28
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

In the field of image technology, existing hyperspectral reconstruction methods, particularly machine learning-based methods, suffer from problems such as low image quality, large training data requirements, long processing time, and insufficient exploration of local regional information and long-range dependencies between pixels.

Method used

A hyperspectral reconstruction method based on a single RGB image is adopted. It combines a cone-shaped multi-scale feature extraction module, a multi-scale adaptive residual attention module, and a combined processing module. It uses traditional methods such as interpolation, principal component analysis, and pseudo-inverse method. The cone-shaped multi-scale feature extraction module performs feature mapping, the multi-scale adaptive residual attention module performs feature processing, and the pixel self-attention module captures the relationship between pixels, and finally generates a hyperspectral image.

Benefits of technology

It improves the reconstruction accuracy of hyperspectral images, enhances the ability to model the relationships between different regions of the image, and generates higher quality hyperspectral images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399021B_ABST
    Figure CN119399021B_ABST
Patent Text Reader

Abstract

The application provides a hyperspectral reconstruction method based on a single RGB image, and has the following characteristics: a conical multi-scale feature extraction module is used to extract features of a feature map of a single RGB image after slice partition processing, to obtain shallow region aggregation features; a plurality of sequentially connected multi-scale adaptive residual attention modules are used to process the shallow region aggregation features, to obtain deep features; a combination processing module is used to perform convolution and activation function processing on the deep features, to obtain a hyperspectral image, wherein the multi-scale adaptive residual attention module comprises a conical multi-scale feature extraction module, an optimal non-local module, a pixel self-attention module, a LayerNorm module and a multi-layer perception module connected in sequence. In summary, the method can improve the accuracy of the reconstructed hyperspectral image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a hyperspectral reconstruction method, specifically a hyperspectral reconstruction method based on a single RGB image. Background Technology

[0002] Hyperspectral (HS) imaging technology records the reflectance or transmittance of an object, combining imaging and spectroscopic techniques to detect the two-dimensional geometric space and one-dimensional spectral information of the target, thereby obtaining a continuous narrowband image with high spectral resolution. Hyperspectral images (HSI) acquired using this technology typically possess multiple spectral bands from infrared to ultraviolet. The rich spectral features have been widely used in remote sensing fields, such as geological exploration, marine environmental monitoring, ship detection, military reconnaissance, ecological research, and crop growth. However, in recent years, the development of HS imaging has encountered bottlenecks because the manufacturing of existing hyperspectral imaging equipment involves optics, mechanics, and computer science, resulting in high integration complexity and large device size. Furthermore, due to limitations in imaging technology, acquiring such high spatiotemporal resolution hyperspectral images containing rich spectral information is time-consuming, inevitably limiting the application scope of hyperspectral imaging technology, especially in portable platforms and high-speed mobile scenarios. One approach to address this problem is to generate such hyperspectral images by recovering lost spectral information from a given RGB image; this method can be called RGB image-based hyperspectral reconstruction. It refers to the inverse process of hyperspectral imaging by discovering an inverse response function. However, since the number of HSIs can be projected onto any RGB input, there may be multiple reasonable combinations of HSIs for the same RGB image. How to reconstruct the optimal HSI has become the core and hot topic of research.

[0003] To address this problem, numerous spectral reconstruction (SR) methods have been proposed, which can be broadly categorized into traditional methods, machine methods, and deep learning methods.

[0004] Existing traditional methods include interpolation techniques, principal component analysis, pseudo-inverse methods, Wiener's method, etc., but these traditional methods have poor expressive power and limited generalization ability, thus limiting their good performance on images.

[0005] Existing machine learning-based spectral reconstruction methods mainly focus on building sparse coding or relatively shallow learning models from specific hyperspectral data. These methods still suffer from problems such as susceptibility to external factors, large training data requirements, long training time, and the need to improve the quality of reconstructed images.

[0006] Existing deep learning methods have achieved good performance, but the accuracy of reconstructed hyperspectral images still needs further improvement. Furthermore, to obtain more advanced feature representations, most methods focus on designing deeper network structures, but lack exploration of information between local regions and long-range dependencies between pixels, limiting the network's ability to follow details and its discriminative learning.

[0007] In conclusion, there is still room for improvement in the quality of generated images using existing spectral reconstruction methods. Summary of the Invention

[0008] This invention is made to solve the above-mentioned problems, and aims to provide a hyperspectral reconstruction method based on a single RGB image.

[0009] This invention provides a hyperspectral reconstruction method based on a single RGB image. The method inputs a single RGB image into a hyperspectral reconstruction model to obtain the corresponding hyperspectral image. The hyperspectral reconstruction model includes: a cone-shaped multi-scale feature extraction module for extracting features from the feature map of the single RGB image after slice partitioning, obtaining shallow region aggregated features; multiple sequentially connected multi-scale adaptive residual attention modules for processing the shallow region aggregated features to obtain deep features; and a combined processing module for performing convolution and activation function processing on the deep features to obtain the hyperspectral image. The multi-scale adaptive residual attention module includes a cone-shaped multi-scale feature extraction module, an optimal nonlocal module, a pixel self-attention module, a LayerNorm module, and a multilayer perceptron module, all connected sequentially. The cone-shaped multi-scale feature extraction module uses the feature map of the single RGB image after slice partitioning or the input of the multi-scale adaptive residual attention module containing the cone-shaped multi-scale feature extraction module as the input feature F. in and the input features F in The feature F is obtained by sequentially performing multi-scale detailed feature extraction and multi-layer concatenation. Left Then, for feature F Left The shallow region aggregate features are obtained by sequentially extracting multi-scale detailed features and stitching multiple layers together, or the input of the optimal non-local module connected to the cone-shaped multi-scale feature extraction module is used as the output feature F. out The optimal nonlocal module is used to convert the output feature F out The feature sequence T is obtained by dividing the features into four equal-sized regions and then enhancing the connections between the features in these regions. i The pixel self-attention module is used to process the feature sequence T. i The long-range dependencies between pixels are captured to obtain the fused representation T. out The LayerNorm module is used to eliminate the fused representation T out and feature F LeftThe adverse effects caused by outlier data in the summation result are addressed by using the summation result as a feature. The multilayer perceptron module is used to perform feature transformation and information reorganization on the output of the LayerNorm module to obtain features. In each multi-scale adaptive residual attention module, the input feature F of the multi-scale adaptive residual attention module is... in ,feature and characteristics The connection results are obtained by performing the connection, and are used as the input of the next multi-scale adaptive residual attention module that is sequentially connected to the first multi-scale adaptive residual attention module. The connection result of the last multi-scale adaptive residual attention module is the deep feature.

[0010] The hyperspectral reconstruction method based on a single RGB image provided by this invention may also have the following features: the cone-shaped multi-scale feature extraction module includes a Left sub-module and a Right sub-module connected in sequence. The Left sub-module includes four Left channel domains and four corresponding convolutional layers of different scales. The channels of the four Left channel domains are all input features F of the cone-shaped multi-scale feature extraction module. in A quarter of the channel, the Nth left The kernel size of the convolution corresponding to the left channel domain of the layer is k1 = 2N. left +1, input feature F in Input the Left submodule to obtain feature F Left The formula for expressing it is: F Left =Relu(BN(F') out In the formula, ∑cat represents a multi-level concatenation operation. For the Nth left The i-th feature map of the Left channel domain, w k1,i For the i-th convolutional kernel with kernel size k1, the N-th... left The number of channels in the Left channel domain of the layer is one-quarter of the original input feature F in Both the feature maps and their corresponding convolutional kernels are divided into m groups. Batch normalization (BN) is used for batch normalization, and ReLU is used for activation. The Right submodule includes four Right channel domains and four corresponding convolutional layers at different scales. The channels of the four Right channel domains are all input features F of the input cone-shaped multi-scale feature extraction module. in One-quarter of the channel, the four Right channel domains correspond to the four Left channel domains in sequence, and the Nth right The kernel size of the convolution corresponding to the right channel domain of the layer is k² = 11-2N. right , feature F LeftThe input Right submodule yields the output feature F. out The formula for expressing it is: F out =Relu(BN(F”) out ), where ∑cat represents multi-level concatenation operations, G i,j For the Nth right The i-th feature map of the layer channel domain, w k2,i For the i-th convolutional kernel with kernel size k2, the N-th... right The feature F in the layer channel domain has one-quarter of the original number of channels. Left The feature maps and their corresponding convolutional kernels are divided into m groups, BN is the batch normalization layer, and ReLU is the activation function.

[0011] The hyperspectral reconstruction method based on a single RGB image provided by this invention may also have the following feature: wherein the output feature F out The feature sequence T is obtained by inputting the optimal nonlocal module. i The formula for expressing T is: i =ONB(F out {lu,ld,ru,rd}), where ONB is the optimal nonlocal module, F out {lu,ld,ru,rd} represents the feature F out It is divided into four equal-sized areas: upper left, lower left, upper right, and lower right.

[0012] The hyperspectral reconstruction method based on a single RGB image provided by this invention may also have the following feature: wherein the pixel self-attention module focuses on the feature sequence T i The fused representation T is obtained through processing. out The specific steps are as follows: Step S1, for feature T i A combination of convolution-activation function-convolution is used to obtain preliminary features, which are then divided into Q, K, and V for the attention mechanism; step S2, the weight of each pixel region is calculated based on Q and K; step S3, based on the feature sequence T... i The fused representation T is obtained by calculating the weights and V. out .

[0013] The hyperspectral reconstruction method based on a single RGB image provided by this invention may also have the following feature: In step S1, the formulas for Q, K, and V are as follows: Q = Conv(ρ(Conv(T)) i K = Conv(ρ(Conv(T)) i V = Conv(ρ(Conv(T)) iIn the formula, ρ(·) is the activation function Prelu and Conv is the convolution operation.

[0014] The hyperspectral reconstruction method based on a single RGB image provided by this invention may also have the following feature: wherein, in step S2, the weight formula is as follows: U i =Corr(Q i ,Conv(K i )), μ i =softmax(U i ), where U i Let Q be the attention value for the i-th pixel region, and Corr be the value calculated from the correlation matrix. i Let Q and K be the values ​​of the i-th pixel region. i Let K be the region of the i-th pixel, softmax be the softmax function operation, and μ be the value of K. i Let be the weight of the i-th pixel region.

[0015] The hyperspectral reconstruction method based on a single RGB image provided by this invention may also have the following feature: wherein, in step S3, the fusion representation T out The formula is as follows: T o ={T o1 ,T o2 ,T o3 ,...T oi ,...T on}, T out =T i +δ(Conv(T o )), where T oi This is the output for the i-th pixel region. For transpose multiplication, V i Let V and T be the region of the i-th pixel. o The set of outputs for all pixel regions, where δ is a learnable parameter.

[0016] The role and effect of invention

[0017] The hyperspectral reconstruction method based on a single RGB image according to the present invention improves the accuracy of the reconstructed hyperspectral image by, on the one hand, interleaving the extraction of detailed features at different scales across four different channels using a cone-shaped multi-scale feature extraction module 10, thereby obtaining aggregated local information and relatively global information including edge spatial information, ultimately achieving the goal of extracting inter-regional contextual features at multiple scales; on the other hand, the pixel self-attention module 202 extracts the inter-pixel dependencies in local regions while capturing the relationships between pixels in that region and all pixels in other distant regions, focusing on key information and reducing the attention given to other irrelevant information, thus enabling better modeling of the relationships between different regions of the image. Therefore, the hyperspectral reconstruction method based on a single RGB image of the present invention can improve the accuracy of the reconstructed hyperspectral image. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the framework of the hyperspectral reconstruction model in an embodiment of the present invention;

[0019] Figure 2 This is a schematic diagram illustrating the structure and working process of the cone-shaped multi-scale feature extraction module in an embodiment of the present invention;

[0020] Figure 3 A schematic diagram illustrating the structure and operation of the multi-scale adaptive residual attention module in an embodiment of the present invention;

[0021] Figure 4 This is a schematic diagram illustrating the working principle of the pixel self-attention module in an embodiment of the present invention;

[0022] Figure 5 The fusion representation T obtained in the embodiments of the present invention out A flowchart;

[0023] Figure 6 This is a schematic diagram of typical example data ARAD_HS_0453 in an embodiment of the present invention;

[0024] Figure 7 This is a schematic diagram comparing the spectral response curves of each model at pixel A in an embodiment of the present invention;

[0025] Figure 8 This is a schematic diagram comparing the spectral response curves of each model at pixel B in an embodiment of the present invention;

[0026] Figure 9 This is a schematic diagram of typical example data ARAD_HS_0459 in an embodiment of the present invention;

[0027] Figure 10 This is a schematic diagram comparing the spectral response curves of each model at pixel C in an embodiment of the present invention;

[0028] Figure 11 This is a schematic diagram comparing the spectral response curves of each model at pixel D in an embodiment of the present invention;

[0029] Figure 12 This is a schematic diagram of typical example data ARAD_HS_0456 in an embodiment of the present invention;

[0030] Figure 13 This is a schematic diagram comparing the spectral response curves of each model at pixel E in an embodiment of the present invention;

[0031] Figure 14 This is a schematic diagram comparing the spectral response curves of each model at pixel F in an embodiment of the present invention;

[0032] Figure 15 This is a schematic diagram of typical example data ARAD_HS_0464 in an embodiment of the present invention;

[0033] Figure 16 This is a schematic diagram comparing the spectral response curves of each model at pixel G in an embodiment of the present invention;

[0034] Figure 17 This is a schematic diagram comparing the spectral response curves of each model at pixel H in an embodiment of the present invention. Detailed Implementation

[0035] To make the technical means, creative features, objectives and effects of this invention easy to understand, the following embodiments, in conjunction with the accompanying drawings, specifically illustrate the hyperspectral reconstruction method based on a single RGB image of this invention.

[0036] In this embodiment, a corresponding hyperspectral image is obtained by inputting a single RGB image into a hyperspectral reconstruction model.

[0037] Figure 1 This is a schematic diagram of the framework of the hyperspectral reconstruction model in an embodiment of the present invention.

[0038] like Figure 1 As shown, the hyperspectral reconstruction model 100 includes a cone-shaped multi-scale feature extraction module 10, M sequentially connected multi-scale adaptive residual attention modules 20, and a combined processing module 30.

[0039] The cone-shaped multi-scale feature extraction module 10 is used to extract features from the feature mapping of a single RGB image after slice partitioning, to obtain shallow region aggregation features. In this embodiment, the input feature F of the cone-shaped multi-scale feature extraction module 10 is... in For feature mapping, output feature F out This represents the aggregation characteristics of shallow regions.

[0040] Figure 2This is a schematic diagram illustrating the structure and working process of the cone-shaped multi-scale feature extraction module in an embodiment of the present invention.

[0041] like Figure 2 As shown, the cone-shaped multi-scale feature extraction module 10 includes a Left sub-module 101 and a Right sub-module 102 connected in sequence. The cone-shaped multi-scale feature extraction module 10 extracts the input features F. in The output feature F is obtained through processing. out The specific work process is as follows:

[0042] The Left submodule 101 includes four Left channel domains and four corresponding convolutional layers of different scales, the Nth... left The kernel size of the convolution corresponding to the left channel domain of the layer is k1 = 2N. left +1, increase the input feature F of size H*W*C. in Divide the channel into four equal parts as the input features for each Left channel domain, and then use the Nth channel as the input feature. left The input features and corresponding convolutional kernels of the left channel domain are both divided into In each Left channel domain, the grouped input features of size H*W*(C / 4m) are further processed using a convolutional kernel of size k1*k1*(C / 4m). That is, the i-th group of convolutional kernels only processes the i-th group of input features, resulting in m convolutional features. Then, all the convolutional features of the four Left channel domains are concatenated to obtain feature F' of size H*W*C. out Then, for feature F' out After processing with batch normalization layers and activation functions, feature F of size H*W*C is obtained. Left As the output of the Left submodule 101, the Left submodule 101 is based on the input feature F in Obtain feature F Left The formula for expressing it is:

[0043]

[0044]

[0045] F Left =Relu(BN(F') out ),

[0046] In the formula, ∑cat represents a multi-level concatenation operation. For the Nth left The i-th feature map of the Left channel domain, w k1,i For the i-th convolutional kernel with kernel size k1, the N-th... lef The number of channels in the left channel domain of layer t is one-quarter of the original input feature F.in The feature maps and their corresponding convolutional kernels are divided into m groups, BN is the batch normalization layer, and ReLU is the activation function.

[0047] The Right submodule 102 includes four Right channel domains and four corresponding convolutional layers of different scales. The four Right channel domains correspond to the four Left channel domains in sequence. The Nth... right The kernel size of the convolution corresponding to the right channel domain of the layer is k² = 11-2N. right , input features F Left Divide the channel into four equal parts as the input features for each Right channel domain, and then divide the Nth channel into four equal parts. right The input features and corresponding convolutional kernels of the left channel domain are both divided into In each right channel domain, the grouped input features of size H*W*(C / 4n) are further processed using a convolutional kernel of size k2*k2*(C / 4n). That is, the i-th group of convolutional kernels only processes the i-th group of input features, resulting in n convolutional features. Then, all the convolutional features of the four right channel domains are concatenated to obtain a feature F of size H*W*C. out Then, for feature F” out After processing by batch normalization layers and activation functions, the output feature F of size H*W*C is obtained. out As the output of Right submodule 102, i.e., the output of cone-shaped multi-scale feature extraction module 10, the aforementioned Right submodule 102, based on the input feature F Left Obtain the output feature F out The formula for expressing it is:

[0048]

[0049]

[0050] F out =Relu(BN(F”) out ),

[0051] In the formula, ∑cat represents a multi-level concatenation operation, and G i,j For the Nth right The i-th feature map of the layer channel domain, w k2,i For the i-th convolutional kernel with kernel size k2, the N-th... right The feature F in the layer channel domain has one-quarter of the original number of channels. Left The feature maps and their corresponding convolutional kernels are divided into m groups, BN is the batch normalization layer, and ReLU is the activation function.

[0052] In this embodiment, the input feature F of the cone-shaped multi-scale feature extraction module 10 is... in and output features F out Since the convolution kernel sizes are inconsistent and all are odd, a half-padding method is used: the padding size is set to [k / 2]k = k1, k2, and the compensation for the four convolution layers is set to the same size.

[0053] In this embodiment, the batch normalization layer and activation function ensure that the entire processing of the cone-shaped multi-scale feature extraction module 10 adapts to nonlinear problems, accelerates convergence to obtain complex deep features, and improves the network's recognition performance.

[0054] In this embodiment, the cone-shaped multi-scale feature extraction module 10, on the one hand, extracts input features F from different areas through four layers of channel domains with progressively increasing kernel sizes in the Left submodule 101. in The module captures local details at different levels. Specifically, it extracts detailed information from smaller local regions using smaller convolutional kernels, and captures detailed information from larger regions, including contextual information, using relatively larger convolutional kernels. The output features of the four channel domains are then concatenated to merge local and relatively global information at the channel level. On the other hand, the Right submodule 102 has four channel domains corresponding to the Left submodule 101, with the corresponding convolutional kernels decreasing sequentially. This allows for the extraction of information from larger regions in the Left submodule 101 using larger convolutional kernels, while the Right submodule 102 captures detailed information from smaller regions within the same channel using smaller convolutional kernels. This interleaved extraction across channels ultimately achieves complementary information at multiple scales. Through the Left and Right submodules 101 and 102, the cone-shaped multi-scale feature extraction module 10 obtains aggregated local information and relatively global information including edge spatial information, ultimately achieving the goal of extracting inter-regional contextual features at multiple scales.

[0055] Figure 3 A schematic diagram illustrating the structure and operation of the multi-scale adaptive residual attention module in an embodiment of the present invention.

[0056] like Figure 3 As shown, the multi-scale adaptive residual attention module 20 includes a conical multi-scale feature extraction module 10, an optimal non-local module 201, a pixel self-attention module 202, a LayerNorm module 203, and a multilayer perceptron module 204 connected in sequence. Its working process is as follows:

[0057] In this multi-scale adaptive residual attention module 20, the Left submodule 101 of the cone-shaped multi-scale feature extraction module 10 processes the input features F. inThe feature F is obtained by sequentially performing multi-scale detailed feature extraction and multi-layer concatenation. Left Then, the feature F is processed through the Right submodule 102. Left The output feature F is obtained by sequentially extracting multi-scale detailed features and concatenating multiple layers. out .

[0058] The optimal nonlocal module 201 will output feature F out The feature sequence T is obtained by dividing the features into four equal-sized regions and then enhancing the connections between the features in these regions. i .

[0059] Among them, the output feature F out Inputting the optimal nonlocal module 201 yields the feature sequence T. i The formula for expressing it is:

[0060] T i =ONB(F out {lu,ld,ru,rd}),

[0061] In the formula, ONB is the optimal nonlocal module 201, F out {lu,ld,ru,rd} represents the feature F out It is divided into four equal-sized areas: upper left, lower left, upper right, and lower right.

[0062] Pixel self-attention module 202 pairs of feature sequences T i The long-range dependencies between pixels are captured to obtain the fused representation T. out .

[0063] Figure 4 This is a schematic diagram illustrating the working principle of the pixel self-attention module in an embodiment of the present invention.

[0064] like Figure 4 As shown, the pixel self-attention module 202 processes the input feature sequence T i The sequence is processed sequentially with convolution, activation function, and convolution again. The results are then divided into Q, K, and V for the attention mechanism. Convolution is then performed on K and V respectively, decomposing the feature information of each channel into information components on the convolution kernel. Each convolution kernel generates new channel information. At this point, the sequence is in the form of (K... i V iElements stored in the form of () contribute differently to key information and contain different amounts of information. Therefore, weight coefficients are needed to represent the importance of relevance to key information. After convolution with K, a correlation matrix is ​​calculated with Q. The result of the correlation matrix is ​​then processed through softmax to obtain attention values, which serve as weights—the weight coefficients for each channel and pixel region. These weights are then multiplied by the transpose of the result of convolution with V, and after further convolution, they are combined with the feature sequence T. i Adding them together yields the fused representation T. out .

[0065] Figure 5 The fusion representation T obtained in the embodiments of the present invention out A flowchart.

[0066] like Figure 5 As shown, the pixel self-attention module 202 focuses on the feature sequence T. i The fused representation T is obtained through processing. out The specific steps are as follows:

[0067] Step S1, for feature T i A combination of convolution-activation function-convolution is used to obtain preliminary features, which are then divided into Q, K, and V for the attention mechanism.

[0068] The formulas for Q, K, and V are as follows:

[0069] Q = Conv(ρ(Conv(T) i ))),

[0070] K = Conv(ρ(Conv(T) i ))),

[0071] V=Conv(ρ(Conv(T i ))),

[0072] In the formula, ρ(·) is the activation function Prelu, and Conv is the convolution operation.

[0073] Step S2: Calculate the weight of each pixel region based on Q and K.

[0074] The formula for the weights is as follows:

[0075] U i =Corr(Q i ,Conv(K i )),

[0076] μ i =softmax(U i ),

[0077] In the formula Ui Let Q be the attention value for the i-th pixel region, and Corr be the value calculated from the correlation matrix. i Let Q and K be the values ​​of the i-th pixel region. i Let K be the region of the i-th pixel, softmax be the softmax function operation, and μ be the value of K. i Let be the weight of the i-th pixel region.

[0078] Step S3, based on the feature sequence T i The fused representation T is obtained by calculating the weights and V. out .

[0079] In step S3, fusion represents T out The formula is as follows:

[0080]

[0081] T o ={T o1 ,T o2 ,T o3 ,...T oi ,...T on},

[0082] T out =T i +δ(Conv(T o )),

[0083] In the formula T oi This is the output for the i-th pixel region. For transpose multiplication, V i Let V and T be the region of the i-th pixel. o δ is the set of outputs from all pixel regions. In this embodiment, the learnable parameter δ is initialized to 0 during training. The model network mainly relies on neighborhood features in the initial stage of training, and then gradually increases the weights of the more distant regions.

[0084] In this embodiment, the pixel self-attention module 202 extracts the inter-pixel dependencies in a local region while capturing the relationships between pixels in that region and all pixels in other, more distant regions. To enhance the capture of long-range inter-pixel dependencies, it uses a self-attention mechanism to focus on key information and reduce the attention given to irrelevant information. Furthermore, since the self-attention mechanism processes the same sequence, the similarity between any two elements in the sequence, with the entire region as the observation scope, can be obtained after a single matrix calculation, eliminating the need to calculate attention values ​​multiple times as the window moves, thus reducing computational load. In summary, the pixel self-attention module 202 can better model the relationships between different regions of an image, improving the accuracy of the generated hyperspectral image.

[0085] Then the fusion representation T out and feature F Left The sum of the results is used as a feature Input LayerNorm module 203. LayerNorm module 203 eliminates the adverse effects caused by outlier data in the summation result.

[0086] The multilayer perceptron module 204 performs feature transformation and information reorganization on the output of the LayerNorm module to obtain features.

[0087] The multi-scale adaptive residual attention module 20 finally inputs the feature F. in ,feature and characteristics The connection is performed to obtain the connection result, which is used as the output of the multi-scale adaptive residual attention module 20, i.e., feature F. i .

[0088] In this embodiment, the shallow region aggregation features are used as the input features F of the first multi-scale adaptive residual attention module 20. in That is, the input of the cone-shaped multi-scale feature extraction module 10 in the first multi-scale adaptive residual attention module 20 is the shallow region aggregated feature.

[0089] In this embodiment, from the first multi-scale adaptive residual attention module 20 to the last second multi-scale adaptive residual attention module 20, the feature F output by each multi-scale adaptive residual attention module 20 is... i The input feature F of the next multi-scale adaptive residual attention module 20 connected to the multi-scale adaptive residual attention module 20 in .

[0090] In this embodiment, the feature F output by the last multi-scale adaptive residual attention module 20 is... i As a deep feature.

[0091] The combined processing module 30 is used to perform convolution and activation function processing on deep features to obtain hyperspectral images. In this embodiment, the combined processing module includes a first convolutional layer, an activation function, and a second convolutional layer connected in sequence, thereby forming a linear-nonlinear combination to sort out the model network and enhance the fitting ability of the model network.

[0092] In this embodiment, the activation functions in the cone-shaped multi-scale feature extraction module 10, the multi-scale adaptive residual attention module 20, and the combined processing module 30 are all parameter-corrected linear units, i.e., Prelude.

[0093] In this embodiment, to test the performance of the hyperspectral reconstruction model 100 of the present invention (i.e., the model of the present invention), three existing deep learning-based spectral reconstruction methods—HRNet, AWAN, and DGCAMN—were used to construct corresponding models, namely the HRNet model, AWAN model, and DGCAMN model. These three models were then retrained with the model of the present invention on the same hardware, programming environment, and training set of the NTIRE 2020 hyperspectral dataset. The RMSE and MRAE metrics of the four trained models were calculated using the typical example data ARAD_HS_0453 in the validation set of the NTIRE 2020 hyperspectral dataset, as shown in the table below.

[0094] Model Name RMSE indicator MRAE index HRNet model 0.01964 0.07037 AWAN model 0.01662 0.04932 DGCAMN model 0.01474 0.04461 This invention model 0.01299 0.03797

[0095] The first column of the table lists the names of each model, the second column lists the RMSE index calculation results of the corresponding model on the typical example data ARAD_HS_0453, and the third column lists the MRAE index calculation results of the corresponding model on the typical example data ARAD_HS_0453. For example, the cell in the fifth row and second column indicates that the RMSE index calculation result of the model of this invention on the typical example data ARAD_HS_0453 is 0.01299. The RMSE index is used to calculate the square root of the average of the sum of squares of the differences between the predicted and the true values. The smaller the calculated RMSE index, the higher the accuracy of the model. The MRAE index is used to measure the relative error between the predicted and the true values. The smaller the calculated MRAE index, the higher the accuracy of the model. Therefore, on the NTIRE2020 hyperspectral dataset, the model of this invention has better accuracy than the other three models, that is, it can reconstruct higher-precision hyperspectral images.

[0096] Figure 6 This is a schematic diagram of typical example data ARAD_HS_0453 in an embodiment of the present invention.

[0097] like Figure 6 As shown, pixels A and B are randomly selected from the typical example data ARAD_HS_0453, and the spectral response curves of the four models mentioned above at these two pixels are compared.

[0098] Figure 7 This is a schematic diagram comparing the spectral response curves of each model at pixel A in an embodiment of the present invention.

[0099] like Figure 7As shown, (a) is the spectral response curve of the HRNet model at pixel A, (b) is the spectral response curve of the AWAN model at pixel A, (c) is the spectral response curve of the DGCAMN model at pixel A, (d) is the spectral response curve of the model of this invention at pixel A, (e) is the ground real spectral response curve at pixel A, and (f) is a comparison diagram of each spectral response curve. The horizontal axis of (a), (b), (c), (d), (e) and (f) are all band numbers, and the vertical axis is the normalized intensity of the spectral response value. In this embodiment, there are a total of 31 band channels. The horizontal axis is 1, which means the first band channel.

[0100] Figure 8 This is a schematic diagram comparing the spectral response curves of each model at pixel B in an embodiment of the present invention.

[0101] like Figure 8 As shown, (a) is the spectral response curve of the HRNet model at pixel B, (b) is the spectral response curve of the AWAN model at pixel B, (c) is the spectral response curve of the DGCAMN model at pixel B, (d) is the spectral response curve of the model of this invention at pixel B, (e) is the ground real spectral response curve at pixel B, and (f) is a comparison diagram of each spectral response curve. The horizontal axis of (a), (b), (c), (d), (e) and (f) are all band numbers, and the vertical axis is the normalized intensity of the spectral response value.

[0102] according to Figure 7 and Figure 8 It can be seen that at pixel A, the spectral response curve of the model of the present invention is basically consistent with the real spectral response curve of the ground. At pixel B, the spectral response curve of the model of the present invention is basically consistent with the real spectral response curve of the ground in the 1-7 and 14-22 band channels, namely 400nm-460nm and 530-610nm. In the other bands, it is also closer to the real spectral response curve of the ground than the other three models.

[0103] Figure 9 This is a schematic diagram of typical example data ARAD_HS_0459 in an embodiment of the present invention.

[0104] like Figure 9 As shown, pixels C and D are randomly selected from the typical example data ARAD_HS_0459, and the spectral response curves of the four models mentioned above at these two pixels are compared.

[0105] Figure 10 This is a schematic diagram comparing the spectral response curves of each model at pixel C in an embodiment of the present invention.

[0106] like Figure 10 As shown, (a) is the spectral response curve of the HRNet model at pixel C, (b) is the spectral response curve of the AWAN model at pixel C, (c) is the spectral response curve of the DGCAMN model at pixel C, (d) is the spectral response curve of the model of this invention at pixel C, (e) is the ground real spectral response curve at pixel C, and (f) is a comparison diagram of each spectral response curve. The horizontal axis of (a), (b), (c), (d), (e) and (f) are all band numbers, and the vertical axis is the normalized intensity of the spectral response value.

[0107] Figure 11 This is a schematic diagram comparing the spectral response curves of each model at pixel D in an embodiment of the present invention.

[0108] like Figure 11 As shown, (a) is the spectral response curve of the HRNet model at pixel D, (b) is the spectral response curve of the AWAN model at pixel D, (c) is the spectral response curve of the DGCAMN model at pixel D, (d) is the spectral response curve of the model of this invention at pixel D, (e) is the ground real spectral response curve at pixel D, and (f) is a comparison diagram of each spectral response curve. The horizontal axis of (a), (b), (c), (d), (e) and (f) are all band numbers, and the vertical axis is the normalized intensity of the spectral response value.

[0109] according to Figure 10 and Figure 11 It can be seen that at pixel C, the spectral response curve of the HRNet model in the 7-17 band, i.e., 460nm-560nm, differs significantly from the actual ground spectral response curve. At pixel D, the spectral response curves of the HRNet model, AWAN model, and DGCAMN model all fluctuate with the actual ground spectral response curve, while the spectral response curve of the model of this invention still has the best fit with the actual ground spectral response curve.

[0110] In summary, the hyperspectral reconstruction method based on a single RGB image of the present invention can generate higher quality hyperspectral images compared with existing hyperspectral reconstruction methods.

[0111] To verify the effectiveness of the pixel self-attention module 202 and the cone-shaped multi-scale feature extraction module 10 proposed in the hyperspectral reconstruction method based on a single RGB image of this invention, the base model used for hyperspectral reconstruction was adjusted as follows: the pixel self-attention module 202 was adaptively adjusted and added to the base model to obtain model Ea; the cone-shaped multi-scale feature extraction module 10, which only contains a single component for obtaining complementary multi-scale information, was adaptively adjusted and added to the base model to obtain model Eb; and the cone-shaped multi-scale feature extraction module 10 was adaptively adjusted and added to the base model to obtain model Ec. Then, the three models were iterated 30 times each using validation set data. The best RMSE and MRAE indices for each model during this process are shown in the table below.

[0112] Model Name RMSE indicator MRAE index Ea model 0.01901 0.05940 Eb model 0.01704 0.04725 Ec model 0.01299 0.03797

[0113] The first column of the table lists the names of each model, the second column lists the RMSE index calculation results for the corresponding model, and the third column lists the MRAE index calculation results for the corresponding model. For example, the cell in the second row and second column shows that the best RMSE index calculation result for the Ea model is 0.01901. As can be seen from the table above, compared with the basic model, the pixel self-attention module 202 added to the Ea model captures long-range dependencies by utilizing features from all locations and highlights the features of important regions, thereby improving the expressive power of spatial features. The Eb model captures detailed features at different regional levels by changing the area of ​​input feature information obtained by each channel group through the single-component cone-shaped multi-scale feature extraction module 10, realizing the powerful learning ability of the network. The Ec model captures interleaved information of regional information of different sizes on the channel through the complete cone-shaped multi-scale feature extraction module 10, which can achieve the purpose of information complementarity. By aggregating local information and relatively global information containing edge spatial information, the RMSE index and MRAE index are improved.

[0114] To more intuitively demonstrate the differences in the hyperspectral images generated by the three models, two random pixels were selected from the typical example data ARAD_HS_0456 and ARAD_HS_0464, respectively, and the spectral response curves were compared.

[0115] Figure 12 This is a schematic diagram of typical example data ARAD_HS_0456 in an embodiment of the present invention.

[0116] like Figure 12 As shown, pixels E and F are selected in the typical example data ARAD_HS_0456 for comparison of the spectral response curves of each model.

[0117] Figure 13This is a schematic diagram comparing the spectral response curves of each model at pixel E in an embodiment of the present invention.

[0118] like Figure 13 As shown, (a) is the spectral response curve of the Ea model at pixel E, (b) is the spectral response curve of the Eb model at pixel E, (c) is the spectral response curve of the Ec model at pixel E, (d) is the ground true spectral response curve at pixel E, and (e) is a comparison of the various spectral response curves. The horizontal axis of (a), (b), (c), (d) and (e) are all band numbers, and the vertical axis is the normalized intensity of the spectral response value.

[0119] Figure 14 This is a schematic diagram comparing the spectral response curves of each model at pixel F in an embodiment of the present invention.

[0120] like Figure 14 As shown, (a) is the spectral response curve of the Ea model at pixel F, (b) is the spectral response curve of the Eb model at pixel F, (c) is the spectral response curve of the Ec model at pixel F, (d) is the ground true spectral response curve at pixel F, and (e) is a comparison of the various spectral response curves. The horizontal axis of (a), (b), (c), (d) and (e) are all band numbers, and the vertical axis is the normalized intensity of the spectral response value.

[0121] according to Figure 13 and Figure 14 It can be seen that at pixels E and F, the spectral response curves of the Ea model are consistent with the ground's true spectral response curve in terms of direction, but there is a certain deviation in value. In the 1-23 band, i.e., 400nm-620nm, the spectral response curves of the Eb model and the Ec model are basically consistent with the ground's true spectral response curve. In the 24-31 band, i.e., 630nm-700nm, the spectral response curve of the Ec model is closer to the ground's true spectral response curve. That is, among the three models, the Ec model has the best hyperspectral reconstruction effect.

[0122] Figure 15 This is a schematic diagram of typical example data ARAD_HS_0464 in an embodiment of the present invention.

[0123] like Figure 15 As shown, pixels G and H are selected in the typical example data ARAD_HS_0464 for comparison of the spectral response curves of each model.

[0124] Figure 16 This is a schematic diagram comparing the spectral response curves of each model at pixel G in an embodiment of the present invention.

[0125] like Figure 16As shown, (a) is the spectral response curve of the Ea model at pixel G, (b) is the spectral response curve of the Eb model at pixel G, (c) is the spectral response curve of the Ec model at pixel G, (d) is the ground true spectral response curve at pixel G, and (e) is a comparison of the various spectral response curves. The horizontal axis of (a), (b), (c), (d) and (e) are all band numbers, and the vertical axis is the normalized intensity of the spectral response value.

[0126] Figure 17 This is a schematic diagram comparing the spectral response curves of each model at pixel H in an embodiment of the present invention.

[0127] like Figure 17 As shown, (a) is the spectral response curve of the Ea model at pixel H, (b) is the spectral response curve of the Eb model at pixel H, (c) is the spectral response curve of the Ec model at pixel H, (d) is the ground true spectral response curve at pixel H, and (e) is a comparison of the various spectral response curves. The horizontal axis of (a), (b), (c), (d) and (e) are all band numbers, and the vertical axis is the normalized intensity of the spectral response value.

[0128] according to Figure 16 and Figure 17 It can be seen that at pixel G, in the 1-25 band (400nm-640nm), the spectral response curve of the Ec model perfectly matches the ground-based true spectral response curve. In the 26-31 band (650nm-700nm), when the spectral response curves of the three models deviate from the ground-based true spectral response curve, the spectral response curve of the Ec model is closer to the ground-based true spectral response curve. At pixel H, in the 24-31 band (630nm-700nm), the spectral response curves of the Ea and Eb models deviate from the ground-based true spectral response curve, but the spectral response curve of the Eb model is closer to the ground-based true spectral response curve, while the curve of the Ec model still basically matches the ground-based true spectral response curve. Therefore, among the three models, the Ec model has the best hyperspectral reconstruction effect.

[0129] In summary, the cone-shaped multi-scale feature extraction module 10 is more ideal in tracking the lost information between local regions during the feature extraction process. The fusion of information from small regions and relatively large regions can help the model reconstruct a more accurate hyperspectral image.

[0130] The role and effect of the embodiments

[0131] According to the hyperspectral reconstruction method based on a single RGB image involved in this embodiment, on the one hand, the cone-shaped multi-scale feature extraction module 10 extracts detailed features at different scales on four different channels in an alternating manner, thereby obtaining aggregated local information and relatively global information including edge spatial information, ultimately achieving the purpose of extracting inter-regional contextual features at multiple scales; on the other hand, the pixel self-attention module 202 extracts the inter-pixel dependencies in local regions while capturing the relationships between pixels in that region and all pixels in other more distant regions, focusing on key information and reducing the attention given to other irrelevant information, thereby enabling better modeling of the relationships between different regions of the image. In summary, this method can improve the accuracy of the reconstructed hyperspectral image.

[0132] The above embodiments are preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention.

Claims

1. A hyperspectral reconstruction method based on a single RGB image, for inputting a single RGB image into a hyperspectral reconstruction model to obtain a corresponding hyperspectral image, characterized in that, The hyperspectral reconstruction model comprises: a conical multi-scale feature extraction module, configured to perform feature extraction on a feature map of the single RGB image after slice partition processing, to obtain shallow regional aggregation features; a plurality of sequentially connected multi-scale adaptive residual attention modules, configured to process the shallow regional aggregation features to obtain deep features; a combination processing module, configured to perform convolution and activation function processing on the deep features to obtain the hyperspectral image, wherein the multi-scale adaptive residual attention module comprises a conical multi-scale feature extraction module, an optimal non-local module, a pixel self-attention module, a LayerNorm module and a multi-layer perception module connected in sequence, The cone-shaped multi-scale feature extraction module is configured to take the feature mapping of the single RGB image after slice partition processing as input feature F in , and the input of the multi-scale adaptive residual attention module containing the cone-shaped multi-scale feature extraction module as input feature F in , and sequentially perform multi-scale detail feature preliminary extraction and multi-layer splicing to obtain feature F Left , and then perform multi-scale detail feature interleaved extraction and multi-layer splicing on the feature F Left to obtain the shallow region aggregation feature or the input of the optimal non-local module connected with the cone-shaped multi-scale feature extraction module as output feature F out , The optimal non-local module is used to divide the output feature F out into four equally sized regions, and the connection between the partition features is further enhanced to obtain a feature sequence T i , The pixel self-attention module is used to capture long-range dependencies between pixels in the feature sequence T i , to obtain a fused representation T out , The LayerNorm module is used to eliminate the adverse effects of singular sample data in the addition result of the fusion representation T out and the feature F Left , and the addition result is taken as a feature The multi-layer perception module is used for feature conversion and information reorganization on the output of the LayerNorm module, to obtain features In each of the multi-scale adaptive residual attention modules, the input features F in , the features , and the features are connected to obtain a connection result as the input of a next multi-scale adaptive residual attention module connected in sequence with the multi-scale adaptive residual attention module. and the connection result of the last multi-scale adaptive residual attention module is the deep features.

2. The hyperspectral reconstruction method based on a single RGB image according to claim 1, wherein: wherein the conical multi-scale feature extraction module comprises a Left sub-module and a Right sub-module connected in sequence, The Left submodule includes four Left channel domains and four corresponding convolutional layers at different scales. Each channel in the four Left channel domains is input to the input feature F of the cone-shaped multi-scale feature extraction module. in A quarter of the channel, the Nth left The kernel size of the convolution corresponding to the left channel domain of the layer is k1 = 2N. left +1, the input feature F in The feature F is obtained by inputting the Left submodule. Left The formula for expressing it is: F Left = Relu(BN(F out ), where ∑cat is a multi-layer concatenation operation, is the Nth left is the i-th group of feature maps of the Left channel domain of the Nth k1,i is the i-th group of convolution kernels with the size of k1, the Nth left is the input feature F in the Left channel domain of the Nth in As the feature maps and the corresponding convolution kernels are both divided into m groups, BN is a batch normalization layer, and Relu is an activation function, The Right submodule includes four layers of Right channel domains and corresponding four layers of different scale convolutions, the channels of the four layers of Right channel domains are each one quarter of the channels of the input feature F input to the conical multi-scale feature extraction module in , the four layers of Right channel domains correspond to the four layers of Left channel domains in sequence, the Nth layer of Right channel domain corresponds to a convolution kernel size k2=11-2N right of the convolution right , the expression formula of the feature F Left input to the Right submodule to obtain the output feature F out is: F out =Relu(BN(F″ out ), where ∑cat is a multi-layer concatenation operation, G i,j is the i-th group of feature maps of the N right -th layer channel domain, w k2,i is the i-th group of convolution kernels with k2 as the size of the convolution kernel, the N right -th layer channel domain, the channel number of the feature F Left As the feature map and the corresponding convolution kernel are divided into m groups, BN is the batch normalization layer, and Relu is the activation function.

3. The hyperspectral reconstruction method based on a single RGB image according to claim 1, wherein: wherein, The output feature F out The input of the optimal non-local module obtains the feature sequence T i The expression formula of the feature sequence T T i = ONB(F out {lu, ld, ru, rd}), where ONB is the optimal non-local module, F out {lu, ld, ru, rd} are the four equal-sized regions into which the feature F out is split into the top-left, bottom-left, top-right, and bottom-right.

4. The method of claim 1, wherein the method is based on a single RGB image. characterized in that: The pixel self-attention module processes the feature sequence T i to obtain a fusion representation T out The specific steps are as follows: Step S1, obtaining the feature T i The combination of convolution-activation function-convolution is used to process to obtain a preliminary feature, and the preliminary feature is divided into Q, K and V of the attention mechanism. step S2, calculating the weight of each pixel region according to Q and K; Step S3, computing the fusion representation T i from the feature sequence T out , the weights and V.

5. The hyperspectral reconstruction method based on a single RGB image according to claim 4, wherein: wherein, in the step S1, the formulas of Q, K and V are as follows: Q = Conv(p(Conv(T i ))), K = Conv(p(Conv(T i ))), V = Conv(p(Conv(T i ))), wherein ρ(·) is an activation function Prelu, and Conv is a convolution operation.

6. The hyperspectral reconstruction method based on a single RGB image according to claim 4, wherein: wherein, in the step S2, the formula of the weight is as follows: U i = Corr(Q i , Conv(K i )) μ i = softmax(U i ), wherein U i is the attention value for the i-th pixel region, Corr is the correlation matrix computation, Q i is Q for the i-th pixel region, K i is K for the i-th pixel region, softmax is the softmax function operation, μ i is the weight for the i-th pixel region.

7. The hyperspectral reconstruction method based on a single RGB image according to claim 4, wherein: wherein In said step S3, said fusion representation T out The formula is as follows: T o = {T o1 , T o2 , T o3 ,...T oi ,...T on}, T out = T i + δ(Conv(T o )), where T oi is the output of the i-th pixel region, is the transpose multiplication, V i is the V of the i-th pixel region, T o is the set of outputs of all pixel regions, and δ is a learnable parameter.

Citation Information

Patent Citations

  • Hyperspectral image reconstruction method based on double-ghost attention mechanism network

    CN112819910A

  • Hyperspectral image super-resolution reconstruction method based on multi-scale space-spectrum feature learning

    CN115272078A