ResCNT-based hyperspectral image reconstruction method
By introducing the ResCNT network in hyperspectral image reconstruction, combining the convolutional neural network and the Transformer model, using the multi-residual jump connection and the Neighborhood Attention module, the problems of insufficient reconstruction accuracy, high computational complexity and poor robustness in the existing technology are solved, and efficient and accurate hyperspectral image reconstruction is achieved.
Patent Information
- Application Number
- CN202510440140.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-30
AI Technical Summary
When existing hyperspectral image reconstruction methods process low-quality, high noise or low-resolution images, the reconstruction accuracy is insufficient, the computational complexity is high, and the robustness is poor, making it difficult to effectively capture global and local information in hyperspectral images.
Using a hyperspectral image reconstruction method based on ResCNT, combining convolutional neural network and Transformer model, the image details recovery ability and global information modeling ability are improved through multi-residual jump connection and Neighborhood Attention module.
It significantly improves the accuracy and robustness of hyperspectral image reconstruction, reduces the consumption of computing resources, and maintains good reconstruction quality under noise interference, which is suitable for use in environments with limited computing resources.
Smart Images

Figure CN120070771A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a method for hyperspectral image (HSI) reconstruction based on ResCNT (Multi-Residual Skip Connections Neighborhood Transformer). Background Art
[0002] With the continuous progress of technology, hyperspectral imaging technology has been widely applied in many fields, such as remote sensing, agricultural monitoring, environmental protection, medical imaging, and industrial inspection. Hyperspectral images (HSI) can capture detailed spectral information of objects by acquiring image data of multiple continuous bands, thereby providing richer image features and information than traditional RGB images. This makes hyperspectral images have significant advantages in tasks such as object recognition, classification, detection, and environmental monitoring.
[0003] However, hyperspectral imaging technology faces some challenges in practical applications. Especially when acquiring hyperspectral images, due to the limitations of imaging devices and environmental factors, problems such as image reconstruction or restoration often occur. These problems may include image noise, missing data, low resolution, etc., resulting in a decline in image quality and affecting the accuracy of subsequent processing and analysis. Therefore, how to effectively reconstruct higher-quality images from low-quality, high-noise, or low-resolution hyperspectral images is an important research topic in the field of hyperspectral image processing currently.
[0004] Traditional hyperspectral image reconstruction methods mainly include classical algorithms based on image restoration, such as Least Squares, Sparse Representation, Dictionary Learning, etc. Although these methods can restore image quality to a certain extent, they usually face the following problems:
[0005] 1. Insufficient reconstruction accuracy: Since the algorithms adopted are based on manually designed features, it is difficult to handle complex hyperspectral image data, and the reconstruction effect is often limited.
[0006] 2. High computational complexity: These classical methods have a large amount of computation, especially when dealing with large-scale image data, the efficiency is low.
[0007] 3. Poor robustness: In the case of noise interference or severe image loss, the performance of traditional algorithms will decline significantly, resulting in an unsatisfactory restoration effect.
[0008] In recent years, with the rapid development of deep learning, methods based on deep neural networks have become a new direction for hyperspectral image reconstruction. Deep learning methods, especially Convolutional Neural Networks (CNNs), have been widely applied to image processing tasks, including image restoration, denoising, super-resolution, etc., due to their powerful feature learning ability. However, although convolutional neural networks have achieved good results in some tasks, they still have difficulty effectively capturing the complex global and detailed information in hyperspectral images because of their limited ability to model long-range dependencies when dealing with hyperspectral images.
[0009] To address these issues, in recent years, hybrid architectures that combine convolutional neural networks and Transformer models have gradually become a new trend. Due to its self-attention mechanism, the Transformer model can effectively capture long-range dependencies in images, thus providing a more accurate global information modeling ability. However, the application of the Transformer model in hyperspectral image reconstruction still faces certain challenges, mainly manifested as follows:
[0010] 1. High computational resource requirements: The Transformer model is generally more computationally complex than CNN, especially when dealing with large-scale hyperspectral image data, which may bring a large computational overhead.
[0011] 2. Low information transfer efficiency: In hyperspectral images, due to the strong correlation between each band, simply relying on the self-attention mechanism may lead to over-propagation of information or unnecessary calculations, reducing efficiency.
[0012] Therefore, how to design a method that can effectively combine the advantages of convolutional neural networks and Transformer models, while ensuring high efficiency, better capture the global and local information in hyperspectral images, remains a research problem in the current field of hyperspectral image reconstruction. Summary of the Invention
[0013] Aiming at the defects of the existing technology, the present invention provides a hyperspectral image reconstruction method based on ResCNT. By introducing a hybrid structure of convolutional neural network (CNN) and Transformer, while maintaining the ability of the convolutional layer to capture local information, the self-attention mechanism of the Transformer is used to model the global information of the image. In this way, the quality of image reconstruction can be effectively improved, while reducing the consumption of computational resources and having strong robustness, especially in the face of noise or missing image data.
[0014] To achieve the above invention objectives, the technical solutions adopted by the present invention are as follows:
[0015] A hyperspectral image reconstruction method based on ResCNT includes the following steps:
[0016] Obtain compressed two-dimensional measurements generated by an encoded aperture snapshot spectral imaging (CASSI) system;
[0017] Use a ResCNT (Residual Convolutional Network Transformer) network for reconstruction. The ResCNT network includes multiple convolutional layers and residual modules, combines the Transformer structure, and recovers the three-dimensional information of the hyperspectral image from the compressed two-dimensional measurements through multi-scale learning;
[0018] During the reconstruction process, adopt Multi-Residual Skip Connections (MRSC) to enhance the ability to restore image details and strengthen the multi-scale representation of feature maps; by introducing skip connections between different network layers, optimize the gradient path in backpropagation, thereby improving the convergence of the network;
[0019] Based on the Neighborhood Attention (NA) module of the self-attention mechanism, dynamically assign different weights to each pixel at each stage of image reconstruction to enhance the consistency of local region details and global information;
[0020] Use the backpropagation algorithm to train the network, optimize the network parameters by minimizing the error between the reconstructed image and the real hyperspectral image, and finally restore the hyperspectral image.
[0021] Furthermore, the method further includes: applying a Degradation-Aware Unfolding Framework (DAUF) in the denoising framework to unfold the compressed two-dimensional measurements to improve the reconstruction effect.
[0022] Furthermore, the specific structure of the ResCNT network includes:
[0023] A convolutional encoder module for extracting low-level image features from the compressed two-dimensional measurements, including multiple convolutional layers, Batch Normalization, and activation functions;
[0024] A residual block for strengthening the transmission of deep features through skip connections and reducing the problem of gradient disappearance in the training of deep networks;
[0025] A Transformer module applies the self-attention mechanism in the network, enhances the long-range dependencies of the image by weighted averaging of features, and models the global information of the features;
[0026] A convolutional decoder module gradually restores the detailed information of the hyperspectral image through multiple deconvolution operations;
[0027] An output layer compares the reconstruction result with the real hyperspectral image by combining a pixel-level loss function to optimize the network output.
[0028] Furthermore, the residual module avoids the attenuation of information in deep propagation by introducing multiple skip connections, while retaining the detailed information, ensuring the accuracy of image restoration.
[0029] Furthermore, the Transformer module establishes dynamic relationships between features at different levels through the self-attention mechanism, weights different regions using adaptive weights, thereby optimizing the global information transmission in the image.
[0030] Furthermore, the NA module calculates the similarity between features using a sliding window around each pixel and dynamically assigns different attention weights to each pixel to enhance the details and image consistency.
[0031] Furthermore, the reconstruction process also includes denoising using the image prior information to further improve the quality of the reconstructed image and reduce the noise impact during the compression process.
[0032] Furthermore, the compressed two-dimensional measurement values are modulated by the physical mask of the spectral encoding aperture, displaced along the spectral dimension by a disperser, and finally obtained through integration.
[0033] Furthermore, the hyperspectral image through the action of the sensing matrix, combines the compressed two-dimensional measurement values and noise to recover the original three-dimensional hyperspectral image.
[0034] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned hyperspectral image reconstruction method is implemented.
[0035] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned hyperspectral image reconstruction method is implemented.
[0036] Compared with the prior art, the advantages of the present invention are:
[0037] 1. The proposed ResCNT in the present invention significantly improves the accuracy of hyperspectral image reconstruction by introducing multi-residual skip connections (MRSC) and neighborhood attention mechanism (NA). Experiments show that compared with traditional deep unfolding methods and other model-based reconstruction methods, ResCNT performs better in terms of PSNR and SSIM scores and can better restore the details and textures of images.
[0038] 2. Although ResCNT outperforms many existing methods in terms of performance, its computational cost increases less. Especially after introducing MRSC and NA, while maintaining high reconstruction accuracy, the number of parameters and floating-point operations (FLOPS) of the model are significantly lower than some other methods, thus effectively reducing the demand for computing resources and making it suitable for use in environments with limited computing resources.
[0039] 3. Under different noise levels, ResCNT shows higher stability compared with other SOTA methods. Especially under high noise interference, ResCNT can still maintain good reconstruction quality, demonstrating its strong noise robustness. This enables the method to handle more types and intensities of noise interference in practical applications, enhancing its practicality and reliability.
[0040] 4. By introducing the neighborhood attention mechanism (NA), the present invention enhances the flexibility of feature learning and window translation equivalence, and optimizes the local modeling ability of features. Compared with the traditional self-attention mechanism, NA can dynamically adjust the window size when processing hyperspectral data, thus effectively improving the performance of image reconstruction.
[0041] 5. The design of the ResCNT method is not limited to hyperspectral image reconstruction, and in other computer vision tasks, especially in the fields of image denoising, image restoration, etc., good results can also be obtained after appropriate adjustment. Therefore, this method has strong versatility and can be widely applied to multiple fields of image processing.
[0042] 6. Compared with other methods, ResCNT significantly reduces the memory usage while retaining high performance. This makes the method suitable for devices with limited memory resources and expands its application scenarios. Brief Description of the Drawings
[0043] Figure 1 is the architecture diagram of the hyperspectral image reconstruction method based on ResCNT in the embodiment of the present invention; where (a) is DAUF, (b) is ResCNT, (c) is MRSC, (d) is NAB, (e) is FFN, and (f) is NA;
[0044] Figure 2 is the simulation result diagram of HSI reconstruction in the embodiment of the present invention;
[0045] Figure 3 It is the result graph of the real HSI reconstruction Scene 1 in the embodiment of the present invention;
[0046] Figure 4 It is the result graph of the real HSI reconstruction Scene 3 in the embodiment of the present invention;
[0047] Figure 5 It is the result graph of the real HSI reconstruction Scene 4 in the embodiment of the present invention;
[0048] Figure 6 It is the result graph of the step-by-step ablation experiment of the module in the embodiment of the present invention;
[0049] Figure 7 It is the result graph of the ablation experiment of the network architecture in the embodiment of the present invention;
[0050] Figure 8 It is the comparison graph of the self-attention mechanism in the embodiment of the present invention. Detailed implementation manners
[0051] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention with reference to the drawings and by way of examples.
[0052] As Figure 1 shown, the present invention proposes a hyperspectral image (HSI) reconstruction method, which combines a degradation-aware unfolding framework (DAUF) and a multi-residual jump connection neighborhood transformer (ResCNT). This method first realizes the reconstruction of hyperspectral images by solving the following optimization problem through DAUF;
[0053] I. As Figure 1 (a) shown, the DAUF framework is used in hyperspectral image reconstruction, and the core objective is to solve the degradation modeling and optimization of the image reconstruction problem. Specifically, the reconstruction task is solved through the following optimization problem:
[0054]
[0055] where τ is a non-negative trade-off parameter used to balance the relationship between the fidelity term (the first term) and the penalty term (the second term). To better decouple the variables x and z in the optimization, an auxiliary variable z is introduced to transform this optimization problem into a form with constraints:
[0056]
[0057] On this basis, the variables x and z are optimized through a two-step update process, namely the linear projection and Gaussian denoising steps, which are updated through the following formulas respectively:
[0058]
[0059] Among them, μ is a parameter used to control the hypothesis penalty coefficient, and k represents the current iteration number. Through the Bayesian probability framework, the above formula can be regarded as the process of denoising the image x k+1 under Gaussian noise.
[0060] To achieve the adaptive estimation of variables μ and σ, the DAUF framework uses a parameter estimator. Through the following formula, the updated parameters can be obtained:
[0061] (μ,σ)&=ε(y,Φ)
[0062]
[0063] Among them, represents the linear projection operation, represents the Gaussian denoising operation. Figure 1 (a) shows the design of DAUF. The initial state z 0 is obtained by shifting y and combined with Φ through a convolution operation for iterative solution.
[0064] II. Multi-Residual Skip Connection Neighborhood Transformer (ResCNT)
[0065] To further improve the denoising effect, the present invention proposes a new denoiser structure - Multi-Residual Skip Connection Neighborhood Transformer (ResCNT), which combines Multi-Residual Skip Connections (MRSC) and Neighborhood Attention (NA) to enhance the hyperspectral image reconstruction effect.
[0066] 1) Network architecture
[0067] As Figure 1 (b) shows, the ResCNT network adopts a U-shaped architecture. In this architecture, Multi-Residual Skip Connections (MRSC) are extended for connecting feature maps at multiple scales. The specific process is as follows:
[0068] The image X k at the current iteration +1 is concatenated with the degradation parameters (such as the noise level μ k / τ k ), and after a conv3×3 convolution operation, a preliminary feature map is obtained Then, X 0 undergoes multi-stage processing in the encoder to generate hierarchical features where i represents the number of features at each stage of the encoder (one NAB block, one downsampling module, one NAB block, and one downsampling module).
[0069] Introduce a bottleneck structure in the middle layer, which contains a NAB and an MRSC. By compressing the feature dimension and introducing the NA and MRSC modules, the network's expression of low-frequency and high-frequency information is further strengthened.
[0070] Apply a symmetric decoder to generate a series of hierarchical feature maps
[0071] Residual image generation and weighting: Based on the output image in the decoder stage, generate a residual image And add it element-wise to the current image x k to obtain the final denoised image z k .
[0072] 2) MRSC is one of the key innovations of ResCNT. Its design purpose is to enhance local accuracy and improve the quality of the reconstructed image. The main idea of MRSC (Multi-Residual Skip Connection) is to improve the localization accuracy and the final feature representation by adding skip connections in the U-shaped structure. Specifically, first, the work of the present invention is to logically extend the conventional lateral skip connections in the U-shaped architecture. Theoretically, adding any hierarchical connections in the U-shaped architecture can enhance the performance of the conventional lateral skip connections, such as dense connections. However, the present invention believes that for hyperspectral image (HSI) reconstruction, the connections between layers must be restricted. Specifically, it has been proven that the additional connections from low-resolution features to high-resolution features only enhance the feature semantics and do not improve the local features, and may even have the opposite effect. Therefore, the present invention connects features of multiple scales in the decoder by connecting with the high-resolution features of the encoder. In addition, it makes sense to select appropriate operations to implement MRSC. Any operation that changes the size can adjust the high-resolution features to the same size as the features connected in the decoder, such as using strided convolution or pooling, while element-wise addition or concatenation can merge multiple feature maps into one. To reduce the computational burden, the present invention uses convolution, max pooling, and element-wise addition in MRSC, which has been proven to be able to maintain spatial accuracy and computational efficiency. Finally, theoretically, the more additional connections at higher levels will shift the focus to local features. Therefore, MRSC contains features of equal resolution and all higher-resolution features. As Figure 1 (c) shows, MRSC is described in detail, how the features are connected with features of the same resolution, and all higher-resolution features are connected after downsampling. The downsampling operation is achieved by max pooling.
[0073] In implementation, MRSC enhances the feature maps through the following steps:
[0074] 1. In each decoder stage, in addition to connecting with the feature map of the current layer, it also connects with higher-resolution features from the encoder at the same time.
[0075] 2. Use the conv1×1 convolution operation and pooling operation to adjust the feature dimensions between different scales to ensure that they can maintain a consistent scale when connected.
[0076] 3. Element-wise addition and concatenation operations combine feature maps from different scales, thereby effectively enhancing the details and structure of the final image.
[0077] The design of MRSC enables ResCNT to obtain sufficient information from feature maps of different scales, avoiding the problem of information loss caused by a single scale.
[0078] 3) Neighborhood Attention (NA)
[0079] In traditional attention mechanisms, although the global attention mechanism has achieved great success, there are still some problems, mainly manifested in the limitations of computational complexity and memory usage. To this end, the present invention proposes a new neighborhood attention (NA) mechanism, which can effectively address this problem.
[0080] Different from the traditional global attention mechanism, the NA mechanism calculates attention within a local region, enabling each feature point to only focus on other points in its neighborhood, thereby reducing the computational burden and being able to better maintain translational equivalence. The calculation process of NA is based on the following formula:
[0081] Figure 1 (f) shows the Neighborhood Attention (NA) used in ResCNT. The present invention defines the input as and the linear projections of \(\mathbf{X}_{in}\), which are respectively and Then, a dot product operation is performed on the query Q and its k nearest neighbors K at position ij to obtain the attention weights with a neighborhood size of k at position ij The formula is as follows:
[0082]
[0083] where, \(\rho\) k (ij) represents the k-th nearest neighbor at position i, j, and B represents the relative position bias. Corresponding to the neighborhood value is also a matrix, which consists of the values valueV of the k nearest neighbors of the input at position ij. The formula is expressed as:
[0084]
[0085] Subsequently, the calculation formula of the NA mechanism with a neighborhood size of k at position ij is as follows:
[0086]
[0087] Among them, represents the dimension parameter of the key vector. For each input position ij, its neighborhood attention is calculated, and the final output is obtained as follows
[0088] The introduction of the NA mechanism enables ResCNT to focus on more relevant features within a local range, thereby improving the reconstruction accuracy of the image.
[0089] III. Experiments
[0090] In this embodiment, ResCNT is implemented using the Pytorch framework and trained on an RTX 4090 GPU. During the training process, the Adam optimizer is adopted, where β1 and β2 are set to 0.9 and 0.999 respectively, the learning rate adopts the CosineAnnealing scheduling strategy, the initial learning rate is set to 4×10^-4, and ε is set to 10^-6. The cropped sizes of the training data are 256×256 and 384×384, and the data is augmented by random rotation and flipping.
[0091] During the training process, the batch size is set to 2, and the ResCNT model has 9 recovery stages. Except for the first and the last stages, the other stages share weights. The loss function is defined as the root mean square error (RMSE) between the reconstruction result and the original hyperspectral image.
[0092] To quantitatively study the effectiveness of ResCNT, the present invention compares it in ten test scenarios on a simulated dataset, specifically comparing performance metrics such as the number of parameters (Params), floating-point operations (FLOPS), peak signal-to-noise ratio (PSNR), and structural similarity (SSIM). The experimental results are compared with other advanced hyperspectral image reconstruction methods, including model-based algorithms (such as TwIST, GAP-TV, DesCI), end-to-end methods (such as λ-Net, TSA-Net, HDNet, MST-L, BIRNAT), and deep unfolding methods (such as DSSP, DNU, DGSMP, DAUHST, GAP-Net, RCUMP). To ensure the fairness of the comparison, all methods are tested under the same experimental settings.
[0093] The experimental results are as follows:
[0094] PSNR: The ResCNT method of the present invention achieved 39.16 dB in PSNR, exceeding the existing two SOTA deep unfolding methods, DAUHST and RCUMP, by 0.80 dB and 0.20 dB respectively, demonstrating its advantage in reconstruction quality.
[0095] SSIM: In terms of the SSIM metric, ResCNT reached 0.971, also showing significant improvement compared to DAUHST and RCUMP, with increases of 0.015 and 0.005 respectively.
[0096] Comparison with other methods: Compared with the CNN-based E2E method HDNet, ResCNT increased by 4.19 dB; compared with the Transformer-based E2E method MST-L, ResCNT increased by 4.19 dB; compared with the RNN-based E2E method BIRNAT, ResCNT increased by 1.58 dB.
[0097] Figure 2 Shows the comparison of the reconstruction performance of ResCNT and other methods on the simulated dataset. In this embodiment, a yellow box was selected from the reconstructed HSI image for further analysis. The enlarged selected area shows that the texture of the image restored by ResCNT is more detailed and closer to the real value. In addition, for the spectral density curves shown in the two green boxed areas in the image RGB, the spectral density curve of ResCNT has a stronger correlation with the real value, verifying its superior performance.
[0098] This embodiment also further verified its effect through the reconstruction results of ResCNT and other methods on the real dataset. 11-bit shot noise was added to the 2D measurement to simulate the noise generated during the actual imaging process. In this embodiment, the model was retrained on the real dataset using a real mask, and the reconstruction results of Scene 1, 3, and 4 are shown in Figure 3 、 4 、5. Especially in Figure 4 , ResCNT restored more details and textures in the eye area of the boy's face, providing an impressive reconstruction effect. The experimental results show that ResCNT can restore more detailed textures and exhibits stronger robustness under noise interference, demonstrating its superior performance.
[0099] Module-by-module ablation experiment:
[0100] To verify the contribution of each module to the performance improvement, this embodiment conducted a step-by-step ablation experiment on the simulation dataset. First, this embodiment removed MRSC and NA from ResCNT-2stg to obtain the baseline model - 1. The results of the step-by-step ablation are as shown in Figure 6As shown in the figure. The PSNR of the baseline model - 1 is 33.34 dB and the SSIM is 0.952. After adding MRSC and NA respectively, the PSNR is improved by 0.41 dB and 0.22 dB respectively, and the SSIM is improved by 0.007 and 0.007 respectively. When using MRSC and NA jointly, the PSNR is improved by 0.41 dB and the SSIM is improved by 0.009. Whether used alone or jointly, the performance improvement proves the effectiveness of the MRSC and NA modules.
[0101] Network architecture ablation experiment:
[0102] To verify the advantages of the multi - residual skip connection (MRSC), this embodiment conducts ablation experiments on three different network architectures (U - shaped architecture, architecture with lateral skip connection (LSC), and architecture with multi - residual skip connection (MRSC)). The lateral skip connection is implemented through concatenation and conv1×1 operations. All models are retrained on the simulation dataset, and each model has two recovery stages. The experimental results are as Figure 7 shown. Compared with the U - shaped architecture, the network with lateral skip connection improves by 0.39 dB in PSNR, but has fewer additional parameters and computational costs. The MRSC architecture significantly exceeds by 0.80 dB in PSNR and has a lower additional computational cost, showing a higher performance improvement.
[0103] Comparison of self - attention mechanisms:
[0104] To verify the superiority of Neighborhood Attention (NA), this embodiment makes a comparison between NA and other MSA (multi - head self - attention mechanisms). The baseline model ResCNT - 1stg removes MRSC and NA. The experimental results are as Figure 8 shown. The PSNR of ResCNT - 1stg is 32.79 dB, the number of parameters is 0.40M, and the FLOPS is 6.85G. After adding G - MSA, SW - MSA, S - MSA, HS - MSA, and NA respectively, the PSNR is improved by 0.84, 0.96, 1.03, 1.26, and 1.73 dB respectively, and the computational cost increase of NA is the least, demonstrating its significant advantages. This superiority mainly stems from the window flexibility and translational equivalence of NA.
[0105] IV. Summary
[0106] By introducing the MRSC and NA modules, ResCNT can effectively improve the accuracy of hyperspectral image reconstruction, especially showing stronger robustness in the face of noise. Its innovative architecture design enables the network to extract information from multi-scale features and further enhance the expression of local features through the neighborhood attention mechanism. Experimental results show that ResCNT outperforms existing state-of-the-art methods on multiple datasets, demonstrating its effectiveness and potential in hyperspectral image reconstruction tasks.
[0107] In another embodiment of the present invention, a terminal device is provided. The terminal device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of the hyperspectral image reconstruction method, including the following steps:
[0108] Obtain compressed two-dimensional measurement values generated by an encoded aperture snapshot spectral imaging (CASSI) system;
[0109] Use a ResCNT (Residual Convolutional Network Transformer) network for reconstruction. The ResCNT network includes multiple convolutional layers and residual modules, combines the Transformer structure, and recovers the three-dimensional information of the hyperspectral image from the compressed two-dimensional measurement values through multi-scale learning;
[0110] During the reconstruction process, adopt multi-residual skip connections (MRSC) to improve the image detail recovery ability and enhance the multi-scale representation of the feature map; by introducing skip connections between different network layers, optimize the gradient path in backpropagation, thereby improving the convergence of the network;
[0111] The Neighborhood Attention (NA) module based on the self-attention mechanism dynamically assigns different weights to each pixel at each stage of image reconstruction, enhancing the consistency of details in the local area and global information;
[0112] The network is trained using the backpropagation algorithm. By minimizing the error between the reconstructed image and the real hyperspectral image, the network parameters are optimized, and finally the hyperspectral image is restored.
[0113] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is the memory device in the terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides storage space, and this storage space stores the operating system of the terminal. And, in this storage space, one or more instructions suitable for being loaded and executed by the processor are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory.
[0114] One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the hyperspectral image reconstruction method in the above embodiment; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps:
[0115] Obtain the compressed two-dimensional measurement values, which are generated by an encoded aperture snapshot spectral imaging (CASSI) system;
[0116] Use a ResCNT (Residual Convolutional Network Transformer) network for reconstruction. The ResCNT network includes multiple convolutional layers and residual modules, combines the Transformer structure, and recovers the three-dimensional information of the hyperspectral image from the compressed two-dimensional measurement values through multi-scale learning;
[0117] During the reconstruction process, multi-residual skip connections (MRSC) are adopted to enhance the detail recovery ability of the image and strengthen the multi-scale representation of the feature map; by introducing skip connections between different network layers, the gradient path in backpropagation is optimized, thereby improving the convergence of the network;
[0118] The Neighborhood Attention (NA) module based on the self-attention mechanism dynamically assigns different weights to each pixel at each stage of image reconstruction, enhancing the details of the local area and the consistency of global information;
[0119] The network is trained using the backpropagation algorithm. By minimizing the error between the reconstructed image and the real hyperspectral image, the network parameters are optimized, and finally the hyperspectral image is restored.
[0120] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0121] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0122] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0124] Those of ordinary skill in the art will realize that the embodiments described herein are to assist the reader in understanding the implementation methods of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not deviate from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A hyperspectral image reconstruction method based on ResCNT, characterized in that: The following steps are involved: acquiring compressed two-dimensional measurements generated by a coded aperture snapshot spectral imaging system; Reconstruction is performed using a ResCNT network, which includes multiple convolutional layers and residual modules, combined with a Transformer structure, to restore three-dimensional information of a hyperspectral image from compressed two-dimensional measurements through multi-scale learning; In the reconstruction process, multiple residual jump connections (MRSC) are used to improve the detail recovery ability of the image and enhance the multi-scale representation of the feature map. By introducing jump connections between different network layers, the gradient path in back propagation is optimized, thereby improving the convergence of the network. The Neighborhood Attention (NA) module based on the self-attention mechanism dynamically assigns different weights to each pixel at each stage of image reconstruction to improve the consistency of local area details and global information; The network is trained using the back-propagation algorithm to optimize the network parameters by minimizing the error between the reconstructed image and the true hyperspectral image, and finally restore the hyperspectral image.
2. The hyperspectral image reconstruction method according to claim 1, characterized in that: The method also includes: applying a degradation-aware unfolding framework DAUF in a denoising framework to unfold the compressed two-dimensional measurement values to improve the reconstruction effect.
3. The hyperspectral image reconstruction method according to claim 1, characterized in that: The specific structure of the ResCNT network includes: A convolutional encoder module to extract low-level image features from compressed 2D measurements, consisting of multiple convolutional layers, batch normalization, and activation functions; A residual module, which is used to enhance the transfer of deep features through skip connections and reduce the gradient vanishing problem in deep network training; A Transformer module that applies a self-attention mechanism in the network to enhance the long-range dependencies of the image by weighted averaging the features and modeling the global information of the features; A convolutional decoder module, which gradually recovers detailed information of the hyperspectral image after multiple deconvolution operations; An output layer that compares the reconstruction result with the true hyperspectral image by combining a pixel-level loss function to optimize the network output.
4. The hyperspectral image reconstruction method according to claim 3, characterized in that: The residual module avoids the attenuation of information in deep propagation by introducing multiple skip connections, while retaining detail information and ensuring the accuracy of image restoration.
5. The hyperspectral image reconstruction method according to claim 3, characterized in that: The Transformer module establishes dynamic relationships between features at different levels through a self-attention mechanism and weights different regions using adaptive weights, thereby optimizing global information transfer in the image.
6. The hyperspectral image reconstruction method according to claim 1, characterized in that: The NA module calculates the similarity between features using a sliding window around each pixel and dynamically assigns different attention weights to each pixel to enhance details and image consistency.
7. The hyperspectral image reconstruction method according to claim 1, characterized in that: The reconstruction process also includes using image prior information to perform denoising, further improving the quality of the reconstructed image and reducing the impact of noise during the compression process.
8. The hyperspectral image reconstruction method according to claim 1, characterized in that: The compressed two-dimensional measurement value is modulated by a physical mask of a spectrally coded aperture, shifted along the spectral dimension by a disperser, and finally obtained by integration.
9. The hyperspectral image reconstruction method according to claim 1, characterized in that: The hyperspectral image is obtained through the perception matrix The compressed two-dimensional measurements and noise recovery are combined to obtain the original three-dimensional hyperspectral image.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the hyperspectral image reconstruction method according to one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Hyperspectral reconstruction method based on self-attention and deep convolution parallelism
CN116665063A
Mask uncertainty hyperspectral image reconstruction method based on self-attention
CN117237432A
Image reconstruction method and system based on FCTFT
CN117557476A
Two-stage multi-scale hyperspectral snapshot compression imaging image reconstruction method
CN117974909A
Convolution and transformer based compressive sensing
US20240153161A1
Cited By
Multi-spatial resolution remote sensing simulation and error influence stripping method
CN120911142A
Multi-spatial resolution remote sensing simulation and error influence stripping method
CN120911142B