Transform-based hyperspectral image reconstruction method
By introducing a Transformer-based method in hyperspectral image reconstruction, combining FLA and hybrid rearrange rectangular attention mechanism, the challenges in computing efficiency and detail recovery in the prior art are solved, and efficient hyperspectral image reconstruction is achieved.
Patent Information
- Application Number
- CN202510439917.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-03
AI Technical Summary
Existing hyperspectral image reconstruction algorithms have challenges in computing efficiency and image detail recovery, especially in the case of insufficient light or rapid movement of objects.
The two-dimensional compressed observation data is reconstructed through the FHSRT module using a hyperspectral image reconstruction method based on Transformer, combined with the focused linear attention (FLA) and a hybrid rearranged rectangular attention mechanism.
It effectively improves the reconstruction quality and computing efficiency of hyperspectral images, can better restore image details and reduce noise, and adapt to the processing needs of large-scale image data.
Smart Images

Figure CN120088408A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hyperspectral image reconstruction, and particularly to a hyperspectral image reconstruction method based on Transformer. Background Art
[0002] Hyperspectral images (HSIs) are spatial-spectral data obtained through multiple narrow spectral bands, which can provide richer spectral information than ordinary RGB images. These information make hyperspectral images have important application values in the fields of medicine, agriculture, remote sensing, etc. However, the acquisition of hyperspectral images is limited by imaging time and hardware costs. To overcome these problems, a snapshot compressive imaging (SCI) system is proposed, which modulates spectral data through a physical mask and a prism, enabling multiple compressed spectral bands to be captured using a single two-dimensional sensor.
[0003] Although the SCI technology effectively solves the problem of hyperspectral image acquisition efficiency, due to the ill-posedness of the reconstruction problem, the current reconstruction algorithms still face challenges in terms of computational efficiency and image detail recovery. Although the existing deep learning-based reconstruction methods can improve the reconstruction performance, they have high computational complexity, especially in the case of insufficient illumination or fast object movement, and the detail recovery effect is not good. Therefore, how to improve the quality of image reconstruction on the premise of ensuring efficient calculation, especially in terms of noise and detail recovery, has become a key issue in this field. Summary of the Invention
[0004] Aiming at the defects of the prior art, the present invention provides a hyperspectral image reconstruction method based on Transformer. By combining the Focused Linear Attention (FLA) and the Hybrid Rearranged Rectangular Attention mechanism, the reconstruction quality and computational efficiency are effectively improved.
[0005] In order to achieve the above invention purposes, the technical solutions adopted by the present invention are as follows:
[0006] A hyperspectral image reconstruction method based on Transformer, comprising:
[0007] Performing a snapshot compressive imaging (SCI) operation on the hyperspectral image to obtain two-dimensional compressed observation data;
[0008] Based on the Transformer architecture, using the Focused Hybrid Rearranged Rectangular Transformer (FHSRT) module to reconstruct the compressed observation data;
[0009] The FHSRT module includes a Focused Hybrid Rearranged Rectangular Multi-Head Self-Attention (FHSR-MSA) module, and the Focused Hybrid Rearranged Rectangular Multi-Head Self-Attention (FHSR-MSA) module includes: a Focused Linear Attention (FLA) module and a Hybrid Rearranged Rectangular Self-Attention (HSRT) module, where:
[0010] The FLA module optimizes the attention weight distribution using a mapping function and enhances feature diversity through a rank recovery module;
[0011] The HSRT module effectively captures local and non-local features in the spatial domain through rectangular self-attention in the horizontal and vertical directions.
[0012] After the spatial structure and spectral information of the two-dimensional compressed observation data are processed by the FHSRT module, the output is a reconstructed hyperspectral image.
[0013] Furthermore, in local features, matrix multiplication is used to calculate the similarity between the query and the key, and the attention output is calculated in combination with the value vector;
[0014] In non-local features, the FLA module is used to enhance the focusing ability of the attention mechanism by adjusting the directions of the features of the query and the key, thereby improving the representation ability of the global context;
[0015] Finally, the outputs of the local and non-local features are weighted and fused to obtain the final Transformer output.
[0016] Furthermore, the FHSRT module is integrated through a Deep Unfolding Framework (DAUF).
[0017] Furthermore, the FLA module uses the following mapping function f p to achieve the direction adjustment of features:
[0018] Sim(Q i ,K j )=φ p (Q i )φ p (K j ) T
[0019] where, x + =ReLU(x), Q i is the i-th element in the query vector, K j is the j-th element in the key vector, φ p (Q i ) is the transformation function for the query vector Q i . φ p (K j) is the transformation function for the key vector K j , φ p (K j ), T and φ p (K j ) is the transpose of φ i (K j ). Sim(Q i , K j ) is the similarity score between the query vector Q
[0020] Further, the improvement of the FLA module solves the problem of feature diversity of linear attention by adding a local interaction module (DWC), and the output is expressed as:
[0021] FLA(Q, K, V) = Sim(Q, K)V = φ p (Q)φ p (K T )V + DWC(V)
[0022] Further, the rank recovery module in the FLA module adopts depthwise separable convolution, significantly reducing the computational complexity while improving the expressive ability of self-attention.
[0023] Further, the working process of the FHSR-MSA module is as follows:
[0024] Project the input feature map into queries, keys, and values through a linear transformation. This process maps the original features to different representation spaces for self-attention calculation.
[0025] Divide the queries, keys, and values into two parts along the channel dimension on average: one part is for the local branch, and the other part is for the non-local branch. The local branch is responsible for capturing local context information, while the non-local branch is used to model global dependencies.
[0026] In the local branch, use focused linear attention (FLA) to perform self-attention calculation on the input queries, keys, and values, focusing on feature interaction within the local region. The task of this branch is to capture the detailed information of each local region in the image. The local information is divided into multiple heads, and each head independently processes the local region.
[0027] In the non-local branch, first divide the queries, keys, and values into multiple non-overlapping windows, and perform a rotation operation on these windows. The rotation operation enables the tokens between different windows to exchange positions, thereby establishing dependencies between windows. This branch captures broader global features through hybrid re-arranged rectangle self-attention (HSR). The rotation operation effectively promotes global information interaction across windows.
[0028] In each self-attention head, the local and non-local branches process their corresponding queries, keys, and values respectively. Each head processes a smaller dimension, enabling the model to compute multiple attention patterns in parallel and capture different types of features.
[0029] The self-attention outputs of the local and non-local branches are weighted and summed respectively. The output of each head is weighted by the learned weights, and finally the local and global information is fused to obtain the final output of the model.
[0030] Compared with the prior art, the advantages of the present invention are as follows:
[0031] 1. By introducing the Focused Hybrid Rearranged Rectangle Transformer (FHSRT) module, the present invention can capture both local and non-local features of the image while reconstructing hyperspectral images, thereby effectively improving the accuracy of image reconstruction. When dealing with complex hyperspectral images, this method can improve the computational efficiency while ensuring the reconstruction quality, meeting the processing requirements of large-scale image data.
[0032] 2. By introducing the Rank Recovery Module, the present invention optimizes the feature representation ability of the self-attention mechanism, enabling the model to recover more high-rank information and enhance the diversity of features when processing high-dimensional and complex data. This makes the model more robust in dealing with various different image scenarios, avoiding problems such as information loss or over-simplification.
[0033] 3. Through the local and non-local branch processing of the input features, the FHSRT module can effectively capture the local details and global dependencies in the image. The local branch captures spatial local details through Focused Linear Attention (FLA), while the non-local branch captures long-range non-local dependencies through Hyperspectral Rectangle Self-Attention (HSRT). This efficient information fusion strategy ensures that neither details are ignored nor global information is wasted during the image reconstruction process, thus improving the performance of the model.
[0034] 4. The present invention adopts the rectangle self-attention mechanism in the FHSRT module and combines rotation operations to model non-local dependencies, significantly improving the computational efficiency of the self-attention mechanism on large-scale data. In particular, through branch processing and head division, the computational complexity is reduced, enabling the model to maintain high computational efficiency when processing large-size hyperspectral images.
[0035] 5. The present invention adopts the Focused Hybrid Rearranged Rectangular Multi-Head Self-Attention (FHSR-MSA) mechanism in the Transformer architecture, which can efficiently process large-scale hyperspectral image data. By dividing the image into multiple local and non-local processing branches, the model can process the information of different regions in parallel, thereby improving the overall computing efficiency and meeting the requirements of big data processing for hyperspectral images. Description of the Drawings
[0036] Figure 1 is the overall framework diagram of the embodiment of the present invention. Among them, (a) is EHQS, (b) is FHSRT, (c) is FHSRAB, (d) is FFN, (e) is FHSR-MSA,
[0037] Figure 2 is the result diagram of simulation reconstruction using 4 channels out of 28 spectral channels in the embodiment of the present invention.
[0038] Figure 3 is the comparison diagram of hyperspectral images reconstructed by FHSRT and existing methods in the embodiment of the present invention. Detailed Embodiment
[0039] To make the purpose, technical solutions and advantages of the present invention clearer, the following takes the drawings as examples and lists embodiments to further elaborate on the present invention in detail.
[0040] As Figure 1 shown, a hyperspectral image reconstruction method based on Transformer includes the following:
[0041] I. Define the compressed spectral image degradation model
[0042] Use the compressed spectral image degradation model to perform compressed observation on the original spectral image, where the compressed spectral image degradation model is:
[0043] y = Φx + g
[0044] where y represents the compressed observation image, Φ represents the sensing matrix, x represents the original spectral image signal, and g represents the imaging noise.
[0045] Perform hyperspectral image reconstruction by solving the optimization problem, and the optimization objective is:
[0046]
[0047] where, is the data fidelity term and r(x) is the regularization term, where λ is the balance coefficient.
[0048] The half - quadratic splitting (HQS) method is introduced, and the alternating direction method of multipliers (ADMM) is used to solve the above - mentioned optimization problem. An auxiliary variable z is introduced through the variable - splitting method, satisfying the constraint condition x = z. Then the above - mentioned optimization problem is rewritten as:
[0049]
[0050] The purpose of this penalty term is to force x and z to converge to the same iterative point. Then this formula is decomposed into the following two sub - problems for iterative solution:
[0051] (1) Solve for x:
[0052]
[0053] (2) Solve for z:
[0054]
[0055] where i = 0, 1, 2, …, K represents the number of iterations.
[0056] Note that the formula for solving x is a regularized least - squares problem, and it has a closed - form solution:
[0057] x i+1 =(Φ T Φ + τI) -1 (φ T y+τz i )
[0058] To simplify this term, since this formula satisfies the Woodbury matrix identity condition, in this embodiment, we have:
[0059] (A + UV T ) -1 =A -1 -A -1 U(I + V T A -1 U) -1 V T A -1
[0060] where and A + UV T is invertible. Let τI = A, U = V = Φ T , in this embodiment, we can simplify (Φ T Φ + τI) -1 to:
[0061]
[0062] From the model part, in this embodiment, we know that the matrix Φ has the following properties:
[0063] ΦΦ T ≈ diag(σ 1 , σ 2 , …, σ t )
[0064] where σ i is the i-th diagonal element of the diagonal matrix. Substituting this into the above formula, this embodiment can derive:
[0065]
[0066] Furthermore, substituting this into the update formula of x, this embodiment can obtain the update operation of x as:
[0067]
[0068] where [Φz i k represents the k-th term of the vector Φz i . Solving the equation of z can be regarded as a denoising problem based on x i+1 . Generally, a denoiser D τ / λ is designed to solve this problem, where τ / λ is related to the noise amplitude.
[0069] II. Constructing the EHQS Network Framework
[0070] As shown in Figure 1 (a), the spectral compressive image reconstruction network architecture in the EHQS expansion framework includes: an initialization network I composed of 1×1 convolutional networks, and a parameter estimator composed of 1×1 convolution, 3×3 stride convolution, global average pooling layer, and three fully connected layers. The measurement value y and the sensing matrix Φ are processed by the initialization network and the parameter estimator, respectively, to obtain the initialization result z 0 and the estimated parameters α and β. The latter captures the key features of the CASSI system by learning the degradation patterns and ill-conditioning degrees caused by mask modulation and dispersion integration, so as to provide noise level information for the denoising prior to guide the iterative learning process. The linear projection layer follows the program flow of the HQS algorithm expansion framework.
[0071] To address the challenges of SCI reconstruction, this embodiment develops the FHSRT module, where the hybrid-random rectangular attention and focused linear attention can be well applied to image reconstruction tasks such as SCI.
[0072] III. Designing the Focused Linear Attention Mechanism
[0073] Modern vision transformers mainly adopt the Softmax attention mechanism, which usually leads to high computational costs. Although the linear attention mechanism has the advantage of linear computational complexity, there is often a significant performance degradation when replacing the Softmax attention. In this section, this embodiment analyzes the reasons for these performance issues, with particular attention to the lack of focusing ability and feature diversity. To this end, this embodiment proposes a Focused Linear Attention (FLA) mechanism to address these deficiencies, achieving high efficiency and enhanced expressive power.
[0074] An important advantage of Softmax attention is that it provides a non-linear reweighting mechanism, which enables it to easily concentrate the attention distribution on key features. Without Softmax, the distribution of linear attention is relatively smooth, and its result is closer to the average of all features, resulting in a significant loss of accuracy.
[0075] The key to solving this problem lies in how to make similar query-key value pairs closer while pushing away dissimilar query-key value pairs. As a remedy, this embodiment proposes a simple and effective solution by adjusting the directions of each query and key feature. Specifically, this embodiment considers a simple mapping function f p :
[0076] Sim(Q i , K j ) = φ p (Q i ) φ p (K j ) T
[0077] where, x + = ReLU(x), x **p represents the element-wise p of x. This modification only adjusts the direction of the feature without changing its magnitude.
[0078] Under certain reasonable assumptions, the mapping function f p will affect the attention distribution.
[0079] Proposition 1 (Feature direction adjustment using f p ) Let
[0080] Assume that x and y have a single maximum value x m and y n .
[0081] For a pair of features {x, y}, when m = n:
[0082]
[0083] For a pair of features {x, y}, when m ≠ n:
[0084]
[0085] Therefore, with an appropriate p value, the mapping function f p (·) can effectively distinguish between similar query-key pairs and dissimilar query-key pairs, thus enhancing the attention distribution, similar to the effect of the original Softmax function.
[0086] In addition to the focusing ability, feature diversity is also one of the factors limiting the expressive power of linear attention. One possible reason may be attributed to the rank of the attention matrix.
[0087] However, this is difficult to achieve in the case of linear attention. In fact, the rank of the attention matrix in linear attention is restricted by the number of tokens N and the channel dimension d of each head:
[0088] rank(φ(Q)φ(K) T ) ≤ min{rank(φ(Q)), rank(φ(N))} ≤ min{N, d}
[0089] Among common vision Transformer designs, d is usually smaller than N. For example, in DeiT, d = 64, N = 196; in Swin Transformer, d = 32, N = 49; in the hyperspectral reconstruction scenario, d = 28, N = 256. In this case, the upper bound of the rank of the attention matrix is restricted to a lower ratio, indicating that many rows of the attention map converge severely. Since the output of self-attention is a weighted sum of the same set of V, the convergence of attention weights inevitably leads to the similarity of aggregated features.
[0090] As a remedy, this embodiment proposes a simple and effective solution to address this limitation of linear attention. Specifically, a depthwise convolution (DWC) module is added to the attention matrix, and the output can be expressed as:
[0091] FLA(Q, K, V) = Sim(Q, K)V = φ p (Q)φ p (K) T V + DWC(V),
[0092] where DWC(V) increases local interaction. This method increases the rank potential of the attention matrix and maintains feature diversity similar to Softmax attention.
[0093] The focus linear attention module of this embodiment is designed as a plug-and-play replacement for the traditional attention mechanism in the Transformer architecture, which balances computational efficiency and expressive power:
[0094] FLA(Q, K, V) = Sim(Q, K)V = φ p (Q)φ p (K) T V + DWC(V),
[0095] which provides improvements in both the sharpness of attention and the diversity of features.
[0096] IV. Focus Hybrid Rearranged Rectangular Multi-Head Self-Attention (FHSR-MSA)
[0097] As shown in Figure 1 (e) the upper path. The HS-MSA model first projects the input feature X in into queries Q, keys K, and values V:
[0098] Q = X in W Q , K = X in W K , V = X in W V ,
[0099] where W Q , W K , W V are learnable parameters, and the bias terms are omitted for simplicity. Next, the model divides these projections into two branches for different processing strategies.
[0100] Specifically, Q, K, V are evenly divided into two parts along the channel dimension:
[0101] Q = [Q l , Q nl , K = [K l , K nl , V = [V l , V nl ,
[0102] where is input to the local branch to capture local content, while Q nl , K nl , model non-local dependencies through the non-local branch.
[0103] The local branch focuses on capturing context details within a specific spatial window, while the non-local branch enables interaction between windows through rotation operations.
[0104] In each head i, the local branch adopts focused linear attention (FLA), which is calculated as follows:
[0105]
[0106] where is a single head, which is divided into h heads along the channel dimension by Q l , K l , V l . The dimension of each head is Note that Figure 1 (d) shows the case of h = 1, and some details are omitted for simplicity.
[0107] The non-local branch calculates the interaction between windows through a rotation operation inspired by ShuffleNet. Specifically, first, Q nl , K nl , are divided into non-overlapping windows of size M×M. Then their shapes are transposed from to to rotate the positions of the tokens and establish dependencies between windows. Q nl , K nl , V nl are divided into h heads: and Then the non-local self-attention is calculated for each head
[0108]
[0109] Subsequently, the rotation is cancelled by transposing to the shape .
[0110] Finally, the outputs of the two branches are aggregated:
[0111]
[0112] This aggregation effectively integrates the local and global dependencies required for the reconstruction task, demonstrating the computational efficiency and effectiveness of the FHSR-MSA model.
[0113] V. Developing the FHSRT network architecture:
[0114] Figure 1 (b) shows the architecture framework of the Focus Hybrid ShuffleRectangle Transformer (FHSRT), which adopts a UNet-based structure. In the k-th iteration, the input contains xk and the splicing with the constant image constructed by β k This spliced input undergoes a 3×3 convolution operation and is mapped to Subsequently, z in is successively processed through an encoder, a bottleneck layer, and a decoder architecture.
[0115] The encoder and decoder structures consist of multiple levels, and each level contains an FHSRT block integrating a downsampling or upsampling module. As shown in Figure 1 (c), the Focus Hybrid ShuffleRectangle Attention Block (FHSRAB) contains two layer normalization operations, a Focus Hybrid Shuffle Rectangle multi-head self-attention (FHSR-MSA) mechanism, and a feed-forward network (FFN).
[0116] For the downsampling operation in the encoder, a 4×4 convolution with a stride of 2 is adopted, while the decoder uses a 2×2 transposed convolution for upsampling. The final output z out is added element-wise to x k to generate the final output z of the denoising layer k+1 .
[0117] This complete architecture demonstrates the integration of modern Transformer mechanisms and classical convolution operations, achieving efficient feature extraction and reconstruction at multiple scales. This delicate design promotes the effective information flow and feature transformation throughout the network levels.
[0118] VI. Experiments:
[0119] 1) Experimental settings:
[0120] In this embodiment, 28 wavelengths are selected for hyperspectral image (HSI) data in the wavelength range of 450 nm to 650 nm through spectral interpolation. The experiment simultaneously includes the testing of simulation datasets and real datasets to verify the effectiveness of the proposed method.
[0121] 2) Simulation datasets:
[0122] In this embodiment, two simulation datasets are used: CAVE and KAIST. The CAVE dataset contains 32 hyperspectral images (HSIs) with a size of 512×512 pixels, while the KAIST dataset contains 30 HSIs, each with a size of 2704×3376 pixels. According to the configuration in MST, the CAVE dataset is used for training, and 10 scenes selected from the KAIST dataset are used for testing.
[0123] 3) Real datasets
[0124] For the real tests, five HSIs obtained by the CASSI system are used in this embodiment.
[0125] Implementation details:
[0126] The FHSRT model is implemented using Pytorch. In this embodiment, the Adam optimizer (parameters β 1 = 0.9 and β 2 = 0.999) is used to train the FHSRT model, and the cosine annealing schedule is adopted when running for 300 epochs on an RTX 4090 GPU. The initial learning rate is set to 4×10 -4 . For the simulation and real tests, in this embodiment, image patches of sizes 256×256 and 660×660 are randomly extracted from a 3D HSI cube with 28 channels for training. The color step size d is set to 2. The batch size is kept as 2, and in this embodiment, the basic number of channels C = N λ = 28 is used for the HSI data. The weights of D are different in each stage. The data augmentation strategy includes random rotation and flipping. The objective is to minimize the root mean square error (RMSE) between the reconstructed HSI and the actual HSI.
[0127] In the simulation scenario, the performance of the FHSRT method of this embodiment is compared with that of 9 state-of-the-art methods. The algorithms compared include 4 CNN methods, 4 Transformer methods, 1 Fourier method, and 1 RNN method. All algorithms are tested using the same settings as the MST method.
[0128] The experimental results show that the best model FHSRT-9stg (9-stage FHSRT) of this embodiment achieves a PSNR of 38.50 dB and an SSIM of 0.972, which is better than other competing methods. Compared with the best CNN method, FHSRT improves by 3.53 dB in PSNR and 0.039 in SSIM. Compared with the state-of-the-art RNN method, the method of this embodiment improves by 0.92 dB in PSNR and 0.012 in SSIM. In addition, compared with the state-of-the-art Transformer method, the FHSRT method of this embodiment improves by 0.14 dB in PSNR and 0.005 in SSIM, while only using 64.16% of the number of parameters. The experiment also shows that the method of this embodiment is significantly lower than other competing methods in terms of computational cost (FLOPs).
[0129] 4) Qualitative results
[0130] Figure 2The results of simulation reconstruction using 4 out of 28 spectral channels are shown. By magnifying the image, more details can be seen. The FHSRT model can better fit the spectral feature curve during the reconstruction process, showing less detail loss. Compared with other methods, the FHSRT model exhibits clearer edge details and textures in each spectral channel.
[0131] 5) Real HSI reconstruction:
[0132] Figure 3 The hyperspectral images reconstructed by the FHSRT method and existing methods are shown. The results indicate that the FHSRT method of this embodiment can better restore the high-frequency details of the image, reduce noise and improve image clarity.
[0133] 6) Ablation experiments:
[0134] Multiple ablation experiments were conducted in this embodiment to verify the impact of each component and design choice on performance.
[0135] Table 1: Performance comparison of different loss functions
[0136]
[0137]
[0138] Table 2: Ablation experiments of each module
[0139]
[0140] Table 3: Ablation experiments of various self-attention mechanisms
[0141]
[0142] To verify the effectiveness of the focused linear attention and the rectangular window, this embodiment starts from the baseline model - 1, which is obtained by removing the focused linear attention and the rectangular window from DAUHST - 2stg, for conducting decomposition ablation experiments. As shown in Table 2, introducing the focused linear attention of this embodiment reduces the FLOPs from 18.44G to 15.07G, indicating a significant reduction in computational complexity, while the PSNR value slightly increases by 0.04. With the same computational complexity maintained, introducing the rectangular window further improves the PSNR by 0.05, finally reaching 36.43. The findings of this embodiment strongly indicate that the focused linear attention and the rectangular window significantly enhance the expressive ability of the linear attention, thereby improving the efficiency and performance of the FHSRRT module.
[0143] To verify the effectiveness of the attention module, in this embodiment, an ablation study was conducted using Baseline Model-2 (obtained by excluding FHSRT from FHSRT-1stg), as shown in Table 2. In this embodiment, different positional embedding strategies were removed to eliminate their influence, and only the attention module was focused on for comparison. To ensure fairness, the number of parameters of the attention module was maintained by keeping the number of channels and the number of heads unchanged. Baseline Model-2 achieved 32.79 dB. In this embodiment, Swin MSA (SW-MSA), Spectral-wise MSA (S-MSA), Hybrid-Shuffle MSA (HS-MSA), and FHSRT were evaluated. According to Table 3, FHSRT showed the most significant improvement, increasing by 1.48 dB, which was 0.52, 0.45, and 0.22 dB higher than SW-MSA, S-MSA, and HS-MSA, respectively.
[0144] To verify the impact of the loss function on the reconstruction quality, in this embodiment, the model proposed in this embodiment was retrained using the Least Absolute Deviation (LAD) loss function. As shown in Table 2, using the LAD loss function can further improve the reconstruction quality.
[0145] VII. Conclusion:
[0146] By combining the advantages of focused linear attention and hybrid-rearranged rectangular windows, the present invention proposes a module called FHSRT, thus effectively solving the limitations of existing Transformer methods in terms of computational efficiency and spatial region selection. Through a large number of experimental verifications, the FHSRT method in this embodiment not only surpasses the existing state-of-the-art technologies in terms of reconstruction quality, but also has higher efficiency in terms of memory and computational resource consumption. The method of the present invention demonstrates excellent performance in the hyperspectral image reconstruction task, especially when dealing with large-scale data, and has strong practicality and broad application prospects.
[0147] The method according to the present invention described above can be implemented in hardware, firmware, or can be implemented as software or computer code stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or can be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a RAM, a ROM, a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, a Transformer-based hyperspectral image reconstruction method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the processing shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the processing shown herein.
[0148] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the implementation methods of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A hyperspectral image reconstruction method based on Transformer, characterized in that: include: Perform snapshot compression imaging operation on the hyperspectral image to obtain two-dimensional compressed observation data; Based on the Transformer architecture, the compressed observation data is reconstructed using a focused hybrid rearranged rectangular FHSRT module; The FHSRT module includes a focused hybrid rearranged rectangular multi-head self-attention FHSR-MSA module, and the FHSR-MSA module includes: a focused linear attention FLA module and a hybrid rearranged rectangular self-attention HSRT module, wherein: The FLA module uses a mapping function to optimize the attention weight distribution and enhances feature diversity through a rank recovery module; The HSRT module effectively captures local and non-local features in the spatial domain through rectangular self-attention in horizontal and vertical directions; The spatial structure and spectral information of the two-dimensional compressed observation data are processed by the FHSRT module and output as a reconstructed hyperspectral image.
2. The hyperspectral image reconstruction method according to claim 1, characterized in that: The data collection specifically includes: in the local features, the similarity between the query and the key is calculated using matrix multiplication, and the attention output is calculated in combination with the value vector; In the non-local features, the FLA module is used to enhance the focusing ability of the attention mechanism by adjusting the direction of the query and key features, thereby improving the representation ability of the global context; Finally, the outputs of local and non-local features are weighted fused to obtain the final Transformer output.
3. The hyperspectral image reconstruction method according to claim 1, characterized in that: The FHSRT module is integrated through a depth unfolding framework.
4. The hyperspectral image reconstruction method according to claim 1, characterized in that: The FLA module uses the following mapping function f p To achieve feature orientation adjustment: Sim(Q i ,K j )=φ p (Q i )φ p (K j ) T in, x + =ReLU(x),Q i is the i-th element in the query vector, K j is the jth element in the key vector, φ p (Q i ) is the query vector Q i The transformation function, φ p (K j ) is the key vector K j The transformation function, φ p (K j ) T Yes p (K j ), Sim(Q i ,K j ) is the query vector Q i and the key vector K j The similarity score between .
5. The hyperspectral image reconstruction method according to claim 4, characterized in that: The improvement of the FLA module solves the feature diversity problem of linear attention by adding a local interaction module, and the output is expressed as: FLA(Q,K,V)=Sim(Q,K)V=φ p (Q)φ p (K) T V+DWC(V) Among them, Q, K, and V are the three core vectors in Transformer: query, key, and value, and DWC(V) represents a deep convolution operation used to enhance the feature interaction of V.
6. The hyperspectral image reconstruction method according to claim 1, characterized in that: The rank recovery module in the FLA module adopts depthwise separable convolution, which significantly reduces the computational complexity while improving the expressive power of self-attention.
7. The hyperspectral image reconstruction method according to claim 1, characterized in that: The workflow of the FHSR-MSA module is as follows: The input feature map is projected into query, key and value through linear transformation; this process maps the original features into a different representation space for self-attention calculation; The query, key, and value are equally divided into two parts along the channel dimension: one part is used for the local branch and the other part is used for the non-local branch; the local branch is responsible for capturing local context information, while the non-local branch is used to model global dependencies; In the local branch, FLA is used to perform self-attention calculations on the input query, key, and value, focusing on the feature interactions within the local area; the task of this branch is to capture the detailed information of each local area in the image; the local information is divided into multiple heads, and each head processes the local area independently; In the non-local branch, the query, key, and value are first divided into multiple non-overlapping windows, and rotation operations are performed on these windows; the rotation operation enables tokens between different windows to exchange positions, thereby establishing dependencies between windows; this branch captures a wider range of global features by mixing and rearranging rectangular self-attention; the rotation operation promotes global information interaction across windows; In each self-attention head, the local and non-local branches process the corresponding query, key, and value respectively; the small dimension processed by each head enables the model to calculate multiple attention modes in parallel and capture different types of features; The self-attention outputs of the local branch and the non-local branch are weighted and summed separately; the output of each head is weighted by the learned weights, and finally the local and global information are fused to obtain the final output of the model.
Citation Information
Patent Citations
Hyperspectral reconstruction method based on self-attention and deep convolution parallelism
CN116665063A
Two-stage multi-scale hyperspectral snapshot compression imaging image reconstruction method
CN117974909A
A laundry treating apparatus
KR1020260059840A
Joint modeling method and apparatus for enhancing local features of pedestrians
US11810366B1
Hyperspectral remote sensing image classification method based on self-attention context network
US20230260279A1
Cited By
Hyperspectral image reconstruction method and system based on flow matching prior assistance
CN121120819A
BERT model and linear attention-based intention recognition method and system
CN121256466A