Transformer-Based Cross-Modal Fusion Hyperspectral Image Super-Resolution Reconstruction Method
Through the transmodal fusion method based on transformer, the feature reconstruction of hyperspectral images is performed using the visual transformer encoder and the Quasi-Recurrent convolution unit, which solves the problem of insufficient information utilization in hyperspectral image super-segment reconstruction, and achieves efficient improvement of spatial resolution of hyperspectral images.
Patent Information
- Application Number
- CN202410803466.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-06-20
AI Technical Summary
The prior art is difficult to effectively utilize the hyperspectral image information with low spatial resolution in hyperspectral image super-segment reconstruction, and the existing cross-modal fusion methods require strict spatial and temporal registration of data and are difficult to obtain.
Transformer-based cross-modal fusion method is adopted, and overlap-free slice splicing is performed through the embedding layer and position coding is added. Feature extraction is used to use the visual transformer encoder, and feature reconstruction is carried out in combination with the self-attention mechanism and the Quasi-Recurrent convolution unit to reduce the requirements for space-time registration.
High spatial resolution reconstruction of hyperspectral images is realized, which can better utilize global spectrum cross-correlation information, reduce computing costs, and adapt to the fusion and enhancement of hyperspectral multi-band data.
Smart Images

Figure CN118628357B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cross-modal data fusion and remote sensing image data enhancement, and particularly relates to a method for super-resolution reconstruction of cross-modal fusion hyperspectral images based on transformers. Background Art
[0002] Deep learning methods based on convolutional neural networks have been widely applied to remote sensing image vision tasks, such as object tracking, object detection, object segmentation, etc. based on hyperspectral images. However, since such technologies rely to a large extent on the spatial features of images, and the spatial resolution of hyperspectral images is usually low, it is difficult for deep learning algorithms to obtain sufficient information. There are some existing methods that fuse RGB images and hyperspectral images to improve the spatial resolution of hyperspectral images. Usually, there are strict spatio-temporal registration requirements for the two-modal data, and it is very difficult to obtain such data. Summary of the Invention
[0003] To solve the above technical problems existing in the prior art, the purpose of the present invention is to provide a method for super-resolution reconstruction of cross-modal fusion hyperspectral images based on transformers. The data of the two modalities are spliced into non-overlapping slices and position encoding is added through an embedding layer. The visual transformer encoder extracts features from the spliced sequence. The self-attention mechanism realizes the self-enhancement of single-modal data features and the information interaction between cross-modalities. After three layers of Quasi-Recurrent convolutional units and one layer of bidirectional Quasi-Recurrent convolutional units for the reconstruction of hyperspectral images, it is more suitable for hyperspectral multi-band data and can better utilize the global spectral cross-correlation information.
[0004] To achieve the above object of the invention, the present invention provides a method for super-resolution reconstruction of cross-modal fusion hyperspectral images based on transformers, including the following steps:
[0005] Step S1: Obtain a fused token sequence based on the RGB image and the hyperspectral image that match the scene;
[0006] Step S2: Adopt a random masking mechanism to jointly train the encoder and decoder of the visual transformer to complete feature reconstruction;
[0007] Step S3: Construct a loss function based on the similarity between the image reconstructed by the decoder and the ground truth image, and optimize the encoder and decoder of the visual transformer;
[0008] Step S4: Use the image reconstruction module to perform super-resolution reconstruction on the features extracted by the encoder of the optimized visual transformer.
[0009] According to a technical solution of the present invention, in the step S1, it specifically includes:
[0010] Step S11: Input the RGB image and the hyperspectral image matching the scene, and preprocess the RGB image and the hyperspectral image respectively to obtain the image block sequences corresponding to the two modalities, including the RGB modality image block sequence and the hyperspectral modality image block sequence.
[0011] The preprocessing process at least includes non-overlapping segmentation and adding position encoding.
[0012] Step S12: Perform linear mapping on the RGB modality image block sequence and the hyperspectral modality image block sequence respectively to obtain token vectors, and then splice the token vectors of the two modalities to obtain a fused token sequence.
[0013] According to a technical solution of the present invention, in the step S2, it specifically includes:
[0014] Step S21: In the training stage, use the preset parameter k as the masking rate to randomly mask the fused block sequence. The visible blocks are embedded by linear mapping, and the masked part uses a learnable masked token.
[0015] Step S22: Input the unmasked token sequence into the encoder of the vision transformer for feature extraction.
[0016] Step S23: Input the masked part and the enhanced features obtained by the encoder of the vision transformer into the decoder of the vision transformer, obtain the relative position of the mask according to the position encoding, and use the decoder of the vision transformer to reconstruct the features extracted in the step S22.
[0017] According to a technical solution of the present invention, in the step S3, it specifically includes:
[0018] In the training stage, use the L1 function to calculate the reconstruction error to optimize the decoder and the encoder simultaneously, and construct the loss function as:
[0019] L HSI = ||F d (F e (unmasked(I fuse )) + masked(I fuse )) - I RGB || 1
[0020] L RGB = ||F d (Fe (unmasked(I fuse ))+masked(I fuse ))-I RGB || 1
[0021] Among them, LHSI represents the loss function for the decoder to reconstruct the hyperspectral features generated by the encoder, LRGB represents the loss function for the decoder to reconstruct the visual features generated by the encoder, Ifuse represents the concatenated hyperspectral-RGB fusion token sequence, Fe(·) represents the encoder function, and Fd(·) represents the decoder function.
[0022] According to a technical solution of the present invention, in step S4, the image reconstruction module is composed of 3 layers of Quasi-Recurrent convolutional units and 1 layer of bidirectional Quasi-Recurrent convolutional units, including:
[0023] In each Quasi-Recurrent convolutional unit, first, two 3D convolutional kernels W z and W f perform two 3D convolutions and then obtain two sets of feature maps Z and F through different activation functions. Then:
[0024] Z = tanh(W z *I), F = sigmoid(W f *I)
[0025] Split the feature map Z and the feature map F in the spectral direction to obtain the z i and f i sequences, input them into the Quasi-Recurrent pooling layer to obtain the fused feature h i , and re-concatenate the fused feature in the spectrum to obtain the reconstructed feature. Among them, the operation in the Quasi-Recurrent pooling layer can be expressed as:
[0026] h i = f i ·h i-1 +(1 - f i )·z i
[0027] In the last layer of bidirectional Quasi-Recurrent convolutional units, for all sequences split in the spectrum, perform Quasi-Recurrent pooling in an alternating forward and backward manner to complete the propagation of global information.
[0028] According to a technical solution of the present invention, in step S4, it specifically includes:
[0029] In the inference stage, the masking process in step S21 is not performed, and the fused token sequence is directly input into the encoder of the trained vision transformer for feature extraction;
[0030] Then, an image reconstruction module is used to perform super-resolution reconstruction on the features extracted by the encoder.
[0031] According to a technical solution of the present invention, in step S22, the encoder of the vision transformer is provided with a total of 4 feature encoding layers, including:
[0032] In any feature encoding layer, first, the fused feature vector is normalized and then used as the Query, Key, and Value matrices for multi-head attention calculation;
[0033] Inside the multi-head attention mechanism, after transposing the fused features and performing matrix multiplication with itself in each head, an attention map is obtained through softmax:
[0034]
[0035] where β i,j represents whether the i-th position should be associated with the j-th position, and each head learns a mapping method h n (·), then the complete output of the attention mechanism of N heads is:
[0036]
[0037] According to an aspect of the present invention, there is provided an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes a super-resolution reconstruction method for cross-modal fusion hyperspectral images as described in any one of the above technical solutions.
[0038] According to an aspect of the present invention, there is provided a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement a super-resolution reconstruction method for cross-modal fusion hyperspectral images as described in any one of the above technical solutions.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] The present invention proposes a cross-modal fusion hyperspectral image super-resolution reconstruction method based on a transformer. The data of two modalities are spliced without overlapping slices and added with positional encoding through an embedding layer, and the spliced sequence is feature-extracted through a vision transformer encoder. The self-attention mechanism is used to realize the self-enhancement of single-modal data features and the information interaction between cross-modalities. The reconstruction of the hyperspectral image is carried out through three layers of Quasi-Recurrent convolutional units and one layer of bidirectional Quasi-Recurrent convolutional units, which is more suitable for hyperspectral multi-band data and can better utilize the global spectral cross-correlation information.
[0041] An encoder based on a vision transformer is adopted, taking advantage of the global receptive field of the attention mechanism, reducing the strong spatio-temporal registration requirements for RGB data and hyperspectral data, and enabling the network to adaptively fuse and enhance cross-modal data. In the training stage, a random masking training mechanism is adopted, enabling the network to learn the key features in the image faster and better while reducing the computational cost.
[0042] The reconstruction of the fused features is realized by adopting a deep learning image reconstruction module designed based on Quasi-Recurrent convolutional units to obtain a hyperspectral image with high spatial resolution. 3D convolution is adopted in the Quasi-Recurrent convolutional units, which is more suitable for hyperspectral multi-band data than ordinary convolution. The Quasi-Recurrent pooling operation adopted therein can better utilize the global spectral cross-correlation information, and the additional bidirectional Quasi-Recurrent convolutional unit can realize the synchronous information propagation of the front and rear time series with less computing power. Brief Description of the Drawings
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0044] Figure 1 Schematically showing the flowchart of the training stage of the cross-modal fusion hyperspectral image super-resolution reconstruction method based on a transformer according to an embodiment of the present invention;
[0045] Figure 2 Schematically showing the flowchart of the inference stage of the cross-modal fusion hyperspectral image super-resolution reconstruction method based on a transformer according to an embodiment of the present invention;
[0046] Figure 3 Schematic diagram showing the working principle of a visual transformer encoder according to an embodiment of the present invention;
[0047] Figure 4 Schematic diagram showing the structural diagram of an image reconstruction module according to an embodiment of the present invention;
[0048] Figure 5 Schematic diagram showing the overall flowchart of a transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method according to an embodiment of the present invention. Detailed implementation manners
[0049] The description of the embodiments of this specification should be combined with the corresponding drawings, which should be part of the complete specification. In the drawings, the shape or thickness of the embodiments may be enlarged and simplified or conveniently marked. Furthermore, the parts of each structure in the drawings will be described separately. It should be noted that the elements not shown or described in words in the drawings are in the forms known to those of ordinary skill in the art.
[0050] Any reference to directions and orientations in the description of the embodiments herein is for convenience of description only and should not be construed as any limitation to the protection scope of the present invention. The following description of the preferred embodiments involves combinations of features, which may exist independently or in combination. The present invention is not particularly limited to the preferred embodiments. The scope of the present invention is defined by the claims.
[0051] As Figure 1 、 Figure 2 and Figure 5 shown, a transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method of the present invention includes the following steps:
[0052] Step S1: Based on the RGB image and the hyperspectral image that are scene-matched, obtain a fused token sequence, specifically including:
[0053] Step S11: Input the RGB image and the hyperspectral image that are scene-matched. The RGB image is a visible light three-band image with a high spatial resolution. The hyperspectral image contains a wide spectral band range and has a high resolution in the spectral dimension, but a low resolution in the spatial dimension. Preprocessing operations such as slicing and adding position encoding are performed on the two images simultaneously to obtain image block sequences corresponding to the two modalities;
[0054] Step S111: Perform non-overlapping segmentation and slicing on the RGB data. The size of each slice is p*p, and the number of channels remains unchanged. Add position encoding;
[0055] Step S112: Segment and cut the hyperspectral data into non-overlapping blocks, each block with a size of p*p, keeping the number of channels unchanged, and adding positional encoding.
[0056] Step S12: Perform linear mapping embedding operations on the RGB modality image block sequence and the hyperspectral modality image block sequence in different ways. Among them, the number of channels after mapping the RGB modality data remains unchanged, while the number of channels after mapping the hyperspectral modality data is made the same as the number of channels of the RGB modality data for the next concatenation operation. Concatenate the two modality token vectors obtained along the channel dimension to obtain a fused token sequence.
[0057] Step S2: Use a random masking mechanism to jointly train the encoder and decoder of the vision transformer to complete feature reconstruction, specifically including:
[0058] Step S21: Use a random masking mechanism to train the vision transformer in the network. During the training phase, use the pre-set parameter k as the masking rate to randomly mask the fused block sequence. The visible blocks use the linear mapping mentioned in Step S12 to embed different modality data respectively, and the masked part directly uses learnable masking tokens. Concatenate the unmasked token vectors along the channel dimension.
[0059] Step S22: Input the unmasked token sequence into the encoder of the vision transformer for feature extraction. The self-attention mechanism in the encoder simultaneously realizes information interaction between the two modalities. Therefore, only a single pipeline working mechanism is needed to achieve cross-modal feature fusion, without the need for additional fusion processing after parallel feature extraction.
[0060] Construct a feature encoder based on the attention mechanism as Figure 3 shown, using the pre-trained weights of ViT, with a total of 4 feature encoding layers.
[0061] In each layer, first normalize the fused feature vector and then use it as the Query, Key, and Value matrices for multi-head attention calculation to achieve self-attention operation, and at the same time realize cross-modal feature interaction in self-attention.
[0062] Inside the multi-head attention mechanism, transpose the fused feature and perform matrix multiplication with itself in each head, and obtain the attention map through softmax:
[0063]
[0064] Among them, β i,jIndicates whether the i-th position should be associated with the j-th position. Each head can learn a mapping method h on its own. n (·). The complete output of the attention mechanism of N heads is:
[0065]
[0066] Step S23: Input the masked part and the enhanced features obtained by the encoder of the vision Transformer into the decoder of the vision Transformer. The structure of the decoder is similar to that of the encoder. Obtain the relative position of the mask according to the position encoding, and use the decoder of the vision Transformer to reconstruct the features extracted in step S22.
[0067] Step S3: In the training stage, construct a loss function based on the similarity between the image reconstructed by the decoder and the ground truth image. Use the input image as the ground truth image for each modality, and optimize the encoder and decoder of the vision Transformer in steps S22 and S23, specifically including:
[0068] In the training stage, use the L1 function to calculate the reconstruction error and optimize the decoder and encoder simultaneously. The constructed loss function is:
[0069] L HSI =||F d (F e (unmasked(I fuse )) + masked(I fuse )) - I RGB || 1
[0070] L RGB =||F d (F e (unmasked(I fuse )) + masked(I fuse )) - I RGB || 1
[0071] Among them, LHSI represents the loss function for the decoder to reconstruct the hyperspectral features generated by the encoder, LRGB represents the loss function for the decoder to reconstruct the visual features generated by the encoder, Ifuse represents the concatenated hyperspectral-RGB fusion token sequence, Fe(·) represents the encoder function, and Fd(·) represents the decoder function.
[0072] Step S4: Use the image reconstruction module to perform super-resolution reconstruction on the features extracted by the optimized encoder of the vision Transformer, specifically including:
[0073] In the inference stage, the masking process in step S21 is not performed, and the fused token sequence is directly input into the encoder of the trained vision transformer for feature extraction;
[0074] Then, the image reconstruction module is used to perform super-resolution reconstruction on the features extracted by the encoder.
[0075] In the training stage, an image reconstruction module with the structure as Figure 4 shown is constructed, and pre-trained weights are adopted. The module consists of 3 layers of Quasi-Recurrent convolutional units and 1 layer of bidirectional Quasi-Recurrent convolutional units;
[0076] In the inference stage, in each Quasi-Recurrent convolutional unit, first, two 3D convolutional kernels W z and W f perform two 3D convolutions, and then two sets of feature maps Z and F are obtained through different activation functions. Then, we have:
[0077] Z = tanh(W z *I), F = sigmoid(W f *I)
[0078] The feature map Z and the feature map F are split in the spectral direction to obtain the z i and f i sequences, which are input into the Quasi-Recurrent pooling layer to obtain the fused feature h i , and the fused feature is re-spliced in the spectrum to obtain the reconstructed feature. Among them, the operation in the Quasi-Recurrent pooling layer can be expressed as:
[0079] h i = f i ·h i-1 +(1 - f i )·z i
[0080] In the last layer of the bidirectional Quasi-Recurrent convolutional unit, for all sequences split in the spectrum, Quasi-Recurrent pooling is performed in an alternating forward and backward manner to complete the propagation of global information.
[0081] According to one aspect of the present invention, there is provided an electronic device, comprising: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes a Transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method according to any one of the above technical solutions.
[0082] According to one aspect of the present invention, there is provided a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement a Transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method according to any one of the above technical solutions.
[0083] The computer-readable storage medium may include any medium capable of storing or transmitting information. Examples of computer-readable storage media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. Code segments may be downloaded via computer networks such as the Internet, intranets, and the like.
[0084] A Transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method of the present invention is directed to a pair of high-spatial-resolution RGB and low-spatial-resolution hyperspectral images in scene pairing. Based on the encoder of the vision Transformer, a single-pipeline working mechanism is adopted to perform feature extraction and fusion on the RGB and hyperspectral image pairs. In the training stage, a random masking mechanism is used to jointly train the encoder-decoder. In the inference stage, a deep learning image reconstruction module designed based on the Quasi-Recurrent convolutional unit is used to reconstruct the fused features extracted by the Transformer. The attention mechanism in the Transformer enables the network to adaptively learn high-resolution spatial features from visible light modality data, and improve the spatial resolution of hyperspectral data without considering the strict registration of the two modality data, solving the problem that hyperspectral data is limited in application in the remote sensing field due to the lack of spatial information, and improving the inference performance of downstream task algorithms such as classification. The method mainly includes the following parts: (1) Feature extraction and fusion based on the Transformer encoder; (2) Optimization of the encoder by the random masking training mechanism; (3) Image reconstruction of the fused features.
[0085] The present invention reduces the requirement for strong data registration in hyperspectral image super-resolution reconstruction, expands the application scenarios of this technology. Meanwhile, the hyperspectral data with improved spatial resolution is beneficial to optimizing the performance of downstream task algorithms, reducing the design difficulty of the model from the perspective of enhancing data, and enabling the hyperspectral data to be adapted to more remote sensing field tasks. During the training phase, a training method that masks the embedded template block with a certain probability is adopted to enable the model to better learn the features matching cross-modal data. During the inference process, a single-pipeline computing architecture is adopted, and the attention mechanism in the transformer is used to achieve bidirectional circulation and adaptive fusion of feature information in the two modalities. A deep learning image reconstruction module designed based on the Quasi-Recurrent convolutional unit is more suitable for hyperspectral multi-band data and can better utilize the global spectral cross-correlation information.
[0086] In addition, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, the embodiments of the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0087] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0088] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide for implementing the steps of the functions specified in one or more processes and / or boxes Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes.
[0089] It should also be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.
[0090] Finally, it should be noted that the above is the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those skilled in the art of this technology, once the basic creative concept of the present invention is known, without departing from the principle described in the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
Claims
1. A transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method, characterized in that: The following steps are involved: Step S1, obtaining a fused token sequence based on the scene matching RGB image and hyperspectral image; Step S2: Use a random mask mechanism to jointly train the encoder and decoder of the visual transformer to complete feature reconstruction, specifically including: Step S21, in the training stage, the fused block sequence is randomly masked using a preset parameter k as the mask rate, the visible blocks are embedded using linear mapping, and the mask part uses a learnable mask token; Step S22: input the unmasked token sequence into the encoder of the visual transformer for feature extraction; Step S23, inputting the mask part and the enhanced features obtained by the encoder of the visual transformer into the decoder of the visual transformer, obtaining the relative position of the mask according to the position encoding, and reconstructing the features extracted in the step S22 by using the decoder of the visual transformer; Step S3, constructing a loss function based on the similarity between the image reconstructed by the decoder and the true value image, and optimizing the encoder and decoder of the visual transformer; Step S4: super-reconstruct the features extracted by the encoder of the optimized visual transformer using an image reconstruction module, wherein the image reconstruction module is composed of three layers of Quasi-Recurrent convolution units and one layer of bidirectional Quasi-Recurrent convolution units, including: In each Quasi-Recurrent convolution unit, firstly W z and W f The two 3D convolution kernels perform two 3D convolutions and then use different activation functions to obtain two sets of feature maps Z and F, then: Z=tanh(W z *I),F=sigmoid(W f *I) Split the feature map Z and feature map F according to the spectral direction to obtain z i and f i sequence, input it into the Quasi-Recurrent pooling layer, and obtain the fusion feature h i , the fused features are reconstructed by rejoining them according to the spectrum, where the operation in the Quasi-Recurrent pooling layer is expressed as: h i =f i ·h i-1 +(1-f i )·z i In the last layer of bidirectional Quasi-Recurrent convolutional units, Quasi-Recurrent pooling is performed in an alternating forward and backward manner for all sequences split according to the spectrum to complete the propagation of global information.
2. The transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method according to claim 1 is characterized in that: The step S1 specifically includes: Step S11: input the scene-matched RGB image and the hyperspectral image, pre-process the RGB image and the hyperspectral image respectively, and obtain image block sequences corresponding to the two modalities, including an RGB modality image block sequence and a hyperspectral modality image block sequence. The preprocessing process at least includes non-overlapping segmentation and block cutting and adding position coding; Step S12: linearly map and embed the RGB modality image block sequence and the hyperspectral modality image block sequence to obtain token vectors, and then concatenate the token vectors of the two modalities to obtain a fused token sequence.
3. The transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method according to claim 1, characterized in that: The step S3 specifically includes: In the training phase, the reconstruction error is calculated using the L1 function to optimize the decoder and encoder simultaneously, and the loss function is constructed as: L HSI =||F d (F e (unmasked(I fuse ))+masked(I fuse ))-I HSI ||1 L RGB =||F d (F e (unmasked(I fuse ))+masked(I fuse ))-I RGB ||1 Among them, L HSI represents the loss function of the decoder to reconstruct the hyperspectral features generated by the encoder, L RGB represents the loss function of the decoder to reconstruct the visual features generated by the encoder, I fuse represents the concatenated hyperspectral-RGB fusion token sequence, Fe(·) represents the encoder function, and Fd(·) represents the decoder function.
4. The transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method according to claim 2, characterized in that: The step S4 specifically includes: In the inference stage, the masking process in step S21 is not performed, and the fused token sequence is directly input into the encoder of the trained visual transformer for feature extraction; The image reconstruction module is then used to super-reconstruct the features extracted by the encoder.
5. The transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method according to claim 1, characterized in that: In step S22, the encoder of the visual transformer is configured with a total of 4 feature encoding layers, including: In any feature encoding layer, the fused feature vector is first normalized and used as the query, key, and value matrices for multi-head attention calculation. Inside the multi-head attention mechanism, in each head, the fused features are transposed and then matrix multiplied with themselves, and the attention map is obtained through softmax: Among them, β i,j Indicates whether the i-th position should be associated with the j-th position. Each head learns a mapping method h by itself. n (·), then the complete output of the attention mechanism of N heads is:
6. An electronic device, characterized in that: include: One or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory, so that the electronic device performs the transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method as described in any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, implement the transformer-based cross-modal fusion hyperspectral image super-resolution reconstruction method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-modal MR image super-resolution method based on gradient attention enhancement
CN116823613A
Hyperspectral reconstruction method and system based on CASSI optical system
CN118052897A