Hyperspectral and multispectral image fusion method based on expansion convolution and cross attention

Through the method based on expansion convolution and cross attention, the problems of feature extraction and information retention in the fusion of hyperspectral and multispectral images are solved, the quality and accuracy of image fusion are improved, and the full fusion of spatial and spectral information is achieved.

CN120298835AActive Publication Date: 2025-07-11XIAN UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510342544.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-11
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

In the prior art, the spatial resolution of high-spectral images is low and the spectral resolution of multi-spectral images is high. However, the existing fusion methods are difficult to take into account the feature extraction of spatial and spectral information, the calculation efficiency and accuracy are difficult to balance, and the feature fusion is insufficient.

Method used

The hyperspectral and multispectral image fusion method based on expansion convolution and cross attention is adopted. The characteristics are extracted through expansion convolution and the cross attention mechanism are introduced, and the characteristic differences of hyperspectral and multispectral images are processed respectively, attention weights are assigned dynamically for feature fusion, and multi-source information is retained during the reconstruction stage.

Benefits of technology

The quality of the fusion image is improved, the richness and accuracy of feature extraction is improved, the information retention ability is enhanced, and better image content understanding and fusion effect is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298835A_ABST
    Figure CN120298835A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral and multispectral image fusion method based on expansion convolution and cross attention, and the method is specifically implemented according to the following steps: 1, carrying out the preprocessing of original data, simulating the generation of a hyperspectral and multispectral image data set, and dividing the generated hyperspectral and multispectral image data set into a training set and a test set; 2, constructing a hyperspectral and multispectral image fusion network model based on expansion convolution and cross attention; 3, training a hyperspectral and multispectral image fusion network model based on expansion convolution and cross attention; and 4, testing the test set by using the trained hyperspectral and multispectral image fusion network model. The method solves the problems that in the prior art, space and spectral information are difficult to consider in feature extraction, calculation efficiency and accuracy are difficult to balance, and feature fusion is insufficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image fusion in image processing, and specifically relates to a hyperspectral and multispectral image fusion method based on dilated convolution and cross attention. Background Art

[0002] In the field of remote sensing, hyperspectral images (HSIs) and multispectral images (MSIs), with their unique data characteristics, have become key data sources for obtaining information about the Earth's surface. Hyperspectral images refer to image data obtained through hyperspectral imaging technology, which images the same target or scene over multiple consecutive narrow bands. This high-resolution spectral information can be used to identify the components of ground objects, accurately distinguish various ground objects such as vegetation and minerals, and provide key evidence for research in multiple fields such as ecology, geology, and the environment. For example, in geological exploration, potential mineral resources can be detected by analyzing the spectral characteristics of different minerals; in agricultural monitoring, the health status of crops can be accurately judged, and early signs of pests and diseases can be identified. However, the spatial resolution of hyperspectral images is relatively low, and the image details are not rich enough to clearly present the spatial distribution and shape characteristics of ground objects. Multispectral images usually consist of several to dozens of relatively wide bands, focusing on obtaining the reflection or radiation differences of different ground objects in specific bands. Their spatial resolution is relatively high, and they can clearly show the spatial position and shape of ground objects, and are widely used in fields such as urban planning and land use monitoring. For example, in urban planning, different functional areas such as residential areas, commercial areas, and industrial areas can be clearly distinguished through multispectral images. However, due to the limited number of bands, multispectral images are insufficient in precise material identification and are difficult to finely classify complex ground objects.

[0003] In order to fully utilize the advantages of the high spectral resolution of hyperspectral images and the high spatial resolution of multispectral images, the hyperspectral and multispectral image fusion technology has emerged. Early fusion methods were mostly based on simple data-level fusion, such as the weighted average method, which directly weighted and combined the pixel values of the two images. Although the operation was simple, the fusion effect was limited, and it was unable to effectively retain the spectral and spatial information of the images. With the development of technology, fusion methods based on the transform domain gradually emerged, such as wavelet transform. By decomposing the images into different frequency sub-bands, fusing the sub-bands of hyperspectral and multispectral images respectively, and then reconstructing the fused image, the fusion quality was improved to a certain extent. In recent years, with the rapid development of deep learning technology, fusion methods based on neural networks have made remarkable progress. By using the powerful feature extraction and learning ability of deep neural networks, the complementary information of the two images can be more effectively mined to achieve more accurate fusion.

[0004] The main methods for hyperspectral and multispectral image fusion based on deep learning are mainly divided into three categories: based on convolutional neural networks (CNNs), generative adversarial networks (GANs), and autoencoders (AEs). The CNN method uses convolutional layers to automatically extract features, and through multi-layer convolution and pooling operations, it extracts and fuses the features of hyperspectral and multispectral images. Its advantage is that it can automatically learn features and effectively extract details, but it has a large amount of computation and is prone to overfitting. The GAN method consists of a generator for fusing images and a discriminator for judging authenticity. Through adversarial training, the generator is optimized to make the generated images more realistic. However, the training is unstable and prone to mode collapse and artifacts. The AE method maps the image to a low-dimensional feature space through an encoder, and then uses a decoder to reconstruct it into a fused image. It has good feature extraction, compression, and denoising capabilities. However, it may lose details and has poor adaptability to complex scenes. Summary of the Invention

[0005] The purpose of the present invention is to provide a hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention, which solves the problems in the prior art that it is difficult to balance spatial and spectral information in feature extraction, it is difficult to balance computational efficiency and accuracy, and feature fusion is insufficient.

[0006] The technical solution adopted by the present invention is that the hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention is specifically implemented according to the following steps:

[0007] Step 1: Preprocess the original data, simulate and generate a hyperspectral and multispectral image dataset, and then divide the generated hyperspectral and multispectral image dataset into a training set and a test set;

[0008] Step 2: Construct a hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention;

[0009] Step 3: Train the hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention;

[0010] Step 4: Use the trained hyperspectral and multispectral image fusion network model to test the test set.

[0011] The characteristics of the present invention also lie in that

[0012] Step 1 is specifically implemented according to the following steps:

[0013] Step 1.1: Preprocess the dataset. Specifically, perform a 5×5 Gaussian filter on the high-spatial-resolution hyperspectral image (HR-HSI) in the dataset to blur the HR-HSI. Then, perform four downsamplings on the blurred HR-HSI to obtain a low-spatial-resolution hyperspectral image (LR-HSI). Next, extract five bands at equal intervals from the HR-HSI to obtain a high-spatial-resolution multispectral image (HR-MSI).

[0014] Step 1.2: Divide the dataset. Specifically, set the 128×128 region at the center of the low-spatial-resolution hyperspectral image (LR-HSI) and the high-spatial-resolution multispectral image (HR-MSI) generated in Step 1.1 as the test set, and use the remaining part of the image outside this central region as the training area. Before training, perform zero-padding on the test area to ensure that the training can cover the entire region of the image. Finally, randomly crop a region of size 128×128 from the training area as the training set for training.

[0015] The hyperspectral and multispectral image fusion network model in Step 2 consists of three parts: a feature extraction module, a feature fusion module, and a feature reconstruction module.

[0016] In Step 2, in the feature extraction module, it is specifically implemented according to the following steps:

[0017] Make corresponding treatments for the different characteristics of the two types of data, HR-MSI and LR-HSI, which are divided into upper and lower branches. For HR-MSI, in the upper branch, first pass through a downsampling module and a band reshaping module. The band reshaping module is a combination of four 3×3 convolutions and the Relu activation function. By downsampling the HR-MSI to reduce the image size and reshaping the bands to match the LR-HSI, then use a cascaded convolutional layer to extract the spatial information of the multispectral image to obtain the feature HR. Specifically, each layer uses a 3×3 convolution and the Relu activation function. For LR-HSI, in the lower branch, first use an upsampling module to make the LR-HSI match the HR-MSI in terms of spatial scale, and then use a multi-scale dilated convolutional block to extract the spectral information of the hyperspectral image to obtain the feature LR. Specifically, each layer uses a 3×3 dilated convolution and the Relu activation function, and the dilation coefficients of the 3×3 dilated convolutions are 2, 4, and 2 respectively.

[0018] In Step 2, in the feature fusion module, it is specifically implemented according to the following steps:

[0019] The HR-MSI features and LR-HSI features obtained from the feature extraction part are input into the cross-attention fusion module. First, the input HR-MSI features and LR-HSI features respectively pass through a linear layer to generate the query matrix Q, key matrix K, and value matrix V required by the attention mechanism. Then, the correlation between Q and K is established in the feature map, which is realized through the calculation of the feature relationship matrix, as shown below:

[0020]

[0021] d k =(dim / heads)(2)

[0022] Among them, A i respectively represent the feature relationship matrices between the queries and keys of the two inputs. Q1, K1, and V1 are the query matrix, key matrix, and value matrix generated by HR-MSI. Q2, K2, and V2 are the query matrix, keyword matrix, and value matrix generated by LR-HSI. T represents the transpose operation. dim is the input feature dimension, heads is the number of heads of multi-head attention, and softmax represents the normalization operation;

[0023] The calculated feature relationship matrix is used as a weight to perform matrix multiplication with the value matrix of the input of the other branch to obtain the feature soft attention, and then added to the original feature of the input of this branch to perform feature fusion across the spatial and spectral domains, as shown below:

[0024]

[0025] Among them, LR and HR respectively represent the hyperspectral feature and multispectral feature input into the cross-attention fusion module. M and N are the fusion features obtained after cross-attention calculation. Finally, a linear layer Linear is used to reshape M and N, and the reshaped results of M and N are added to obtain the final fusion feature F, which is expressed as follows:

[0026] F = Linear(M)+Linear(N) (4).

[0027] In step 2, in the feature reconstruction module, it is specifically implemented according to the following steps:

[0028] First, the output feature F of the cross-attention fusion module in the feature fusion part is concatenated with the two input features HR and LR of the cross-attention fusion module. Subsequently, it enters the mixed-order convolution feature reconstruction module. The features processed by the band reshaping module are input into the cascaded convolution layer and the multi-scale dilated convolution block in two paths for feature reconstruction, and then combined by addition. Subsequently, the addition result is successively input into the band reshaping module and the upsampling module to perform non-linear transformation on the features. Finally, the features processed above are first concatenated with the original input HR-MSI of the overall fusion network, and then the concatenation result is concatenated with LR-HSI↑↑ obtained by upsampling the LR-MSI twice and undergoes band reshaping to reconstruct the high-spatial-resolution hyperspectral image HR-HSI with high quality in both spatial and spectral dimensions.

[0029] Step 3 is specifically implemented according to the following steps:

[0030] Step 3.1: The root mean square error RMSE is used as the loss function, as follows:

[0031]

[0032] where L fus represents the loss function, is the HR-HSI for reference, and represents the reconstructed HR-HSI. H and W respectively represent the height and width of the image, and C represents the number of spectral channels;

[0033] Step 3.2: Adam is used as the optimizer to adjust the learning rate. The initial learning rate is set to 0.0001. The training set generated in Step 1 is used as the input of the hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention in Step 2 for supervised training. Training stops until the loss value converges or reaches the set maximum number of iterations. Finally, a set of weight parameters corresponding to the image fusion network model based on dilated convolution and cross-attention is obtained.

[0034] Step 4 is specifically implemented according to the following steps:

[0035] The hyperspectral and multispectral images of the test set are input into the image fusion network model based on dilated convolution and cross-attention trained in Step 3, and the trained weight parameters are brought in, and finally the fused image is output.

[0036] The beneficial effects of the present invention are as follows. For the hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention, firstly, ordinary convolution and dilated convolution are respectively used for feature extraction according to the characteristics of different images. Dilated convolution can expand the receptive field without increasing the number of parameters and computational complexity, enabling the network to capture image features at different scales, thereby extracting richer context information, which is helpful for a more comprehensive understanding of the image content and greatly contributes to improving the quality of the fused image. Secondly, in the fusion stage, a cross-attention mechanism is introduced. By calculating the attention weights between different inputs, the model can focus on important feature regions, explore the correlations between hyperspectral and multispectral image features, and enable the model to better fuse the advantageous information of both, improving the accuracy and effectiveness of the fusion result. Finally, in the reconstruction stage, skip connections are adopted to continuously splice the processing process of the reconstructed image features with the previous inputs. This operation can fully retain multi-source information, enhance feature representation, avoid losing key information during the model processing, and make the new feature map contain both basic details and high-level semantics, which helps the model better understand the image content and improve the fusion effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 FIG. is a flowchart for implementing the hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention of the present invention;

[0038] Figure 2 FIG. is a framework diagram of the network model adopted by the hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention of the present invention;

[0039] Figure 3 FIG. is a framework diagram of the multi-scale dilated convolution block in the hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention of the present invention;

[0040] Figure 4 FIG. is a framework diagram of the cross-attention fusion module in the hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention of the present invention;

[0041] Figure 5 FIG. is the mixed-order convolution feature reconstruction module in the hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention of the present invention;

[0042] Figure 6 FIG. is the fusion result of the contrast experiment of the network model in the Urban dataset in the hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0044] In the feature extraction stage of the present invention, according to the different characteristics of multi-spectral images and hyperspectral images, ordinary convolution and dilated convolution are respectively used for processing, realizing the targeted and efficient extraction of the features of the two types of images. In the fusion stage, a cross-attention mechanism is introduced to enable the deep interaction between the spatial features of multi-spectral images and the spectral features of hyperspectral images, dynamically allocate attention weights according to the feature importance, fully explore the internal connection between the features of the two types of images, and achieve feature complementarity. This method provides a better solution for the fusion of hyperspectral and multi-spectral images, and is expected to improve the quality of the fused images and the subsequent application effects.

[0045] The hyperspectral and multi-spectral image fusion method based on dilated convolution and cross-attention of the present invention has a process Figure 1 as shown, and is specifically implemented according to the following steps:

[0046] Step 1: Preprocess the original data, simulate and generate a hyperspectral and multi-spectral image dataset, and then divide the generated hyperspectral and multi-spectral image dataset into a training set and a test set;

[0047] Step 1 is specifically implemented according to the following steps:

[0048] Step 1.1: Preprocess the dataset, that is, perform 5×5 Gaussian filtering on the high-spatial-resolution hyperspectral image HR-HSI in the dataset to blur the HR-HSI, then perform four downsamplings on the blurred high-spatial-resolution hyperspectral image HR-HSI to obtain a low-spatial-resolution hyperspectral image LR-HSI, and then equally spaced extract five bands from the HR-HSI to obtain a high-spatial-resolution multi-spectral image HR-MSI;

[0049] Step 1.2: Divide the dataset, that is, set the 128×128 area in the center of the low-spatial-resolution hyperspectral image LR-HSI and the high-spatial-resolution multi-spectral image HR-MSI generated in Step 1.1 as the test set, and use the remaining part of the image outside the center area as the training area. Before training, perform zero-padding processing on the test area to ensure that the training can cover the entire area of the image. It should be clear that there is no overlapping part between the test set and the training set. Finally, randomly crop a 128×128 area from the training area as the training set for training.

[0050] Step 2: Construct a hyperspectral and multi-spectral image fusion network model based on dilated convolution and cross-attention, and determine the structure of each module;

[0051] The hyperspectral and multi-spectral image fusion network model in Step 2 is as Figure 2As shown in the figure, it consists of three parts: a feature extraction module, a feature fusion module, and a feature reconstruction module. The feature extraction module adopts a dual-branch structure to process HR-MSI and LR-HSI separately. The upper branch for extracting HR-MSI features consists of a downsampling module, a band reshaping module, and a cascaded convolutional layer. The lower branch for extracting LR-MSI features consists of an upsampling module and a multi-scale dilated convolution block. The structure of the multi-scale dilated convolution block is as shown in Figure 3 As shown, it consists of three consecutive dilated convolutions and a ReLu activation function. The dilation coefficients of the three dilated convolutions are 2, 4, and 2 in sequence. The feature fusion module inputs the HR-MSI features and LR-HSI features obtained from the upper and lower branches of the feature extraction part into a cross-attention fusion module to obtain fused features. The structure of the cross-attention fusion module is as shown in Figure 4 As shown, it adopts a dual-branch structure to process the input HR-MSI features and LR-HSI features separately. The upper and lower branch structures are the same. First, a query matrix Q, a key matrix K, and a value matrix V are generated through linear layers respectively. Then, a feature relationship matrix A is calculated by performing a feature relationship matrix calculation on the query matrix Q and the key matrix K. Subsequently, the feature relationship matrix A is multiplied with the value matrix V of the other branch, and then the result of the matrix multiplication is added to the original input feature of this branch. Finally, the calculation results of the two branches are output through a linear layer, and the results are added to obtain fused features. The feature reconstruction part consists of a mixed-order convolution feature reconstruction module, a skip connection, and a band reshaping module. The structure of the mixed-order convolution feature reconstruction module is as shown in Figure 5 As shown, taking a feature with a size of (H / 2, W / 2, 3C) as the input, first, the band dimension is adjusted through a band reshaping module. Then, the feature is divided into two paths. One path enters a cascaded convolutional layer, and the other path enters a multi-scale dilated convolution block. After different feature reconstructions, the results of the two paths are added and fused. The fused result is then adjusted in bands through a band reshaping module, and finally, the spatial resolution is enhanced through an upsampling module to output a feature with a size of (H, W, C), realizing feature reconstruction.

[0052] Combined with Figure 2 and Figure 3 , in step 2, in the feature extraction module, it is specifically implemented according to the following steps:

[0053] Corresponding processing is carried out according to the different characteristics of HR-MSI and LR-HSI data, which is divided into upper and lower branches. For HR-MSI, in the upper branch, it first passes through a downsampling module and a band reshaping module. The band reshaping module is a combination of four 3×3 convolutions and the Relu activation function. By downsampling HR-MSI, the image size is reduced, and band reshaping is performed to match LR-HSI, reduce the computational load, and expand the receptive field. Then, a cascaded convolutional layer is used to extract the spatial information of the hyperspectral image to obtain the feature HR. Specifically, each layer uses a 3×3 convolution and the Relu activation function. Such a design can obtain a stronger ability to capture spatial details and has certain advantages in dealing with spatial correlations. For LR-HSI, in the lower branch, it first passes through an upsampling module to make LR-HSI match HR-MSI in terms of spatial scale. Then, a multi-scale dilated convolutional block is used to extract the spectral information of the hyperspectral image to obtain the feature LR, as Figure 3 shown. Specifically, each layer uses a 3×3 dilated convolution and the Relu activation function. The dilation coefficients of the 3×3 dilated convolution are 2, 4, and 2 respectively. Such a design can significantly expand the receptive field without increasing the number of parameters and better capture the global spectral features.

[0054] Combined with Figure 2 and Figure 4 , in step 2, in the feature fusion module, it is specifically implemented according to the following steps:

[0055] The HR-MSI feature and LR-HSI feature obtained from the feature extraction part are input into the cross-attention fusion module, and the cross-attention fusion module is as Figure 4 shown. First, the input HR-MSI feature and LR-HSI feature respectively pass through a linear layer to generate the query matrix Q, key matrix K, and value matrix V required for the attention mechanism. Then, the correlation between Q and K is established in the feature map, which is realized through the feature relationship matrix calculation, as follows:

[0056]

[0057] d k =(dim / heads)(2)

[0058] where A i respectively represent the feature relationship matrices between the queries and keys of the two inputs. Q1, K1, and V1 are the query matrix, key matrix, and value matrix generated by HR-MSI, and Q2, K2, and V2 are the query matrix, keyword matrix, and value matrix generated by LR-HSI. T represents the transpose operation, dim is the input feature dimension, heads is the number of heads of multi-head attention, and softmax represents the normalization operation;

[0059] The calculated characteristic relationship matrix is used as weights to perform matrix multiplication with the value matrix input by another branch to obtain feature soft attention, which is then added to the original features input by this branch to perform feature fusion across the spatial and spectral domains, as follows:

[0060]

[0061] Where LR and HR represent the hyperspectral features and multispectral features input to the cross-attention fusion module respectively, M and N are the fusion features obtained after cross-attention calculation. Finally, a linear layer Linear is used to reshape M and N, and the reshaped results of M and N are added together to obtain the final fusion feature F, which is expressed as follows:

[0062] F = Linear(M) + Linear(N) (4).

[0063] This method can enable more sufficient information interaction and fusion of the features of hyperspectral images and multispectral images in the spectral and spatial dimensions, reduce the amount of information loss, and improve the fusion effect.

[0064] Combined with Figure 2 and Figure 5 , in step 2, in the feature reconstruction module, it is specifically implemented according to the following steps:

[0065] First, the output feature F of the cross-attention fusion module in the feature fusion part is concatenated with the two input features HR and LR of the cross-attention fusion module. This operation retains the basic information and details in the original input features and prevents information loss during fusion. Subsequently, it enters the mixed-order convolutional feature reconstruction module, as Figure 5 shown. The features processed by the band reshaping module will be input into the cascaded convolutional layer and the multi-scale dilated convolutional block in two paths for feature reconstruction, and then combined by addition. During this process, information of different scales and characteristics complements each other, refining the feature representation. Subsequently, the addition result is sequentially input into the band reshaping module and the upsampling module to further perform non-linear transformation on the features and enhance the expression ability of the features. Finally, the features processed above are first concatenated with the HR-MSI of the original input of the overall fusion network, and then the concatenation result is concatenated with the LR-HSI↑↑ obtained by upsampling the LR-MSI twice and undergoes band reshaping, making full use of the feature advantages at different stages, enriching the diversity and integrity of information, and reconstructing the high-spatial-resolution hyperspectral image HR-HSI with high quality in both spatial and spectral dimensions.

[0066] Step 3, train the hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention;

[0067] Step 3 is specifically implemented according to the following steps:

[0068] Step 3.1: Use the Root Mean Square Error (RMSE) as the loss function, as shown below:

[0069]

[0070] where L fus represents the loss function, is the reference HR-HSI, and represents the reconstructed HR-HSI. H and W represent the height and width of the image respectively, and C represents the number of spectral channels;

[0071] Step 3.2: Use Adam as the optimizer to adjust the learning rate. Set the initial learning rate to 0.0001. Use the training set generated in Step 1 as the input for the hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention in Step 2 for supervised training until the loss value converges or reaches the set maximum number of iterations, and then stop training. Finally, obtain a set of weight parameters corresponding to the image fusion network model based on dilated convolution and cross-attention.

[0072] Step 4: Use the trained hyperspectral and multispectral image fusion network model to test the test set.

[0073] Step 4 is specifically implemented according to the following steps:

[0074] Input the hyperspectral and multispectral images of the test set into the image fusion network model based on dilated convolution and cross-attention trained in Step 3, and bring in the trained weight parameters to finally output the fused image.

[0075] For the fused image, compare the obtained fused image with the original HR-HSI image, conduct a quantitative evaluation of the fused image, and select 4 representative evaluation metrics, which are: Root Mean Square Error (RMSE), Peak Signal to Noise Ratio (PSNR), Relative Average Spectral Error (RASE), and Spectral Angle Mapper (SAM). At the same time, conduct a qualitative analysis based on the fusion result map.

[0076] Example 1

[0077] The hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention of the present invention has the following process Figure 1 as shown, and is specifically implemented according to the following steps:

[0078] Step 1: Preprocess the original data, simulate and generate hyperspectral and multispectral image datasets, and then divide the generated hyperspectral and multispectral image datasets into a training set and a test set;

[0079] Step 2: Construct a hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention, and determine the structure of each module;

[0080] Step 3: Train the hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention;

[0081] Step 4: Use the trained hyperspectral and multispectral image fusion network model to test the test set.

[0082] Example 2

[0083] The hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention of the present invention has a process Figure 1 as shown, and is specifically implemented according to the following steps:

[0084] Step 1: Preprocess the original data, simulate and generate hyperspectral and multispectral image datasets, and then divide the generated hyperspectral and multispectral image datasets into a training set and a test set;

[0085] Step 1 is specifically implemented according to the following steps:

[0086] Step 1.1: Preprocess the dataset, that is, perform 5×5 Gaussian filtering on the high-spatial-resolution hyperspectral image HR-HSI in the dataset to blur the HR-HSI, then perform four downsamplings on the blurred high-spatial-resolution hyperspectral image HR-HSI to obtain a low-spatial-resolution hyperspectral image LR-HSI, and then equally spaced extract five bands from the HR-HSI to obtain a high-spatial-resolution multispectral image HR-MSI;

[0087] Step 1.2: Divide the dataset, that is, set the 128×128 area in the center of the low-spatial-resolution hyperspectral image LR-HSI and the high-spatial-resolution multispectral image HR-MSI generated in Step 1.1 as the test set, and use the remaining part of the image outside this central area as the training area. Before training, perform zero-padding processing on the test area to ensure that the training can cover the entire area of the image. It should be clear that there is no overlap between the test set and the training set. Finally, randomly crop a 128×128 area from the training area as the training set for training.

[0088] Step 2: Construct a hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention, and determine the structure of each module;

[0089] Step 3: Train a hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention;

[0090] Step 4: Use the trained hyperspectral and multispectral image fusion network model to test the test set.

[0091] Example 3

[0092] The hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention of the present invention has a process Figure 1 as shown, and is specifically implemented according to the following steps:

[0093] Step 1: Preprocess the original data, simulate and generate a hyperspectral and multispectral image data set, and then divide the generated hyperspectral and multispectral image data set into a training set and a test set;

[0094] Step 1 is specifically implemented according to the following steps:

[0095] Step 1.1: Preprocess the data set, that is, perform 5×5 Gaussian filtering on the high-spatial-resolution hyperspectral image HR-HSI in the data set to blur the HR-HSI, then perform four downsamplings on the blurred high-spatial-resolution hyperspectral image HR-HSI to obtain a low-spatial-resolution hyperspectral image LR-HSI, and then equally-spaced extract five bands from the HR-HSI to obtain a high-spatial-resolution multispectral image HR-MSI;

[0096] Step 1.2: Divide the data set, that is, set the 128×128 area in the center of the low-spatial-resolution hyperspectral image LR-HSI and the high-spatial-resolution multispectral image HR-MSI generated in Step 1.1 as the test set, and use the remaining part of the image outside the center area as the training area. Before starting the training, perform zero-padding processing on the test area to ensure that the training can cover the entire area of the image. It should be clear that there is no overlap between the test set and the training set. Finally, randomly crop a 128×128 area from the training area as the training set for training.

[0097] Step 2: Construct a hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention, and determine the structure of each module;

[0098] In Step 2, the hyperspectral and multispectral image fusion network model is as Figure 2 shown, and includes three parts: a feature extraction module, a feature fusion module, and a feature reconstruction module.

[0099] Combined with Figure 2 and Figure 3 , in Step 2, in the feature extraction module, it is specifically implemented according to the following steps:

[0100] Corresponding processing is performed according to the different characteristics of HR-MSI and LR-HSI data, which is divided into upper and lower branches. For HR-MSI, in the upper branch, it first passes through a downsampling module and a band reshaping module. The band reshaping module is a combination of four 3×3 convolutions and the Relu activation function. By downsampling HR-MSI, the image size is reduced, and band reshaping is performed to match LR-HSI, reduce the computational load, and expand the receptive field. Then, a cascaded convolutional layer is used to extract the spatial information of the hyperspectral image to obtain the feature HR. Specifically, each layer uses a 3×3 convolution and the Relu activation function. Such a design can obtain a stronger ability to capture spatial details and has certain advantages in dealing with spatial correlation. For LR-HSI, in the lower branch, it first passes through an upsampling module to make LR-HSI match HR-MSI in terms of spatial scale, and then a multi-scale dilated convolution block is used to extract the spectral information of the hyperspectral image to obtain the feature LR, as Figure 3 shown. Specifically, each layer uses a 3×3 dilated convolution and the Relu activation function, and the dilation coefficients of the 3×3 dilated convolution are 2, 4, and 2 respectively. Such a design can significantly expand the receptive field without increasing the number of parameters and better capture the global spectral features.

[0101] Step 3: Train a hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention;

[0102] Step 4: Use the trained hyperspectral and multispectral image fusion network model to test the test set.

[0103] Example 4

[0104] The hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention of the present invention has a process Figure 1 as shown, and is specifically implemented according to the following steps:

[0105] Step 1: Preprocess the original data, simulate and generate a hyperspectral and multispectral image dataset, and then divide the generated hyperspectral and multispectral image dataset into a training set and a test set;

[0106] Step 1 is specifically implemented according to the following steps:

[0107] Step 1.1: Preprocess the dataset, that is, perform 5×5 Gaussian filtering on the high-spatial-resolution hyperspectral image HR-HSI in the dataset to blur HR-HSI, then perform four downsamplings on the blurred high-spatial-resolution hyperspectral image HR-HSI to obtain a low-spatial-resolution hyperspectral image LR-HSI, and then equally-spaced extract five bands from HR-HSI to obtain a high-spatial-resolution multispectral image HR-MSI;

[0108] Step 1.2: Divide the dataset. That is, set the 128×128 area at the center of the low-spatial-resolution hyperspectral image LR-HSI and the high-spatial-resolution multispectral image HR-MSI generated in Step 1.1 as the test set, and use the remaining part of the image outside this central area as the training area. Before training, perform zero-padding on the test area to ensure that the training can cover the entire area of the image. It should be clear that there is no overlapping part between the test set and the training set. Finally, randomly crop a 128×128 area from the training area as the training set for training.

[0109] Step 2: Construct a hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention, and determine the structure of each module;

[0110] In Step 2, the hyperspectral and multispectral image fusion network model is as Figure 2 shown, which consists of three parts: a feature extraction module, a feature fusion module, and a feature reconstruction module.

[0111] Combined with Figure 2 and Figure 3 , in Step 2, in the feature extraction module, it is specifically implemented according to the following steps:

[0112] Make corresponding treatments for the different characteristics of the two types of data, HR-MSI and LR-HSI, which are divided into upper and lower branches. For HR-MSI, in the upper branch, it first passes through a downsampling module and a band reshaping module. The band reshaping module is a combination of four 3×3 convolutions and the Relu activation function. By downsampling HR-MSI to reduce the image size and reshaping the bands, the purpose of matching LR-HSI, reducing the computational load, and expanding the receptive field is achieved. Then, use a cascaded convolutional layer to extract the spatial information of the multispectral image to obtain the feature HR. Specifically, each layer uses a 3×3 convolution and the Relu activation function. Such a design can obtain stronger spatial detail capture ability and has certain advantages in dealing with spatial correlation. For LR-HSI, in the lower branch, first use an upsampling module to make LR-HSI match HR-MSI in the spatial scale, and then use a multi-scale dilated convolution block to extract the spectral information of the hyperspectral image to obtain the feature LR, as Figure 3 shown. Specifically, each layer uses a 3×3 dilated convolution and the Relu activation function, and the dilation coefficients of the 3×3 dilated convolution are 2, 4, and 2 respectively. Such a design can significantly expand the receptive field without increasing the number of parameters and better capture the global spectral features.

[0113] Combined with Figure 2 and Figure 4 , in Step 2, in the feature fusion module, it is specifically implemented according to the following steps:

[0114] The HR-MSI features and LR-HSI features obtained from the feature extraction part are input into the cross-attention fusion module, and the cross-attention fusion module is as shown in Figure 4 . First, the input HR-MSI features and LR-HSI features respectively pass through a linear layer to generate the query matrix Q, key matrix K, and value matrix V required by the attention mechanism. Then, the correlation between Q and K is established in the feature map, which is realized by calculating the feature relationship matrix, as follows:

[0115]

[0116] d k =(dim / heads)(2)

[0117] where, A i respectively represent the feature relationship matrices between the queries and keys of the two inputs. Q1, K1, and V1 are the query matrix, key matrix, and value matrix generated by HR-MSI, and Q2, K2, and V2 are the query matrix, keyword matrix, and value matrix generated by LR-HSI. T represents the transpose operation, dim is the input feature dimension, heads is the number of heads in multi-head attention, and softmax represents the normalization operation;

[0118] The calculated feature relationship matrix is used as a weight to perform matrix multiplication with the value matrix of the input of the other branch to obtain the feature soft attention, and then added to the original feature of the input of this branch to perform feature fusion across the spatial and spectral domains, as follows:

[0119]

[0120] where, LR and HR respectively represent the hyperspectral feature and multispectral feature input into the cross-attention fusion module. M and N are the fusion features obtained after cross-attention calculation. Finally, a linear layer Linear is used to reshape M and N, and the reshaped results of M and N are added to obtain the final fusion feature F, which is expressed as follows:

[0121] F = Linear(M)+Linear(N)(4).

[0122] This method can enable the features of hyperspectral images and multispectral images to achieve more sufficient information interaction and fusion in the spectral and spatial dimensions, reduce the amount of information loss, and improve the fusion effect.

[0123] Step 3: Train the hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention;

[0124] Step 4: Use the trained hyperspectral and multispectral image fusion network model to test the test set.

[0125] Example 5

[0126] The hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention of the present invention has a process Figure 1 as shown, and is specifically implemented according to the following steps:

[0127] Step 1. Preprocess the original data, simulate and generate a hyperspectral and multispectral image dataset, and then divide the generated hyperspectral and multispectral image dataset into a training set and a test set;

[0128] Step 1 is specifically implemented according to the following steps:

[0129] Step 1.1. Preprocess the dataset, that is, perform a 5×5 Gaussian filter on the high-spatial-resolution hyperspectral image HR-HSI in the dataset to blur the HR-HSI, then perform four downsamplings on the blurred high-spatial-resolution hyperspectral image HR-HSI to obtain a low-spatial-resolution hyperspectral image LR-HSI, and then equally-spaced extract five bands from the HR-HSI to obtain a high-spatial-resolution multispectral image HR-MSI;

[0130] Step 1.2. Divide the dataset, that is, set the 128×128 area in the center of the low-spatial-resolution hyperspectral image LR-HSI and the high-spatial-resolution multispectral image HR-MSI generated in Step 1.1 as the test set, and use the remaining part of the image outside this central area as the training area. Before starting the training, perform zero-padding processing on the test area to ensure that the training can cover the entire area of the image. It should be clear that there is no overlapping part between the test set and the training set. Finally, randomly crop a 128×128 area from the training area as the training set for training.

[0131] Step 2. Build a hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention, and determine the structure of each module;

[0132] In Step 2, the hyperspectral and multispectral image fusion network model is as Figure 2 shown, and includes three parts: a feature extraction module, a feature fusion module, and a feature reconstruction module.

[0133] Combined with Figure 2 and Figure 5 , in Step 2, in the feature reconstruction module, it is specifically implemented according to the following steps:

[0134] First, splice the output feature F of the cross-attention fusion module in the feature fusion part with the two input features HR and LR of the cross-attention fusion module. This operation retains the basic information and details in the original input features and prevents information loss during the fusion. Subsequently, enter the mixed-order convolution feature reconstruction module, as Figure 5As shown in the figure, the features processed by the band reshaping module will be input into the cascaded convolutional layer and the multi-scale dilated convolutional block in two paths for feature reconstruction, and then combined by addition. In this process, information with different scales and characteristics complements each other, refining the feature representation. Subsequently, the addition result is input into the band reshaping module and the upsampling module in sequence to further perform non-linear transformation on the features and enhance the expression ability of the features. Finally, the features processed above are first spliced with the original input HR-MSI of the overall fusion network, and then the splicing result is spliced with the LR-HSI↑↑ obtained by upsampling the LR-MSI twice and band reshaping is performed, making full use of the feature advantages at different stages, enriching the diversity and integrity of information, and reconstructing the high spatial resolution hyperspectral image HR-HSI with high quality in both spatial and spectral dimensions.

[0135] Step 3: Train the hyperspectral and multispectral image fusion network model based on dilated convolution and cross attention;

[0136] Step 3 is specifically implemented according to the following steps:

[0137] Step 3.1: Use the root mean square error RMSE as the loss function, as follows:

[0138]

[0139] where L fus represents the loss function, is the HR-HSI for reference, and represents the reconstructed HR-HSI, H and W respectively represent the height and width of the image, and C represents the number of spectral channels;

[0140] Step 3.2: Use Adam as the optimizer to adjust the learning rate, set the initial learning rate to 0.0001, and use the training set generated in Step 1 as the input of the hyperspectral and multispectral image fusion network model based on dilated convolution and cross attention in Step 2 for supervised training, and stop training until the loss value converges or reaches the set maximum number of iterations. Finally, a set of weight parameters corresponding to the image fusion network model based on dilated convolution and cross attention is obtained.

[0141] Step 4: Use the trained hyperspectral and multispectral image fusion network model to test the test set.

[0142] Example 6

[0143] The following further illustrates the effect of the present invention in combination with experiments:

[0144] 1. Experimental conditions:

[0145] The experiments of the present invention used NVIDIA GeForce RTX 4090D (24G), the deep learning framework used was PyTorch 2.1.0, and the programming language was Python 3.10.

[0146] 2. Dataset Introduction

[0147] The Urban dataset used in the experiments of the present invention was obtained by the Hydice sensor when photographing Copper Tree Bay in Texas, USA in 1995. The image size of this dataset was 307×307, and the original data covered 210 bands. After completing preprocessing operations such as noise elimination and water absorption band removal, usually 162 bands were retained for subsequent processing and analysis work.

[0148] 3. Parameter Settings:

[0149] The present invention selected the Adam algorithm as the optimizer, and the initial learning rate was set to 0.0001.

[0150] 4. Experimental Comparison

[0151] To verify the effectiveness of the method proposed in the present invention, the present invention was compared with 7 existing hyperspectral and multispectral image fusion methods, and the fusion results of the proposed method were quantitatively evaluated and qualitatively analyzed. The evaluation index results of different fusion methods on the Urban dataset are shown in Table 1, and the fusion results of different fusion methods on the Urban dataset are as Figure 6 shown.

[0152] Table 1 Fusion Performance of Different Methods on the Urban Dataset

[0153] Method RMSE↓ PSNR↑ ERGAS↓ SAM↓ MSDCNN 3.1317 35.9676 1.7652 3.1199 TFNet 3.0405 36.2243 1.7146 2.8524 ResTFNet 2.8916 36.6604 1.598 2.7355 SSFCNN 8.5801 27.2133 4.2196 8.7402 ConSSFCNN 4.9788 34.1876 1.8615 3.1634 MCT 2.5517 37.7465 1.34 2.3299 Proposed 2.4556 38.0799 1.2776 2.2704

[0154] Table 1 shows that the method proposed in the present invention has achieved strong competitive results in all four indicators. Even though the number of bands in the Urban dataset is large, the method still achieved relatively ideal fusion results, fully demonstrating the effectiveness of the method in maintaining spectral information and spatial information. Figure 6 The fusion result diagram more intuitively shows the fusion effect. Through comparison, it is found that the method proposed in the present invention has the best effect in retaining the image spatial structure and detail features, that is, the features such as the contours and textures of ground objects are the clearest and are basically the same as the original image. Therefore, the method proposed in the present invention has good performance in hyperspectral and multispectral image fusion.

Claims

1. A hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention, characterized in that The implementation is specifically carried out according to the following steps: Step 1: Preprocess the original data, simulate and generate hyperspectral and multispectral image datasets, and then divide the generated hyperspectral and multispectral image datasets into training sets and test sets; Step 2: Construct a hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention; Step 3: Train the hyperspectral and multispectral image fusion network model based on dilated convolution and cross-attention; Step 4: Use the trained hyperspectral and multispectral image fusion network model to test the test set.

2. The hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention according to claim 1, wherein The specific implementation of Step 1 is carried out according to the following steps: Step 1.1: Preprocess the dataset, that is, perform 5×5 Gaussian filtering on the high-spatial-resolution hyperspectral image HR-HSI in the dataset to blur the HR-HSI, then perform four downsamplings on the blurred high-spatial-resolution hyperspectral image HR-HSI to obtain a low-spatial-resolution hyperspectral image LR-HSI, and then extract five bands at equal intervals from the HR-HSI to obtain a high-spatial-resolution multispectral image HR-MSI; Step 1.2: Divide the dataset, that is, set the 128×128 area in the center of the low-spatial-resolution hyperspectral image LR-HSI and the high-spatial-resolution multispectral image HR-MSI generated in Step 1.1 as the test set, and use the remaining part of the image outside this central area as the training area. Before starting training, perform zero-padding processing on the test area to ensure that the training can cover the entire area of the image. Finally, randomly crop a 128×128 area from the training area as the training set for training.

3. The hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention according to claim 2, wherein The hyperspectral and multispectral image fusion network model in Step 2 includes three parts: a feature extraction module, a feature fusion module, and a feature reconstruction module.

4. The hyperspectral and multispectral image fusion method based on dilated convolution and cross attention according to claim 3, characterized in that In Step 2, in the feature extraction module, the specific implementation is carried out according to the following steps: Corresponding processing is performed according to the different characteristics of the two types of data, HR-MSI and LR-HSI, which are divided into upper and lower branches. For HR-MSI, in the upper branch, it first passes through a downsampling module and a band reshaping module. The band reshaping module is a combination of four 3×3 convolutions and the Relu activation function. By downsampling the HR-MSI to reduce the image size and performing band reshaping to match the LR-HSI, and then using a cascaded convolutional layer to extract the spatial information of the multispectral image to obtain the feature HR. Specifically, each layer uses a 3×3 convolution and the Relu activation function. For LR-HSI, in the lower branch, it first passes through an upsampling module to make the LR-HSI match the HR-MSI in the spatial scale, and then uses a multi-scale dilated convolution block to extract the spectral information of the hyperspectral image to obtain the feature LR. Specifically, each layer uses a 3×3 dilated convolution and the Relu activation function, and the dilation coefficients of the 3×3 dilated convolution are 2, 4, and 2 respectively.

5. The hyperspectral and multispectral image fusion method based on dilated convolution and cross attention according to claim 4, wherein In Step 2, in the feature fusion module, the specific implementation is carried out according to the following steps: The HR-MSI features and LR-HSI features obtained from the feature extraction part are input into the cross-attention fusion module. First, the input HR-MSI features and LR-HSI features respectively pass through a linear layer to generate the query matrix Q, key matrix K, and value matrix V required by the attention mechanism. Then, the correlation between Q and K is established in the feature map, which is realized through the calculation of the feature relationship matrix, as shown below: d k = (dim / heads) (2) Among them, A i respectively represent the feature relationship matrices between two input queries and keys. Q1, K1, and V1 are the query matrix, key matrix, and value matrix generated by HR-MSI. Q2, K2, and V2 are the query matrix, keyword matrix, and value matrix generated by LR-HSI. T represents the transpose operation. dim is the input feature dimension, heads is the number of heads for multi-head attention, and softmax represents the normalization operation; The calculated feature relationship matrix is used as a weight to perform matrix multiplication with the value matrix input from the other branch to obtain feature soft attention, and then added to the original features input from this branch to perform feature fusion across the spatial and spectral domains, as shown below: Where LR and HR respectively represent the hyperspectral features and multispectral features input into the cross-attention fusion module, M and N are the fused features obtained after cross-attention calculation. Finally, a linear layer Linear is used to reshape M and N, and the reshaped results of M and N are added to obtain the final fused feature F, which is expressed as follows: F = Linear(M) + Linear(N) (4).

6. The hyperspectral and multispectral image fusion method based on dilated convolution and cross attention according to claim 5, wherein In step 2, in the feature reconstruction module, it is specifically implemented according to the following steps: First, the output feature F of the cross-attention fusion module in the feature fusion part is concatenated with the two input features HR and LR of the cross-attention fusion module. Subsequently, it enters the mixed-order convolution feature reconstruction module. The features processed by the band reshaping module are input into the cascaded convolution layer and the multi-scale dilation convolution block in two paths for feature reconstruction, and then combined by addition. Subsequently, the addition result is successively input into the band reshaping module and the upsampling module to perform non-linear transformation on the features. Finally, the features processed above are first concatenated with the HR-MSI originally input by the overall fusion network, and then the concatenated result is concatenated with the LR-HSI↑↑ obtained by upsampling the LR-MSI twice and the band reshaping is performed to reconstruct the high-spatial-resolution hyperspectral image HR-HSI with high quality in both the spatial and spectral dimensions.

7. The hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention according to claim 6, characterized in that Step 3 is specifically implemented according to the following steps: Step 3.1: Use the root mean square error RMSE as the loss function, as shown below: Among them, L fus represents the loss function, is the HR-HSI for reference, while represents the reconstructed HR-HSI. H and W respectively represent the height and width of the image, and C represents the number of spectral channels; Step 3.2: Use Adam as the optimizer to adjust the learning rate. Set the initial learning rate to 0.0001. Use the training set generated in step 1 as the input of the hyperspectral and multispectral image fusion network model based on dilation convolution and cross-attention in step 2 for supervised training until the loss value converges or reaches the set maximum number of iterations, and finally obtain a set of weight parameters corresponding to the image fusion network model based on dilation convolution and cross-attention.

8. The hyperspectral and multispectral image fusion method based on dilated convolution and cross-attention according to claim 7, characterized in that Step 4 is specifically implemented according to the following steps: Input the hyperspectral and multispectral images of the test set into the image fusion network model based on dilation convolution and cross-attention trained in step 3, and bring in the trained weight parameters, and finally output the fused image.

Citation Information

Patent Citations

  • Hyperspectral and multispectral image fusion method based on attention mechanism

    CN117474781A

  • Multi-source remote sensing image fusion method and system

    CN118968245A

  • Online quality monitoring method and system based on optical multispectral fusion

    CN119198566A