Hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement

By constructing a hyperspectral and multispectral image fusion network based on classification and multi-level residual enhancement, the problem of low image fusion quality in the prior art is solved, and efficient image fusion and fidelity improvement of spectral features is achieved.

CN119399043BActive Publication Date: 2025-05-09QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510005236.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-09
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

Existing hyperspectral and multispectral image fusion methods based on deep learning are difficult to organically combine classification tasks with image fusion tasks, resulting in low image fusion quality.

Method used

A hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement is proposed. By constructing an image fusion network based on classification guidance and multi-level residual enhancement, combining band embedding modules, spatial attention modules, multi-level residual enhancement modules and classification networks, efficient fusion of images is achieved.

Benefits of technology

This method can significantly improve the resolution of hyperspectral images and the fidelity of spectral features, enhance the reconstruction ability of image details, and have strong robustness to background noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399043B_ABST
    Figure CN119399043B_ABST
Patent Text Reader

Abstract

The present application discloses a hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement, and relates to the field of new generation information technology. The reconstructed high-resolution hyperspectral image obtained by the hyperspectral and multispectral image fusion method described in the present application is more accurate, and the deviation between the reconstructed high-resolution hyperspectral image and the label corresponding to the source image is less, which is more prominent in the reconstruction of image details and can have a higher restoration accuracy; moreover, the hyperspectral and multispectral image fusion method described in the present application has strong robustness and adaptability when fusing source images with more background noise, and can effectively suppress the interference of background noise on the image fusion quality; in addition, the hyperspectral and multispectral image fusion method described in the present application also has strong distinguishing ability and reconstruction accuracy for different categories of ground objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of new generation information technology and relates to a hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement. Background Art

[0002] Hyperspectral images have hundreds of continuous spectral bands and can capture rich spectral features. However, due to the limitations of physical equipment, existing hyperspectral sensors face a trade-off between spatial and spectral information when capturing images. Therefore, improving the resolution of hyperspectral images has attracted much attention from researchers in recent years. Multispectral images have higher spatial resolution and fewer spectral bands. Fusion of low-resolution hyperspectral images with high-resolution multispectral images has become an important way to obtain high-resolution hyperspectral images. These fused high-resolution hyperspectral images have a wide range of applications in remote sensing data analysis, environmental change detection, image classification, target recognition and other fields.

[0003] Existing hyperspectral and multispectral image fusion methods based on deep learning perform excellent fusion effects by automatically extracting features from hyperspectral images and multispectral images. However, most existing hyperspectral and multispectral image fusion methods based on deep learning cannot organically combine the classification task with the image fusion task, and use the information in the classification task to guide the image fusion task to improve the image fusion quality. Therefore, this application proposes a hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement. Summary of the invention

[0004] In view of the defects and shortcomings in the above-mentioned prior art, the present application proposes a hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement.

[0005] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme:

[0006] A hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement includes the following steps:

[0007] S1: Preprocess the images in the Pavia University dataset and the Pavia Center dataset to obtain the training set A and the test set;

[0008] S2, construct an image fusion network based on classification guidance and multi-level residual enhancement;

[0009] The image fusion network based on classification guidance and multi-level residual enhancement includes a parallel band embedding module and a spatial attention module, the band embedding module is connected to the first 3×3 convolution layer, the first 3×3 convolution layer is respectively connected to the multi-level residual enhancement module and the classification network, and the image output by the classification network is the output result of the classification task; the classification network and the spatial attention module are both connected to the information interaction module, the multi-level residual enhancement module and the information interaction module are both connected to the first element-by-element multiplication module, and the first element-by-element multiplication module is connected to the second 3×3 convolution layer; in the present application, an upsampling module and a channel attention module are also included, the upsampling module, the channel attention module and the band embedding module are connected in parallel, the channel attention module and the second 3×3 convolution layer are both connected to the second element-by-element multiplication unit, the upsampling module and the second element-by-element multiplication module are both connected to the attention-based multi-scale fusion module, and the output image of the attention-based multi-scale fusion module is the reconstructed high-resolution hyperspectral image;

[0010] S3. Constructing the total generation loss of the image fusion network based on classification guidance and multi-level residual enhancement ;

[0011] S4, obtaining a training set B based on the training set A, and using the training set B to train the image fusion network based on classification guidance and multi-level residual enhancement to obtain an image fusion network model based on classification guidance and multi-level residual enhancement;

[0012] S5. Input the low-resolution hyperspectral image to be processed and the high-resolution multispectral image to be processed into the trained image fusion network model based on classification guidance and multi-level residual enhancement. After forward propagation once, the predicted reconstructed high-resolution hyperspectral image and classification result can be obtained.

[0013] Preferably, step S1 specifically includes the following steps:

[0014] S1-1, based on the images in the Pavia University dataset, obtain the original low-resolution hyperspectral images for training and the original low-resolution hyperspectral images for testing;

[0015] S1-2, based on the images in the Pavia Center dataset, obtain the original low-resolution hyperspectral images for training and the original low-resolution hyperspectral images for testing;

[0016] S1-3, original high-resolution multispectral images for training based on images from the Pavia University dataset and original high-resolution multispectral images for testing;

[0017] S1-4, original high-resolution multispectral images for training based on images in the Pavia Center dataset and original high-resolution multispectral images for testing;

[0018] The original low-resolution hyperspectral images used for training and the high-resolution multispectral images used for training in this application together constitute a training set A; the original low-resolution hyperspectral images used for testing in this application and the high-resolution multispectral images used for testing together constitute a test set.

[0019] Preferably, in step S2, the band embedding module includes a downsampling module, the input end of the downsampling module is connected to the input end of the spatial attention module, the output end of the downsampling module is connected to the first Concat layer, the input end of the first Concat layer is connected to the input end of the channel attention module, the output end of the first Concat layer is connected to the pixel shuffling module, the output end of the pixel shuffling module and the input end of the downsampling module are both connected to the second Concat layer, and the output end of the second Concat layer serves as the output end of the band embedding module.

[0020] Preferably, in step S2, the spatial attention module includes a maximum pooling layer and an average pooling layer, the input end of the maximum pooling layer is connected to the input end of the average pooling layer, the output end of the maximum pooling layer and the output end of the average pooling layer are both connected to the Concat layer, the Concat layer is also connected to the first convolution layer (the convolution kernel size is 3×3), the first Sigmoid layer, the second convolution layer (the convolution kernel size is 3×3) and the second Sigmoid layer in sequence, and the output end of the second Sigmoid layer and the input end of the maximum pooling layer are both connected to the input end of the element-by-element multiplication unit.

[0021] Preferably, in step S2, the channel attention module includes a maximum pooling layer and an average pooling layer, the input end of the maximum pooling layer is connected to the input end of the average pooling layer, the maximum pooling layer is also connected to the first fully connected layer, the first ReLU activation layer, and the second fully connected layer in sequence, the average pooling layer is connected to the third fully connected layer, the second ReLU activation layer, and the fourth fully connected layer in sequence, the output ends of the second fully connected layer and the fourth fully connected layer are both connected to the Add layer, the Add layer is connected to the convolution layer (the convolution kernel size is 3×3) and the Sigmoid layer in sequence, and the output end of the Sigmoid layer and the input end of the maximum pooling layer are both connected to the element-by-element multiplication module.

[0022] Preferably, in step S2, the multi-level residual enhancement module includes three multi-level residual connection modules connected in sequence, the input end of the first multi-level residual connection module and the output end of the third multi-level residual connection module are both connected to the Add layer, and the three multi-level residual connection modules have the same structure and function.

[0023] Preferably, in step S2, the multi-level residual connection module includes four convolution units, and the four convolution units are connected by multi-level residual connections. Each convolution unit includes a 3×3 convolution layer, a RELU activation layer and a Concat layer connected in sequence. The output end of the Concat layer of the last convolution unit (i.e., the fourth convolution unit) is connected to the 3×3 convolution layer, and the output end of the 3×3 convolution layer and the input end of the first convolution unit are both connected to the Add layer.

[0024] Preferably, in step S2, the information interaction module includes a convolution layer (the convolution kernel size is 3×3), a maximum pooling layer and an Add layer connected in sequence, the input end of the convolution layer is connected to the output end of the classification network, the input end of the Add layer is connected to the output end of the spatial attention module, and the output end of the Add layer is connected to the Sigmoid layer.

[0025] Preferably, in step S2, the attention-based multi-scale fusion module includes a 5×5 convolution layer and a first 3×3 convolution layer, the input end of the 5×5 convolution layer and the input end of the first 3×3 convolution layer are both connected to the output end of the second element-by-element multiplication unit connected to the output end of the channel attention module, the output end of the 5×5 convolution layer is connected to the third element-by-element multiplication unit, the output end of the first 3×3 convolution layer is connected to the fourth element-by-element multiplication unit, the output end of the 5×5 convolution layer and the output end of the first 3×3 convolution layer are also connected to the first Add layer, the first Add layer is connected to the spatial attention module, and the spatial attention module is respectively connected to the input end of the third element-by-element multiplication unit and the input end of the fourth element-by-element multiplication unit. The output ends of the third element-by-element multiplication unit and the fourth element-by-element multiplication unit are also connected to the second Add layer, the output end of the second Add layer and the output end of the upsampling module are both connected to the third Add layer, the third Add layer is connected to the channel attention module, the output end of the upsampling module is also connected to the second 3×3 convolution layer, the second 3×3 convolution layer is connected to the fifth element-by-element multiplication unit, the second Add layer is also connected to the sixth element-by-element multiplication unit, the output end of the channel attention module is respectively connected to the input end of the fifth element-by-element multiplication unit and the input end of the sixth element-by-element multiplication unit, the output end of the fifth element-by-element multiplication unit and the output end of the sixth element-by-element multiplication unit are both connected to the fourth Add layer.

[0026] Preferably, in step S3, the total generation loss of the image fusion network based on classification guidance and multi-level residual enhancement is Including fusion loss and classification loss , the total generation loss of image fusion network based on classification guidance and multi-level residual enhancement The calculation method is as shown in formula (1):

[0027] (1)

[0028] In formula (1), and are all hyperparameters, and are set to 0.7 and 0.3 respectively.

[0029] Preferably, step S4 specifically includes the following steps:

[0030] S4-1, obtain training set B based on training set A;

[0031] According to the ground object categories, the original low-resolution hyperspectral images used for training are divided respectively, and then 10% of the images and their corresponding labels are randomly selected from the original low-resolution hyperspectral images corresponding to the 9 ground object categories as the original low-resolution hyperspectral sub-images used for training; according to the ground object categories, the original high-resolution multispectral images used for training are divided respectively, and then 10% of the images and their corresponding labels are randomly selected from the original high-resolution multispectral images corresponding to the 9 ground object categories as the original high-resolution multispectral sub-images used for training; the original low-resolution hyperspectral sub-images used for training and the original high-resolution multispectral sub-images used for training constitute the training set B;

[0032] In the training process of the hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement described in the present application, a training set B is input in each training epoch, and the training sets B input in different training epochs are all obtained by using step S4-1;

[0033] S4-2, input the training set B obtained in step S4-1 into the image fusion network based on classification guidance and multi-level residual enhancement of the present application, perform forward propagation, and calculate the fusion loss respectively and classification loss , in total generated losses Back propagation is performed under the guidance of , and the weight parameters and hyper parameters of the image fusion network based on classification guidance and multi-level residual enhancement are updated. Among them, the fusion loss Update of weight parameters and hyperparameters of constrained multi-level residual enhancement module, information interaction module and attention-based multi-scale fusion module; classification loss The weight parameters and hyperparameters of the constrained classification network are updated; after 2000 epochs, the training process of the image fusion network based on classification guidance and multi-level residual enhancement is completed, and the image fusion network model based on classification guidance and multi-level residual enhancement is obtained; among them, the hyperparameters include learning rate and number of iterations.

[0034] Compared with the prior art, the beneficial technical effects of this application are:

[0035] The hyperspectral and multispectral image fusion method described in the present application is based on the reconstructed high-resolution hyperspectral image obtained from the image used for testing obtained from the Pavia University data set, and the deviation between the reconstructed high-resolution hyperspectral image and the label corresponding to the source image is small. It is more outstanding in reconstructing image details and can have higher restoration accuracy; moreover, the hyperspectral and multispectral image fusion method described in the present application has strong robustness and adaptability when fusing source images with more background noise, and can effectively suppress the interference of background noise on image fusion quality; in addition, the hyperspectral and multispectral image fusion method described in the present application has strong discrimination ability and reconstruction accuracy for different categories of ground objects, and effectively improves the fidelity of spectral characteristics of each category of ground objects in the process of hyperspectral and multispectral image fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a schematic diagram of the structure of the image fusion network based on classification guidance and multi-level residual enhancement in this application;

[0037] Figure 2 This is a schematic diagram of the network structure of the band embedding module in this application;

[0038] Figure 3 This is a schematic diagram of the network structure of the spatial attention module in this application;

[0039] Figure 4 This is a schematic diagram of the network structure of the channel attention module in this application;

[0040] Figure 5 This is a schematic diagram of the network structure of the multi-level residual enhancement module in this application;

[0041] Figure 6 This is a schematic diagram of the network structure of the information interaction module in this application;

[0042] Figure 7 Schematic diagram of the network structure of the attention-based multi-scale fusion module in this application. DETAILED DESCRIPTION

[0043] The present invention will be further described below with reference to the accompanying drawings and examples;

[0044] A hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement includes the following steps:

[0045] S1: Preprocess the images in the Pavia University dataset and the Pavia Center dataset to obtain the training set A and the test set, which specifically includes the following steps;

[0046] The Pavia University dataset and Pavia Center dataset in this application are both existing datasets. The URL for obtaining the Pavia University dataset is: https: / / github.com / li-yapeng / MambaHSI. The Pavia University dataset contains a total of 42,776 pixels of ground objects, and includes 9 ground object categories. The URL for obtaining the PaviaCenter dataset is: https: / / github.com / liangjiandeng / HSRnet. The Pavia Center dataset contains a total of 7,456 pixels of ground objects, and includes 9 ground object categories.

[0047] S1-1. Obtain original low-resolution hyperspectral images for training and original low-resolution hyperspectral images for testing based on images in the Pavia University dataset; specifically, the steps include:

[0048] The images in the Pavia University dataset are processed with Gaussian blur and down-sampled in turn to obtain a preliminary preprocessed image. Sample points are set in the preliminary preprocessed image. The number of sample points is the same as the number of ground object pixels contained in the preliminary preprocessed image.

[0049] 10% of the sample points are randomly selected from the above sample points, and the preliminary preprocessed images are segmented with the selected sample points as the center and the size of 16×16 to obtain the original low-resolution hyperspectral image for training; the remaining 90% of the sample points in the preliminary preprocessed image are segmented with the size of 16×16 to obtain the original low-resolution hyperspectral image for testing;

[0050] In this application, the Pavia University dataset contains a total of 42776 pixels of ground objects, so the number of ground object pixels contained in the preliminary preprocessed image obtained by the Pavia University dataset is 42776, and the number of sample points of the preliminary preprocessed image obtained by the Pavia University dataset is 42776;

[0051] S1-2, based on the images in the Pavia Center dataset, obtaining original low-resolution hyperspectral images for training and original low-resolution hyperspectral images for testing; specifically including the following steps:

[0052] The images in the Pavia Center dataset are processed with Gaussian blur and down-sampled in turn to obtain a preliminary preprocessed image. Sample points are set in the preliminary preprocessed image. The number of sample points is the same as the number of ground object pixels contained in the preliminary preprocessed image.

[0053] 10% of the sample points are randomly selected from the above sample points, and the preliminary preprocessed images are segmented with the selected sample points as the center and the size of 16×16 to obtain the original low-resolution hyperspectral image for training; the remaining 90% of the sample points in the preliminary preprocessed image are segmented with the size of 16×16 to obtain the original low-resolution hyperspectral image for testing;

[0054] In this application, since the Pavia Center dataset contains a total of 7456 pixels of ground objects, the number of ground object pixels contained in the preliminary preprocessed image obtained by the Pavia Center dataset is 7456, and the number of sample points of the preliminary preprocessed image obtained by the Pavia Center dataset is 7456;

[0055] S1-3, original high-resolution multispectral images for training and original high-resolution multispectral images for testing based on images in the Pavia University dataset; specifically includes the following steps:

[0056] The images in the Pavia University dataset are used to obtain a single-process high-resolution multispectral image through a spectral response function; sample points are set in the single-process high-resolution multispectral image, and the number of sample points is the same as the number of ground object pixels contained in the single-process high-resolution multispectral image;

[0057] 10% of the sample points are randomly selected from the above sample points, and the preliminary preprocessed images are segmented with the selected sample points as the center and the size of 64×64 to obtain the high-resolution multispectral image for training; the remaining 90% of the sample points in the once processed high-resolution multispectral image are segmented with the size of 64×64 to obtain the original high-resolution multispectral image for testing;

[0058] The spectral response function used in this application is consistent with the IKONOS2 spectral response function disclosed at https: / / nwp-saf.eumetsat.int / site / software / rttov / download / coefficients / spectral-response-functions / #mw;

[0059] S1-4, original high-resolution multispectral images for training and original high-resolution multispectral images for testing based on images in the Pavia Center dataset; specifically includes the following steps:

[0060] The images in the Pavia Center dataset are used to obtain a single-processed high-resolution multispectral image through the spectral response function; sample points are set in the single-processed high-resolution multispectral image, and the number of sample points is the same as the number of ground object pixels contained in the single-processed high-resolution multispectral image; 10% of the sample points are randomly selected from the above sample points, and the preliminary preprocessed image is segmented with the selected sample points as the center and the size of 64×64 to obtain a high-resolution multispectral image for training; the single-processed high-resolution multispectral image is segmented with the remaining 90% of the sample points in the single-processed high-resolution multispectral image as the center and the size of 64×64 to obtain the original high-resolution multispectral image for testing.

[0061] The original low-resolution hyperspectral images used for training and the high-resolution multispectral images used for training in this application together constitute a training set A; the original low-resolution hyperspectral images used for testing in this application and the high-resolution multispectral images used for testing together constitute a test set.

[0062] S2, construct an image fusion network based on classification guidance and multi-level residual enhancement;

[0063] The structure of the image fusion network based on classification guidance and multi-level residual enhancement is as follows: Figure 1 As shown, Figure 1 middle, represents an element-by-element multiplication module. The image fusion network based on classification guidance and multi-level residual enhancement includes a band embedding module and a spatial attention module, and the input end of the band embedding module is connected to the input end of the spatial attention module;

[0064] The band embedding module is connected to the first 3×3 convolutional layer, and the first 3×3 convolutional layer is respectively connected to the multi-level residual enhancement module and the classification network. The image output by the classification network is the output result of the classification task; the classification network and the spatial attention module are both connected to the information interaction module, and the multi-level residual enhancement module and the information interaction module are both connected to the first element-by-element multiplication module, and the first element-by-element multiplication module is connected to the second 3×3 convolutional layer;

[0065] The image fusion network based on classification guidance and multi-level residual enhancement also includes an upsampling module and a channel attention module. The input end of the upsampling module, the input end of the channel attention module and the input end of the band embedding module are connected, the output end of the channel attention module and the output end of the second 3×3 convolutional layer are both connected to the second element-by-element multiplication unit, the output end of the upsampling module and the second element-by-element multiplication module are both connected to the attention-based multi-scale fusion module, and the output image of the attention-based multi-scale fusion module is the reconstructed high-resolution hyperspectral image.

[0066] Among them, the band embedding module is used to perform preliminary fusion of the input original low-resolution hyperspectral image and the original high-resolution multispectral image to obtain a preliminary fusion image that complements the spatial and spectral information, and transmit it to the first 3×3 convolutional layer;

[0067] The spatial attention module is used to adaptively adjust the weights of different positions in the input original high-resolution multispectral image, focus on the key areas in the original high-resolution multispectral image, improve the ability to extract spatial information in the original high-resolution multispectral image, and obtain a spatial feature map that can highlight important spatial details in the image; and transmit the spatial feature map to the information interaction module;

[0068] The channel attention module is used to adaptively weight each channel of the input original low-resolution hyperspectral image, enhance the spectral information feature representation of the key channels, and suppress unimportant channels to obtain a feature map with enhanced spectral details. This feature map mainly highlights the subtle information in the spectrum, which helps to improve the spectral resolution and spectral authenticity of the image.

[0069] The upsampling module is used to perform bicubic interpolation upsampling on the input original low-resolution hyperspectral image, with the aim of effectively increasing the resolution of the low-resolution hyperspectral image while ensuring the clarity of the original low-resolution hyperspectral image, and obtaining a low-resolution hyperspectral image that is four times larger than the original low-resolution hyperspectral image; and the image is transmitted to the attention-based multi-scale fusion module;

[0070] The first 3×3 convolutional layer is used to perform local feature extraction on the feature map output by the band embedding module to obtain a feature map with local information enhancement;

[0071] The multi-level residual enhancement module is used to further fuse the feature map output by the first 3×3 convolutional layer to extract rich spectral information in the hyperspectral image and rich spatial information in the multispectral image. It can also better handle the problem of information redundancy and obtain a residual map containing rich spatial spectral information.

[0072] The classification network is used to classify the feature map output by the first 3×3 convolutional layer. At the same time, it can also capture the spatial information and spectral information in the feature map output by the first 3×3 convolutional layer, output the classification result and the feature map with spatial information and spectral information; and the feature map with spatial information and spectral information is transmitted to the information interaction module;

[0073] An information interaction module is used to integrate the spatial feature map output by the spatial attention and the feature map output by the classification network to obtain an attention feature map with enhanced features; since the information interaction module integrates the feature map with spatial information and spectral information output by the classification network and the spatial feature map output by the spatial attention to obtain an attention feature map with enhanced features, and the attention feature map with enhanced features can continue to be used in subsequent image fusion processes, therefore, the setting of the information interaction module enables the present application to realize the function of guiding the image fusion task by using the classification task;

[0074] The first element-by-element multiplication module is used to multiply the feature map output by the multi-level residual enhancement module and the feature map output by the information interaction module element by element to obtain a feature map that integrates multiple layers of feature information and highlights important features;

[0075] The second 3×3 convolutional layer is used to perform a convolution operation on the feature map output by the first element-by-element multiplication unit to enhance the detail features and obtain a feature map with more prominent detail features.

[0076] The second element-by-element multiplication unit is used to perform element-by-element multiplication on the feature map output by the channel attention module and the feature map output by the second 3×3 convolutional layer to obtain a feature map with enhanced spatial spectral information, and input it into the attention-based multi-scale fusion module;

[0077] The attention-based multi-scale fusion module is used to capture features at different levels at multiple scales of the low-resolution hyperspectral image output by the upsampling module and the feature map enhanced with spatial spectral information output by the second element-by-element multiplication unit, and effectively weight the features of each scale through the attention mechanism to obtain a reconstructed high-resolution hyperspectral image.

[0078] The structure and function of the band embedding module in this application:

[0079] In this application, the network structure of the band embedding module is as follows: Figure 2 As shown, Figure 2 middle Represents a Concat layer. The band embedding module includes a downsampling module, the input of the downsampling module is connected to the input of the spatial attention module, the output of the downsampling module is connected to the first Concat layer, the input of the first Concat layer is connected to the input of the channel attention module, the output of the first Concat layer is connected to the pixel shuffle module, the output of the pixel shuffle module and the input of the downsampling module are both connected to the second Concat layer, and the output of the second Concat layer serves as the output of the band embedding module; in this application, the pixel shuffle module has the same structure and function as the pixelShuffle module disclosed in the paper "Hyperspectral Image Super-resolution via DeepSpatio-spectral Attention Convolutional Neural Networks".

[0080] In the present application, the functions of the band embedding module are as follows: the downsampling module is used to downsample the original high-resolution multispectral image to obtain a multispectral image, and the downsampling factor is set to 4; the first Concat layer is used to splice the original low-resolution hyperspectral image and the multispectral image output by the downsampling module to obtain a feature map with a size of 16×16×97 and containing spatial information and spectral information; the pixel shuffling module is used to spatially rearrange and upsample the feature map output by the first Concat layer to obtain a feature map with a size of 64×64×93, which contains high-resolution spatial information extended from the low-resolution feature map and retains the detailed structure of the image; the second Concat layer is used to splice the original high-resolution multispectral image and the feature map output by the pixel shuffling module to complete the process of preliminary fusion of the original low-resolution hyperspectral image and the original high-resolution multispectral image, and obtain a preliminary fused image with complementary spatial information and spectral information.

[0081] The structure and function of the spatial attention module in this application:

[0082] The network structure of the spatial attention module in this application is as follows: Figure 3 As shown, Figure 3 middle, Represents the Concat layer, represents the Sgmoid layer, Represents an element-by-element multiplication module. In the present application, the spatial attention module includes a maximum pooling layer and an average pooling layer, the input of the maximum pooling layer is connected to the input of the average pooling layer, the output of the maximum pooling layer and the output of the average pooling layer are both connected to the Concat layer, and the Concat layer is also sequentially connected to the first convolution layer (the convolution kernel size is 3×3), the first Sigmoid layer, the second convolution layer (the convolution kernel size is 3×3) and the second Sigmoid layer, and the output of the second Sigmoid layer and the input of the maximum pooling layer are both connected to the input of the element-by-element multiplication unit.

[0083] In the spatial attention module, the maximum pooling layer is used to perform the maximum pooling operation on the input feature map in the spatial dimension, so that the network can focus on the higher-level abstract features of the key areas in the original high-resolution multispectral image while maintaining the key features; the average pooling layer performs the average pooling operation on the input feature map in the spatial dimension and calculates the average value of all values ​​in the feature window; the Concat layer is used to splice the feature map output by the maximum pooling layer and the feature map output by the average pooling layer to fuse local and global information; the first convolutional layer is used to perform feature fusion and detail extraction operations on the feature map output by the Concat layer to obtain a feature map with local feature enhancement ; The first Sgmoid layer is used to activate the feature map output by the first convolutional layer; the second convolutional layer is used to perform a convolution operation on the feature map output by the first Sgmoid layer to obtain a feature map with a more prominent spatial weight distribution; the second Sgmoid layer is used to activate the feature map output by the second convolutional layer; the element-by-element multiplication unit is used to perform an element-by-element multiplication operation on the feature map of the input spatial attention module and the feature map output by the second Sgmoid layer to adaptively adjust the weights of different positions in the input original high-resolution multispectral image, and obtain a spatial feature map that can highlight important spatial details in the image.

[0084] The structure and function of the channel attention module in this application:

[0085] The network structure of the channel attention module in this application is as follows: Figure 4 As shown, Figure 4 middle Indicates the Add layer, represents the Sgmoid layer, Represents an element-by-element multiplication module. In the present application, the channel attention module includes a maximum pooling layer and an average pooling layer. The input of the maximum pooling layer is connected to the input of the average pooling layer. The maximum pooling layer is also connected to the first fully connected layer, the first ReLU activation layer, and the second fully connected layer in sequence. The average pooling layer is connected to the third fully connected layer, the second ReLU activation layer, and the fourth fully connected layer in sequence. The outputs of the second fully connected layer and the fourth fully connected layer are both connected to the Add layer. The Add layer is connected to the convolution layer (the convolution kernel size is 3×3) and the Sigmoid layer in sequence. The output of the Sigmoid layer and the input of the maximum pooling layer are both connected to the element-by-element multiplication module.

[0086] In the channel attention module of the present application, the maximum pooling layer is used to perform the maximum pooling operation on the input original low-resolution hyperspectral image in the channel dimension, so that the network can focus on the spectral information characteristics of the key channel while maintaining the key features; the average pooling layer performs the average pooling operation on the input original low-resolution hyperspectral image in the channel dimension, and calculates the average value of all values ​​in the feature window; the feature map output by the maximum pooling layer will be processed by the first fully connected layer, the first RELU activation layer and the second fully connected layer in sequence, and the feature map output by the average pooling layer will be processed by the third fully connected layer, the second RELU activation layer and the fourth fully connected layer in sequence. In the present application, the fully connected layer is parameterized and mapped to each channel. The weights are assigned, and the RELU activation layer introduces nonlinear expression capabilities, so that the importance of different channels can be optimized, the spectral information feature representation of key channels can be enhanced, and unimportant channels can be suppressed; the feature map output by the second fully connected layer and the feature map output by the fourth fully connected layer are added using the Add layer; the convolution layer is used to integrate the feature map output by the Add layer, and the Sgmoid layer is used to activate the feature map output by the convolution layer; the element-by-element multiplication module is used to perform element-by-element multiplication of the feature map output by the Sgmoid layer and the original low-resolution high-spectral image of the input channel attention module to obtain a feature map with enhanced spectral details, which mainly highlights the subtle information in the spectrum.

[0087] The structure and function of the multi-level residual enhancement module in this application:

[0088] In this application, the network structure of the multi-level residual enhancement module is as follows: Figure 5 As shown, Figure 5 middle Indicates the Add layer, In the present application, the multi-level residual enhancement module includes three multi-level residual connection modules connected in sequence, the input end of the first multi-level residual connection module and the output end of the third multi-level residual connection module are both connected to the Add layer, and the three multi-level residual connection modules have the same structure and function;

[0089] In the present application, the multi-level residual connection module includes four convolution units, which are connected by multi-level residual connections. Each convolution unit includes a 3×3 convolution layer, a RELU activation layer and a Concat layer connected in sequence. The output end of the Concat layer of the last convolution unit (i.e., the fourth convolution unit) is connected to the 3×3 convolution layer, and the output end of the 3×3 convolution layer and the input end of the first convolution unit are both connected to the Add layer.

[0090] In the present application, the four convolution units are connected by means of multi-level residual connections, as shown below: the input end of the first convolution unit is respectively connected to the input end of the Concat layer in the first convolution unit, the input end of the Concat layer in the second convolution unit, the input end of the Concat layer in the third convolution unit, and the input end of the Concat layer in the fourth convolution unit; the output end of the Concat layer in the first convolution unit is respectively connected to the input end of the Concat layer in the second convolution unit, the input end of the Concat layer in the third convolution unit, and the input end of the Concat layer in the fourth convolution unit; the output end of the Concat layer in the second convolution unit is respectively connected to the input end of the Concat layer in the third convolution unit and the input end of the Concat layer in the fourth convolution unit; the output end of the Concat layer in the third convolution unit is connected to the input end of the Concat layer in the fourth convolution unit.

[0091] In the present application, the function of the multi-level residual connection module is as follows: the input feature map is processed by the convolution layer and RELU activation layer of the first convolution unit to obtain the first layer residual map, and input it into the Concat layer of the first convolution unit; the Concat layer of the first convolution unit splices the first layer residual map and the feature map input to the first convolution unit to obtain the first spliced ​​map;

[0092] The first spliced ​​image is input into the convolution layer of the second convolution unit, the Concat layer of the second convolution unit, the Concat layer of the third convolution unit, and the Concat layer of the fourth convolution unit respectively; the feature map output by the Concat layer in the first convolution unit is processed by the convolution layer and RELU activation layer of the second convolution unit to obtain the second layer residual map, which is input into the Concat layer of the second convolution unit. The Concat layer of the second convolution unit splices the second layer residual map, the first spliced ​​image, and the feature map input into the first convolution unit to obtain the second spliced ​​image;

[0093] The second spliced ​​image is input into the Concat layer of the third convolution unit and the Concat layer of the fourth convolution unit respectively; the second spliced ​​image output by the Concat layer of the second convolution unit is processed by the convolution layer and RELU activation layer of the third convolution unit to obtain the third layer residual image, which is input into the Concat layer of the third convolution unit; the Concat layer of the third convolution unit splices the third layer residual image, the second spliced ​​image, the first spliced ​​image and the feature image input into the first convolution unit to obtain the third spliced ​​image;

[0094] The third splicing image is input into the Concat layer of the fourth convolution unit; the third splicing image output by the third Concat layer is processed by the convolution layer and RELU activation layer of the fourth convolution unit to obtain the fourth layer residual image, which is input into the Concat layer of the fourth convolution unit; the Concat layer of the fourth convolution unit splices the fourth layer residual image, the third splicing image, the second splicing image, the first splicing image and the feature image input into the first convolution unit to obtain the fourth splicing image;

[0095] The fourth spliced ​​image is processed by convolution of 3×3 convolution layers, and then added with the feature map of the first convolution unit input using the Add layer to obtain the output result of the first multi-level residual connection module; in the present application, the multi-level residual connection module gradually extracts rich spectral information in the hyperspectral image and rich spatial information in the multi-spectral image by accumulating and weighted fusion features layer by layer, effectively reducing redundant information. In the present application, the residual connection setting between the multi-level residual connection modules enables the present application to optimize the feature expression through residual learning, and finally obtain a residual map containing rich spatial information and spectral information.

[0096] The structure and function of the classification network in this application:

[0097] The classification network used in this application has the same structure and function as the HybridSpectralNet (HybridSN for short) disclosed in the paper "HybridSN: Exploring 3-D-2-D CNN FeatureHierarchy for Hyperspectral Image Classification".

[0098] The structure and function of the information interaction module in this application:

[0099] The structure of the information interaction module, such as Figure 6 As shown, Figure 6 middle Indicates the Add layer, In this application, the information interaction module includes a convolution layer (the convolution kernel size is 3×3), a maximum pooling layer, and an Add layer connected in sequence. The input end of the convolution layer is connected to the output end of the classification network, the input end of the Add layer is connected to the output end of the spatial attention module, and the output end of the Add layer is connected to the Sigmoid layer.

[0100] The feature map output by the classification network is processed by the convolution layer and the maximum pooling layer in turn, and then added to the feature map output by the spatial attention module through the Add layer to achieve information integration, and then activated by the Sgmoid layer to obtain an attention feature map with feature enhancement.

[0101] The structure and function of the attention-based multi-scale fusion module in this application:

[0102] In this application, the structure of the attention-based multi-scale fusion module is as follows: Figure 7 As shown, Figure 7 middle, Indicates the Add layer, In this application, the attention-based multi-scale fusion module includes a 5×5 convolution layer and the first 3×3 convolution layer. The input end of the 5×5 convolution layer and the input end of the first 3×3 convolution layer are connected to Figure 2 The output of the second element-by-element multiplication unit connected to the output of the neutral and channel attention modules is connected, the output of the 5×5 convolution layer is connected to the third element-by-element multiplication unit, the output of the first 3×3 convolution layer is connected to the fourth element-by-element multiplication unit, the output of the 5×5 convolution layer and the output of the first 3×3 convolution layer are also connected to the first Add layer, the first Add layer is connected to the spatial attention module, the spatial attention module is respectively connected to the input of the third element-by-element multiplication unit and the input of the fourth element-by-element multiplication unit, the output of the third element-by-element multiplication unit and the output of the fourth element-by-element multiplication unit are also connected to the first The output of the second Add layer and the output of the upsampling module are connected to the third Add layer, the third Add layer is connected to the channel attention module, the output of the upsampling module is also connected to the second 3×3 convolution layer, the second 3×3 convolution layer is connected to the fifth element-by-element multiplication unit, the second Add layer is also connected to the sixth element-by-element multiplication unit, the output of the channel attention module is respectively connected to the input of the fifth element-by-element multiplication unit and the input of the sixth element-by-element multiplication unit, the output of the fifth element-by-element multiplication unit and the output of the sixth element-by-element multiplication unit are both connected to the fourth Add layer. The spatial attention module and the multi-scale fusion module in the attention-based multi-scale fusion module of this application Figure 1The structure and function of the spatial attention module in the present application are the same; the channel attention module and Figure 1 The structure of the mid-channel attention module is the same and its function is the same;

[0103] In this application, the functions of the attention-based multi-scale fusion module are:

[0104] The 5×5 convolutional layer is used to Figure 2 The feature map output by the second element-by-element multiplication unit in the image extracts wider features in the image through a larger receptive field and smoothes the image features; the first 3×3 convolutional layer is used to Figure 2 The feature map output by the second element-by-element multiplication unit in the image extracts local detail features and enhances the detail information and texture features in the image; the first Add layer is used to add the feature map output by the first 3×3 convolution layer and the feature map output by the 5×5 convolution layer to achieve the integration of features of different receptive fields; the third element-by-element multiplication module is used to perform element-by-element multiplication of the spatial attention map output by the spatial attention module and the feature map output by the 5×5 convolution layer to achieve the integration of wider spatial information in the image; the fourth element-by-element multiplication module is used to perform element-by-element multiplication of the spatial attention map output by the spatial attention module and the feature map output by the first 3×3 convolution layer to achieve the integration of local spatial information; the second Add layer is used to add the feature map output by the third element-by-element multiplication module and the feature map output by the fourth element-by-element multiplication module to achieve the integration of global information and local information. The third Add layer is used to add the feature maps output by the second Add layer and the upsampling module to complete the spectral information; the second 3×3 convolutional layer is used to perform a convolution operation on the feature map output by the upsampling module to refine the spectral information; the fifth element-by-element multiplication module is used to perform element-by-element multiplication on the channel attention map output by the channel attention module and the feature map output by the second 3×3 convolutional layer to further integrate the spectral information; the sixth element-by-element multiplication module performs element-by-element multiplication on the channel attention map output by the channel attention module and the feature map output by the second Add layer to achieve joint optimization of spatial spectral information; the fourth Add layer is used to perform feature addition on the feature map output by the fifth element-by-element multiplication module and the feature map output by the sixth element-by-element multiplication module to achieve the final optimization of the global spatial spectral information and obtain a reconstructed high-resolution hyperspectral image.

[0105] S3. Constructing the total generation loss of the image fusion network based on classification guidance and multi-level residual enhancement ;

[0106] The total generation loss of the image fusion network based on classification guidance and multi-level residual enhancement in this application Including fusion loss and classification loss , the total generation loss of image fusion network based on classification guidance and multi-level residual enhancement The calculation method is as shown in formula (1):

[0107] (1)

[0108] In formula (1), and All are hyperparameters. and are set to 0.7 and 0.3 respectively.

[0109] In formula (1), the fusion loss is the L1 norm, and its calculation formula is shown in formula (2):

[0110] (2)

[0111] In formula (2), Z represents the reconstructed high-resolution hyperspectral image, X represents the input low-resolution hyperspectral image for training, represents the L1 norm.

[0112] In formula (1), the classification loss is the cross entropy loss, and its calculation formula is shown in formula (3):

[0113] (3)

[0114] In formula (3), is the number of categories, is a hot encoding value of the true label of the sample. If the sample belongs to the category ,but =1, otherwise =0, The classification network in the image fusion network based on classification guidance and multi-level residual enhancement predicts that the sample belongs to the category The above samples refer to the total number of original low-resolution hyperspectral images used for training and high-resolution multispectral images used for training.

[0115] S4, obtaining a training set B based on the training set A, and using the training set B to train the image fusion network based on classification guidance and multi-level residual enhancement to obtain an image fusion network model based on classification guidance and multi-level residual enhancement; specifically comprising the following steps:

[0116] S4-1, obtain training set B based on training set A;

[0117] According to the ground object categories, the original low-resolution hyperspectral images used for training are divided respectively, and then 10% of the images and their corresponding labels are randomly selected from the original low-resolution hyperspectral images corresponding to the 9 ground object categories as the original low-resolution hyperspectral sub-images used for training; according to the ground object categories, the original high-resolution multispectral images used for training are divided respectively, and then 10% of the images and their corresponding labels are randomly selected from the original high-resolution multispectral images corresponding to the 9 ground object categories as the original high-resolution multispectral sub-images used for training; the original low-resolution hyperspectral sub-images used for training and the original high-resolution multispectral sub-images used for training constitute the training set B;

[0118] In the training process of the hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement described in the present application, a training set B is input in each training epoch, and the training sets B input in different training epochs are all obtained by using step S4-1;

[0119] S4-2, input the training set B obtained in step S4-1 into the image fusion network based on classification guidance and multi-level residual enhancement of the present application, perform forward propagation, and calculate the fusion loss respectively and classification loss , in total generated losses Back propagation is performed under the guidance of , and the weight parameters and hyper parameters of the image fusion network based on classification guidance and multi-level residual enhancement are updated. Among them, the fusion loss Update of weight parameters and hyperparameters of constrained multi-level residual enhancement module, information interaction module and attention-based multi-scale fusion module; classification loss The weight parameters and hyperparameters of the constrained classification network are updated; after 2000 epochs, the training process of the image fusion network based on classification guidance and multi-level residual enhancement is completed, and the image fusion network model based on classification guidance and multi-level residual enhancement is obtained; wherein, the hyperparameters include the learning rate and the number of iterations;

[0120] In this embodiment, the training process of the image fusion network based on classification guidance and multi-level residual enhancement is implemented based on the NVIDIA 2080Ti GPU chip, and the Adam optimizer is used to optimize the loss gradient and back propagate. In the training process of the image fusion network based on classification guidance and multi-level residual enhancement, the batch size is set to 64, and the learning rates of the multi-level residual enhancement module, the information interaction module, the attention-based multi-scale fusion module, and the classification network are all set to 1×10 -4 The learning rate mode of the multi-level residual enhancement module, information interaction module, attention-based multi-scale fusion module and classification network adopts exponential learning rate decay, where the decay coefficient gamma is set to 0.998.

[0121] S5. Input the low-resolution hyperspectral image to be processed and the high-resolution multispectral image to be processed into the trained image fusion network model based on classification guidance and multi-level residual enhancement. After forward propagation once, the predicted reconstructed high-resolution hyperspectral image and classification result can be obtained.

[0122] test:

[0123] In order to compare the fusion effect of the hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement described in this application, this application compares the four existing hyperspectral and multispectral image fusion methods with the hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement described in this application, wherein the existing hyperspectral and multispectral image fusion methods include: HSRnet method (from "Hyperspectral Image Super-resolution via Deep Spatio-spectral Attention Convolutional Neural Networks"), U2Net method (from "U2Net: A General Framework with Spatial-Spectral-Integrated Double U-Netfor Image Fusion"), PSRT method (from "PSRT: Pyramid Shuffle-and-Reshuffle Transformer for Multispectral and Hyperspectral Image Fusion"), KNL method (from "KNLConv: Kernel-space Non-local Convolution for Hyperspectral Image Fusion"), and the like. The above four existing hyperspectral and multispectral image fusion methods and the hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement described in this application are tested using the test set constructed in this application to obtain reconstructed high-resolution hyperspectral images. Then, the reconstructed high-resolution hyperspectral images and their corresponding labels are respectively input into the PyTorch-based hyperspectral and multispectral image fusion evaluation tool shown in https: / / github.com / PSRben / U2Net to output the evaluation indicators. 、 、 and , evaluation index 、 、 and The results are shown in Table 1. The hyperspectral and multispectral image fusion method described in this application is represented by Ours method in Table 1.

[0124] Table 1 shows the results of testing different hyperspectral and multispectral image fusion methods

[0125] The evaluation indicators used in this application include 、 、 and ,in, A metric used to measure the quality of each band of a reconstructed high-resolution hyperspectral image, usually expressed in decibels (dB); It is used to measure the angle between the spectral vectors of each pixel, reflecting the similarity of the hyperspectral images in the spectral dimension; Indicators used to comprehensively consider the spatial and spectral information of images; It measures the structural similarity of images and evaluates them by combining brightness, contrast and structural information. Among the above four test indicators, and The smaller the value, the better. The larger the value, the better.

[0126] The most challenging part of this application is to test the images obtained from the Pavia University dataset in the test set, because the background information in these images is complex, including multiple types of objects, such as buildings, roads, vegetation, and water bodies, which are densely distributed and have fuzzy category boundaries. However, from the test results shown in Table 1, the hyperspectral and multispectral image fusion method described in this application performs well in the four evaluation indicators when testing the images obtained from the Pavia University dataset.

[0127] Since the above four existing hyperspectral and multispectral image fusion methods are tested, the PSRT method is 、 、 and The test results are good. Specifically, the PSRT method is used to test the and The lowest value, The value is the highest. Therefore, for the images obtained from the Pavia University dataset for testing, this application focuses on comparing the test results of the PSRT method with the hyperspectral and multispectral image fusion method described in this application, as follows:

[0128] Compared with the PSRT method, the hyperspectral and multispectral image fusion method described in this application is The evaluation index increased by (46.4467-45.0769) / 46.4467=2.95%;

[0129] Compared with the PSRT method, the hyperspectral and multispectral image fusion method described in this application is The evaluation index decreased by (1.1374-0.9492) / 1.1374=16.55%;

[0130] Compared with the PSRT method, the hyperspectral and multispectral image fusion method described in this application is The evaluation index decreased by (1.2511-1.0074) / 1.2511=19.48%;

[0131] Compared with the PSRT method, the hyperspectral and multispectral image fusion method described in this application is The evaluation index increased by (0.9866-0.9881) / 0.9866=0.15%;

[0132] Obviously, from the above test results, it can be seen that the hyperspectral and multispectral image fusion method described in this application can obtain The evaluation index is significantly reduced, which means that the fused image (i.e., the reconstructed high-resolution hyperspectral image) obtained by the hyperspectral and multispectral image fusion method described in the present application when testing the image obtained from the Pavia University dataset for testing is more accurate than the reconstructed high-resolution hyperspectral images obtained by the above four existing hyperspectral and multispectral image fusion methods. The reconstructed high-resolution hyperspectral image obtained by the hyperspectral and multispectral image fusion method described in the present application has less deviation from the label corresponding to the source image, is more outstanding in reconstructing image details, and can have higher restoration accuracy.

[0133] In addition, the hyperspectral and multispectral image fusion method described in this application The evaluation index is also significantly reduced, which shows that the hyperspectral and multispectral image fusion method described in this application has significant advantages in improving the quality of reconstructed high-resolution hyperspectral images. This shows that the hyperspectral and multispectral image fusion method described in this application has strong robustness and adaptability when fusing source images with more background noise, and can effectively suppress the interference of background noise on image fusion quality; in addition, the hyperspectral and multispectral image fusion method described in this application has significant advantages in improving the quality of reconstructed high-resolution hyperspectral images. The significant reduction in the evaluation index also shows that the hyperspectral and multispectral image fusion method described in this application has stronger discrimination ability and reconstruction accuracy for different categories of land objects, and effectively improves the fidelity of the spectral characteristics of each category of land objects in the process of hyperspectral and multispectral image fusion.

[0134] Moreover, when the hyperspectral and multispectral image fusion method described in the present application is tested on the images used for testing obtained from the Pavia University dataset, and There are also significant improvements in these two evaluation indicators, which indicates that the fused image (i.e., the reconstructed high-resolution hyperspectral image) obtained by the hyperspectral and multispectral image fusion method described in this application has excellent performance in terms of the fidelity and structural similarity of spectral information and spatial information. It can not only realize the effective fusion of low-resolution hyperspectral images and high-resolution multispectral images under complex backgrounds, but also guide the optimization of the feature expression of the target area through classification tasks, thereby improving the spectral restoration of different types of objects and the reconstruction quality of spatial details.

[0135] In addition, the hyperspectral and multispectral image fusion method described in this application is tested on the images used for testing obtained from the Pavia Center dataset in the test set, and the test indicators obtained are also significantly improved, as follows:

[0136] For the test of the images obtained from the Pavia Center dataset, the PSRT method was tested among the five existing hyperspectral and multispectral image fusion methods mentioned above. 、 、 and The test results are good. Specifically, the PSRT method is used to test the and The lowest value, The value is the highest. Therefore, for the test of the images obtained from the Pavia Center dataset for testing, this application focuses on comparing the test results of the PSRT method with the hyperspectral and multispectral image fusion method described in this application, as follows:

[0137] Compared with the PSRT method, the hyperspectral and multispectral image fusion method described in this application is The evaluation index increased by (49.3543-49.1757) / 49.3543=0.36%;

[0138] Compared with the PSRT method, the hyperspectral and multispectral image fusion method described in this application is The evaluation index decreased by (1.6911-1.6889) / 1.6911=0.13%;

[0139] Compared with the PSRT method, the hyperspectral and multispectral image fusion method described in this application is The evaluation index decreased by (0.9416-0.9361) / 0.9416=0.58%;

[0140] Compared with the PSRT method, the hyperspectral and multispectral image fusion method described in this application is The evaluation index increased by (0.9947-0.9938) / 0.9947=0.01%.

[0141] Obviously, the hyperspectral and multispectral image fusion method described in this application can also achieve very good results when testing the images obtained from the Pavia Center dataset for testing, especially The evaluation index is significantly reduced, which indicates that the high-resolution hyperspectral image reconstructed by the hyperspectral and multispectral image fusion method described in the present application is improved in global quality and reduced in error when fusing the images used for testing obtained from the Pavia Center dataset, compared with the reconstructed high-resolution hyperspectral images obtained by the above four existing hyperspectral and multispectral image fusion methods. Moreover, the hyperspectral and multispectral image fusion method described in the present application also performs better in detail fidelity, spectral restoration and spatial resolution optimization, and can more accurately restore the true characteristics of different types of land objects.

Claims

1. A hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement, characterized by: The following steps are involved: S1: Preprocess the images in the Pavia University dataset and the Pavia Center dataset to obtain the training set A and the test set; The original low-resolution hyperspectral images used for training and the high-resolution multispectral images used for training together constitute the training set A; S2, construct an image fusion network based on classification guidance and multi-level residual enhancement; The image fusion network based on classification guidance and multi-level residual enhancement includes an upsampling module, a channel attention module, and a parallel band embedding module and a spatial attention module, the band embedding module is connected to the first 3×3 convolution layer, the first 3×3 convolution layer is respectively connected to the multi-level residual enhancement module and the classification network, and the image output by the classification network is the output result of the classification task; the classification network and the spatial attention module are both connected to the information interaction module, the multi-level residual enhancement module and the information interaction module are both connected to the first element-by-element multiplication module, and the first element-by-element multiplication module is connected to the second 3×3 convolution layer; the upsampling module, the channel attention module and the band embedding module are connected in parallel, the channel attention module and the second 3×3 convolution layer are both connected to the second element-by-element multiplication unit, the upsampling module and the second element-by-element multiplication module are both connected to the attention-based multi-scale fusion module, and the output image of the attention-based multi-scale fusion module is the reconstructed high-resolution hyperspectral image; S3. Constructing the total generation loss of the image fusion network based on classification guidance and multi-level residual enhancement ; S4, obtaining a training set B based on the training set A, and using the training set B to train the image fusion network based on classification guidance and multi-level residual enhancement to obtain an image fusion network model based on classification guidance and multi-level residual enhancement; The input data of the band embedding module are the original low-resolution hyperspectral image and the original high-resolution multispectral image, the input data of the upsampling module and the channel attention module are the original low-resolution hyperspectral image, and the input data of the spatial attention module are the original high-resolution multispectral image; S5. Input the low-resolution hyperspectral image to be processed and the high-resolution multispectral image to be processed into the trained image fusion network model based on classification guidance and multi-level residual enhancement. After forward propagation once, the predicted reconstructed high-resolution hyperspectral image and classification result can be obtained.

2. The hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement according to claim 1 is characterized in that: Step S1 specifically includes the following steps: S1-1, based on the images in the Pavia University dataset, obtain the original low-resolution hyperspectral images for training and the original low-resolution hyperspectral images for testing; S1-2, based on the images in the Pavia Center dataset, obtain the original low-resolution hyperspectral images for training and the original low-resolution hyperspectral images for testing; S1-3, original high-resolution multispectral images for training based on images from the Pavia University dataset and original high-resolution multispectral images for testing; S1-4, original high-resolution multispectral images for training based on images in the Pavia Center dataset and original high-resolution multispectral images for testing; The original low-resolution hyperspectral images used for training and the high-resolution multispectral images used for training together constitute the training set A; the original low-resolution hyperspectral images used for testing and the high-resolution multispectral images used for testing together constitute the test set.

3. The hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement according to claim 1 is characterized in that: In step S2, the band embedding module includes a downsampling module, the input end of the downsampling module is connected to the input end of the spatial attention module, the output end of the downsampling module is connected to the first Concat layer, the input end of the first Concat layer is connected to the input end of the channel attention module, the output end of the first Concat layer is connected to the pixel shuffling module, the output end of the pixel shuffling module and the input end of the downsampling module are both connected to the second Concat layer, and the output end of the second Concat layer serves as the output end of the band embedding module.

4. The hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement according to claim 1 is characterized in that: In step S2, the spatial attention module includes a maximum pooling layer and an average pooling layer. The input end of the maximum pooling layer is connected to the input end of the average pooling layer, and the output end of the maximum pooling layer and the output end of the average pooling layer are both connected to the Concat layer. The Concat layer is also connected to the first convolutional layer, the first Sigmoid layer, the second convolutional layer and the second Sigmoid layer in sequence. The output end of the second Sigmoid layer and the input end of the maximum pooling layer are both connected to the input end of the element-by-element multiplication unit.

5. The hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement according to claim 1 is characterized in that: In step S2, the channel attention module includes a maximum pooling layer and an average pooling layer. The input end of the maximum pooling layer is connected to the input end of the average pooling layer. The maximum pooling layer is also connected to the first fully connected layer, the first ReLU activation layer, and the second fully connected layer in sequence. The average pooling layer is connected to the third fully connected layer, the second ReLU activation layer, and the fourth fully connected layer in sequence. The output ends of the second fully connected layer and the fourth fully connected layer are both connected to the Add layer. The Add layer is connected to the convolution layer and the Sigmoid layer in sequence. The output end of the Sigmoid layer and the input end of the maximum pooling layer are both connected to the element-by-element multiplication module.

6. The hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement according to claim 1 is characterized in that: In step S2, the multi-level residual enhancement module includes three multi-level residual connection modules connected in sequence, the input end of the first multi-level residual connection module and the output end of the third multi-level residual connection module are both connected to the Add layer, and the three multi-level residual connection modules have the same structure and function.

7. The method for fusion of hyperspectral and multispectral images based on classification and multi-level residual enhancement according to claim 6, characterized in that: In step S2, the multi-level residual connection module includes four convolution units, which are connected by multi-level residual connections. Each convolution unit includes a 3×3 convolution layer, a RELU activation layer, and a Concat layer connected in sequence. The output end of the Concat layer of the last convolution unit is connected to the 3×3 convolution layer, and the output end of the 3×3 convolution layer and the input end of the first convolution unit are connected to the Add layer.

8. The hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement according to claim 1 is characterized by: In step S2, the information interaction module includes a convolution layer, a maximum pooling layer and an Add layer connected in sequence, the input end of the convolution layer is connected to the output end of the classification network, the input end of the Add layer is connected to the output end of the spatial attention module, and the output end of the Add layer is connected to the Sigmoid layer.

9. The hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement according to claim 1, characterized in that: In step S2, the attention-based multi-scale fusion module includes a 5×5 convolutional layer and a first 3×3 convolutional layer. The input end of the 5×5 convolutional layer and the input end of the first 3×3 convolutional layer are both connected to the output end of the second element-by-element multiplication unit connected to the output end of the channel attention module. The output end of the 5×5 convolutional layer is connected to the third element-by-element multiplication unit. The output end of the first 3×3 convolutional layer is connected to the fourth element-by-element multiplication unit. The output end of the 5×5 convolutional layer and the output end of the first 3×3 convolutional layer are also connected to the first Add layer. The first Add layer is connected to the spatial attention module, and the spatial attention module is respectively connected to the input end of the third element-by-element multiplication unit and the input end of the fourth element-by-element multiplication unit. The output ends of the three element-by-element multiplication units and the output end of the fourth element-by-element multiplication unit are also connected to the second Add layer, the output end of the second Add layer and the output end of the upsampling module are both connected to the third Add layer, the third Add layer is connected to the channel attention module, the output end of the upsampling module is also connected to the second 3×3 convolution layer, the second 3×3 convolution layer is connected to the fifth element-by-element multiplication unit, the second Add layer is also connected to the sixth element-by-element multiplication unit, the output end of the channel attention module is respectively connected to the input end of the fifth element-by-element multiplication unit and the input end of the sixth element-by-element multiplication unit, the output end of the fifth element-by-element multiplication unit and the output end of the sixth element-by-element multiplication unit are both connected to the fourth Add layer.

10. The hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement according to claim 1, characterized in that: Step S4 specifically includes the following steps: S4-1, obtain training set B based on training set A; According to the ground object categories, the original low-resolution hyperspectral images used for training are divided respectively, and then 10% of the images and their corresponding labels are randomly selected from the original low-resolution hyperspectral images corresponding to the nine ground object categories as the original low-resolution hyperspectral sub-images used for training; The original high-resolution multispectral images used for training are divided according to the types of objects, and then 10% of the images and their corresponding labels are randomly selected from the original high-resolution multispectral images corresponding to the nine types of objects as the original high-resolution multispectral sub-images used for training; The original low-resolution hyperspectral sub-images used for training and the original high-resolution multispectral sub-images used for training constitute the training set B; During the training process of the hyperspectral and multispectral image fusion method based on classification and multi-level residual enhancement, a training set B is input in each training epoch, and the training sets B input in different training epochs are all obtained by using step S4-1; S4-2, input the training set B obtained in step S4-1 into the image fusion network based on classification guidance and multi-level residual enhancement of the present application, and perform forward propagation. Back propagation is performed under the guidance of , and the weight parameters and hyperparameters of the image fusion network based on classification guidance and multi-level residual enhancement are updated; after 2000 epochs, the training process of the image fusion network based on classification guidance and multi-level residual enhancement is completed, and the image fusion network model based on classification guidance and multi-level residual enhancement is obtained.

Citation Information

Patent Citations

  • Medium-wave infrared hyperspectral and multispectral image fusion method, system and medium

    CN115018750A

  • Adversarial hyperspectral and multispectral remote sensing fusion method

    CN116468645A