Multispectral and hyperspectral image fusion method

By superimposing the original multispectral and hyperspectral images on the spectral channel and generating fusion images using a VGG-based generator, the spectral distortion problem in the prior art is solved, and high-quality multispectral and hyperspectral images are achieved.

CN119941523APending Publication Date: 2025-05-06CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411811466.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing multispectral and hyperspectral image fusion methods have spectral distortion problems, making it difficult to effectively retain the original image spectrum information.

Method used

A multispectral and hyperspectral image fusion method is proposed. By superimposing the original multispectral image with the original hyperspectral image on the spectral channel, the input data HSI is obtained and inputted into a preset VGG-based generator to generate the fusion image. The generator includes an improved VGG module, a CBAM attention module and a feature recovery module, and optimizes the fusion effect through deep learning technology and dual discriminator structure.

Benefits of technology

Through this method, the spatial and spectral information quality of the fusion image can be effectively improved, spectral distortion can be reduced, the network can learn image details, and the fusion effect can be optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941523A_ABST
    Figure CN119941523A_ABST
Patent Text Reader

Abstract

The invention discloses a multispectral and hyperspectral image fusion method, which comprises the following steps of: superposing an original multispectral image and an original hyperspectral image on a spectral channel to obtain input data HSI; and inputting the input data HSI into a preset VGG-based generator to generate a fused image. The deep learning technology is adopted, so that the space and spectral information of the fused image can be effectively improved; the VGG Block and the CBAM Block are used for feature extraction, so that the learning ability of the model is improved; a double-discriminator structure is introduced, so that space and spectrum adversarial learning is realized, and the fusion effect is further optimized; by adjusting the weight of the loss function, the balance among the losses is realized, and the quality of the fused image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion, and in particular to a multi-spectral and hyper-spectral image fusion method. Background Art

[0002] At present, the fusion methods of hyperspectral and multispectral images mainly include four categories, namely: panchromatic sharpening method, matrix decomposition method, tensor representation method and deep learning-based method. Panchromatic sharpening methods can be divided into component substitution method (CS) and multiscale decomposition method (MSA). The CS-based method generally uses a certain transformation method to transform the LR MSI from the RGB color space to another component space with the original image spatial information, and then replaces this spatial component with a panchromatic image and finally performs an inverse transformation to obtain a fused image. However, this method inevitably changes the spectrum information of the original image, and there is a spectral distortion problem. Summary of the invention

[0003] In order to solve the technical problems existing in the background technology, the present invention proposes a multispectral and hyperspectral image fusion method.

[0004] The present invention proposes a multispectral and hyperspectral image fusion method, comprising:

[0005] The original multispectral image is superimposed with the original hyperspectral image on the spectral channel to obtain the input data HSI;

[0006] The original multispectral image and the original hyperspectral image are input into the preset VGG-based generator to generate a fused image.

[0007] Preferably, the preset VGG-based generator includes an improved VGG module, a CBAM attention module, and a feature recovery module; the inputting the original multispectral image and the original hyperspectral image into the preset VGG-based generator to generate a fused image specifically includes:

[0008] Input the original multispectral image and the original hyperspectral image into the improved VGG module to obtain a first feature map;

[0009] The first feature map is input into the CBAM attention module, and the first feature map is passed through the channel attention unit in the CBAM attention module to obtain a second feature map that pays more attention to the channel;

[0010] The first feature map is multiplied by the second feature map and then input into the spatial attention unit in the CBAM attention module to obtain the third feature map;

[0011] The first feature map is multiplied by the third feature map as the output of the CBAM attention module and input into the feature recovery module for feature recovery processing to obtain a fused image.

[0012] Preferably, the improved VGG module is an improvement on the original VGG module, and the improvement includes: deleting the maximum pooling layer in the original VGG module to keep the spatial scale of the input feature map unchanged; modifying the input of the first convolutional layer of the original VGG module, and the improved input is B=C+c , C is the number of channels of the input data, and c is the number of channels of the original multispectral.

[0013] Preferably, the convolution kernels in the preset VGG-based generator all use 3×3 convolution kernels.

[0014] Preferably, the preset VGG-based generator is obtained by training an initial generative adversarial network using multiple pairs of training samples and target loss functions, and the target loss function includes a generator loss function and a discriminator loss function.

[0015] Preferably, the step of inputting the original multispectral image and the original hyperspectral image into a preset VGG-based generator also includes:

[0016] Acquire multiple pairs of training samples, each pair of training samples includes an original multispectral image and an original hyperspectral image;

[0017] Construct an initial generative adversarial network;

[0018] The initial generative adversarial network includes: a VGG generator, a spatial discriminator D_spa and a spectral discriminator D_spe;

[0019] The VGG generator is used to generate a corresponding virtual low spatial resolution hyperspectral image according to the input original multispectral image and the original hyperspectral image;

[0020] The spatial discriminator D_spa is used to distinguish the spatial information of the fused image from the spatial information of the multispectral image;

[0021] The spectral discriminator D_spe is used to distinguish the spectral information of the fused image from the spectral information of the hyperspectral image;

[0022] Using the multiple pairs of training samples and the target loss function, training the initial generative adversarial network to obtain a trained generative adversarial network;

[0023] The generator loss function is used to train the VGG generator;

[0024] The discriminator loss function is used to train the spatial discriminator D_spa and the spectral discriminator D_spe; the trained VGG generator in the trained generative adversarial network is used as the preset VGG-based generator.

[0025] Preferably, the spatial discriminator D_spa and the spectral discriminator D_spe have exactly the same structure except for the input channel of the first convolutional layer; the input channel of the spatial discriminator D_spa is set to the number of multispectral image channels c, and the input channel of the spectral discriminator D_spe is set to the number of hyperspectral image channels C.

[0026] Preferably, the generator loss function is a weighted sum of adversarial loss, spatial loss and spectral loss, expressed as follows:

[0027] Loss_G=α*L spa +β*L spe +L adv ;

[0028] Among them, Loss_G is the total loss of the generator; L spa is the space loss; L spe is the spectral loss; L adv is the adversarial loss; α and β are hyperparameters used to control the balance among adversarial loss, spatial loss and spectral loss.

[0029] Preferably, the spatial loss is defined by the sum of the MSE loss and the perceptual loss, and the MSE loss and the perceptual loss are expressed as follows:

[0030]

[0031] Where W and H represent the length and width of the image; C and c represent the number of channels of the generated image and the original MSI respectively; G x,y,z (HSI, MSI) is the image generated by the generator; MSI x,y,z (x, y) is the original multispectral image; L mse and L per The initial weights of are 10 and 0.6 respectively to ensure that the two losses have the same order of magnitude; in the formula, Φ x,y represents the pre-trained VGG network, M and N represent the length and width of the feature map after the VGG network respectively;

[0032] The spectral loss is specifically:

[0033]

[0034] Where W and H represent the length and width of the image; C and c represent the number of channels of the generated image and the original MSI respectively; G x,y(HSI, MSI) is the image generated by the generator; MSI x,y (x, y) is the original multispectral image; is the transpose of the original multispectral image; || 2 is the two-norm operation;

[0035] The adversarial loss is specifically:

[0036] L adv =(D spa (G(HSI,MSI))-1) 2 +(D spe (G(HSI,MSI))-1) 2 ;

[0037] Among them, Dspa is the spatial discriminator; Dpse is the spectral discriminator.

[0038] Preferably, the discriminator loss function is specifically:

[0039]

[0040] Among them, G(X) is the fake image generated by the generator, a and b are soft labels used to give the original image and the generated image a classification label, and the loss function architecture is the least squares loss in LSGAN, which can make the network more stable; when i = spa, j = ms, the loss function represents the spatial discriminator loss. For the spatial discriminator, once it cannot distinguish between G(X) and I j , the model can preserve the spatial information of the multispectral image; when i = spe, j = hs, the loss function represents the spectral discriminator loss. In this part, it is mandatory to have G(X) and I j With the same spectral data distribution, the spectral discriminator cannot distinguish between G(X) and I j , the goal is achieved; D spe (G(X)) represents the output of the generated hyperspectral image after the spectral discriminator, which serves as the input of both the spatial and spectral discriminator losses, thus ensuring that adversarial training without reference images can be successfully completed.

[0041] In the present invention, the proposed multispectral and hyperspectral image fusion method superimposes the original multispectral image and the original hyperspectral image on the spectral channel to obtain the input data HSI; the input data HSI is input into a preset VGG-based generator to generate a fused image. By extracting the input features after splicing in the channel dimension, the ability of the network to learn various details in the image is improved, and the training parameters and the amount of calculation are greatly reduced by using the improved VGG module. The network performance is improved by the feature recovery module, and the problem of gradient disappearance is alleviated. The use of deep learning technology can effectively improve the spatial and spectral information of the fused image; the use of VGG Block and CBAM Block for feature extraction improves the learning ability of the model; the introduction of a dual discriminator structure realizes spatial and spectral adversarial learning, and further optimizes the fusion effect; by adjusting the weight of the loss function, a balance between the losses is achieved, and the quality of the fused image is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A schematic diagram of the workflow structure of a multispectral and hyperspectral image fusion method proposed in the present invention;

[0043] Figure 2 A schematic diagram of the overall network structure of a generator for a multispectral and hyperspectral image fusion method proposed in the present invention;

[0044] Figure 3 A schematic diagram of the improved VGG module structure of a multispectral and hyperspectral image fusion method proposed in the present invention;

[0045] Figure 4 A schematic diagram of the CBAM attention module structure of a multispectral and hyperspectral image fusion method proposed in the present invention;

[0046] Figure 5 Schematic diagram of the discriminator network structure of the multispectral and hyperspectral image fusion method proposed in the present invention. DETAILED DESCRIPTION

[0047] Reference Figure 1-5 The present invention proposes a multispectral and hyperspectral image fusion method, comprising the following steps:

[0048] S0, obtain multiple pairs of training samples, each pair of training samples includes an original multispectral image, an input data HSIand a fused image; construct an initial generative adversarial network; use multiple pairs of training samples and target loss functions to train the initial generative adversarial network to obtain a trained generative adversarial network; the initial generative adversarial network includes: a VGG generator, a spatial discriminator D_spa and a spectral discriminator D_spe; the VGG generator is used to generate a corresponding false hyperspectral image according to the input original multispectral image and the input data HSI; the spatial discriminator D_spa is used to distinguish the spatial information of the fused image from the spatial information of the multispectral image; the spectral discriminator D_spe is used to distinguish the spectral information of the fused image from the spectral information of the multispectral image; the generator loss function is used to train the VGG generator; the discriminator loss function is used to train the spatial discriminator D_spa and the spectral discriminator D_spe; the trained VGG generator in the trained generative adversarial network is used as the preset VGG-based generator.

[0049] In this embodiment, the training data image block size is 64×64 original MSI and pre-fused HSI images, the number of data processed in each batch is 16, and the number of training steps is set to 17100 steps (a total of 300 epochs are set). The optimizer of both the generator and the discriminator selects the Adam optimizer, and the hyperparameters α and β are set to 13 and 5 respectively. The soft label a is set to a random number between [0,0.3], and the soft label b is set to a random number between [0.7,1]. The initial learning rate of the generator is set to 0.001, and the initial learning rate of the discriminator is set to 0.0001. Thereafter, the discriminator learning rate is reduced to 0.8 times the original every 50 epochs. The network proposed in this paper is trained on an NVIDIA RTX 4090 GPU using PyTorch 2.0.1 in the Python 3.9.18 environment.

[0050] In this embodiment, if Figure 5 As shown in the figure, the spatial discriminator D_spa and the spectral discriminator D_spe have the same structure except for the input channel of the first convolution layer; the input channel of the spatial discriminator D_spa is set to the number of multispectral image channels c, and the input channel of the spectral discriminator D_spe is set to the number of hyperspectral image channels C. In the figure, k represents the convolution kernel size, n represents the number of convolution kernels, s represents the step size, Conv+BN+ReLU represents the convolution layer with batch normalization and activation function, and MaxPooling represents the maximum pooling layer.

[0051] In this embodiment, the spatial discriminator and the spectral discriminator are represented by D_spa and D_spe. The input of D_spa is the original multispectral image, and the input of D_spe is the original hyperspectral data and the generated fused image. The network structure of the two discriminators can force the image generated by the generator to contain more spatial and spectral information. The VGG feature extraction module is also introduced in the discriminator, and the fully connected layer in the original GAN ​​discriminator is replaced by a 1×1 convolution kernel, which can greatly reduce the number of feature channels and, more importantly, the number of parameters.

[0052] In this embodiment, the preset VGG-based generator is obtained by training the initial generative adversarial network using multiple pairs of training samples and target loss functions, and the target loss function includes a generator loss function and a discriminator loss function.

[0053] In this embodiment, the generator loss function is a weighted sum of adversarial loss, spatial loss, and spectral loss, expressed as follows:

[0054] Loss_G=α*L spa +β*L spe +L adv ;

[0055] Among them, Loss_G is the total loss of the generator; L spa is the space loss; L spe is the spectral loss; L adv is the adversarial loss; α and β are hyperparameters used to control the balance among adversarial loss, spatial loss and spectral loss.

[0056] In this embodiment, the spatial loss between the fake hyperspectral image generated by the generator and the original multispectral image is l spa Defined by the sum of MSE loss and perceptual loss, the MSE loss and perceptual loss are expressed as follows:

[0057]

[0058] Where W and H represent the length and width of the image; C and c represent the number of channels of the generated image and the original MSI respectively; G x,y,z (HSI, MSI) is the image generated by the generator; MSI x,y,z (x, y) is the original multispectral image; L mse and L per The initial weights of are 10 and 0.6 respectively to ensure that the two losses have the same order of magnitude; in the formula, Φ x,y represents the pre-trained VGG network, M and N represent the length and width of the feature map after the VGG network respectively;

[0059] The spectral loss is specifically:

[0060]

[0061] Among them, it can be seen from the definition of the loss function that the characteristic of our discriminator is to force the generator to generate a fused image with more spatial and spectral information in both the original MSI and the original HSI. Therefore, the adversarial loss includes two parts: spectral adversarial loss and spatial adversarial loss. The purpose of spatial adversarial learning is to make D_spa unable to distinguish between real multi-spectra and fake hyperspectra in spatial details, and the purpose of spectral adversarial learning is to make D_spe unable to distinguish between real hyperspectra and generated hyperspectra in spectral resolution. The initial weight in the formula is 0.1. W and H represent the length and width of the image; C and c represent the number of channels of the generated image and the original MSI respectively; G x,y (HSI, MSI) is the image generated by the generator; MSI x,y (x, y) is the original multispectral image; is the transpose of the original multispectral image; || 2 is the two-norm operation.

[0062] The adversarial loss is:

[0063] L adv =(D spa (G(HSI,MSI))-1) 2 +(D spe (G(HSI,MSI))-1) 2 ;

[0064] Among them, Dspa is the spatial discriminator; Dpse is the spectral discriminator.

[0065] In this embodiment, the discriminator loss function is specifically:

[0066]

[0067] Among them, G(X) is the fake image generated by the generator, a and b are soft labels used to give the original image and the generated image a classification label, and the loss function architecture is the least squares loss in LSGAN, which can make the network more stable; when i = spa, j = ms, the loss function represents the spatial discriminator loss. For the spatial discriminator, once it cannot distinguish between G(X) and I j , the model can preserve the spatial information of the multispectral image; when i = spe, j = hs, the loss function represents the spectral discriminator loss. In this part, it is mandatory to have G(X) and I j With the same spectral data distribution, the spectral discriminator cannot distinguish between G(X) and I j , the goal is achieved; D spe(G(X)) represents the output of the generated hyperspectral image after the spectral discriminator, which serves as the input of both the spatial and spectral discriminator losses, thus ensuring that adversarial training without reference images can be successfully completed.

[0068] Specifically, the discriminator part of the generative adversarial network training process requires two inputs, one is the generated image and the other is the input data. This will produce two probability values, which are input into the discriminator loss function, and further iterative training can be used to optimize the generator network parameters so that the distribution of the generated image is closer to the input data. However, the above training process is based on the assumption that there is a hyperspectral image with high spatial resolution and high spectral resolution in the original dataset as a reference image. However, there is no such an ideal hyperspectral image in the reference-free dataset used in this paper, so there is a problem that it cannot be learned in the spatial discriminator. To solve this problem, this paper proposes a dual discriminator adversarial learning method without reference images. First, in the spectral discriminator, consistent with the original training method, the discriminator input is the generated hyperspectrum and the original hyperspectrum, while in the spatial discriminator, the input is only the original high spatial resolution multispectral image. Secondly, the input of the generator is passed through the output of the spectral discriminator and is used as the input of the loss function of the spatial discriminator and the spectral discriminator at the same time to ensure the smooth completion of the adversarial training of the spatial discriminator.

[0069] S1. Superimpose the original multispectral image and the original hyperspectral image on the spectral channel to obtain the input data HSI .

[0070] S2. Input data HSI Input into the preset VGG-based generator to generate the fused image.

[0071] In this embodiment, the preset VGG-based generator includes an improved VGG module, a CBAM attention module, and a feature recovery module; the original multispectral image and the original hyperspectral image are input into the preset VGG-based generator to generate a fused image, specifically including: inputting the original multispectral image and the original hyperspectral image into the improved VGG module to obtain a first feature map; inputting the first feature map into the CBAM attention module, and the first feature map is passed through the channel attention unit in the CBAM attention module to obtain a second feature map that pays more attention to the channel; multiplying the first feature map and the second feature map and inputting the resultant into the spatial attention unit in the CBAM attention module to obtain a third feature map; multiplying the first feature map and the third feature map as the output of the CBAM attention module and inputting the feature recovery module for feature recovery processing to obtain a fused image.

[0072] In this embodiment, the improved VGG module is an improvement on the original VGG module, and the improvements include: deleting the maximum pooling layer in the original VGG module to keep the spatial scale of the input feature map unchanged; modifying the input of the first convolutional layer of the original VGG module, and the improved input is B=C+c, where C is the number of channels of the input data, and c is the number of channels of the original multi-spectrum.

[0073] In this embodiment, the convolution kernels in the preset VGG-based generator all use 3×3 convolution kernels.

[0074] It should be noted that the first two convolutional layers of the improved VGG module use 64 3×3 convolutional kernels, the next two convolutional layers use 128 3×3 convolutional kernels, the next three convolutional layers use 256 3×3 convolutional kernels, and the last four convolutional layers use 512 3×3 convolutional kernels. Each convolutional layer is activated by the ReLU activation function. The two convolutional layers of the channel attention layer in the CBAM attention module use 64 and 512 1×1 convolutional kernels respectively, replacing the fully connected layer to perform channel feature dimension up- and dimension down-scaling. The convolutional layers in the first residual block in the feature recovery module all use 512 3×3 convolutional kernels, the next convolutional layer and the second residual block all use 256 3×3 convolutional kernels, and the convolutional kernels of the last two convolutional layers are 3×3×128 and 3×3×C respectively, where C represents the number of channels of the hyperspectral data.

[0075] In the specific working process of the multispectral and hyperspectral image fusion method of this embodiment, the ability of the network to learn various details in the image is improved by extracting the input features after splicing in the channel dimension, and the training parameters and calculation amount are greatly reduced by using the improved VGG module. The network performance is improved by the feature recovery module to alleviate the problem of gradient disappearance. The use of deep learning technology can effectively improve the spatial and spectral information of the fused image; the use of VGG Block and CBAM Block for feature extraction improves the learning ability of the model; the introduction of a dual discriminator structure realizes spatial and spectral adversarial learning, further optimizing the fusion effect; by adjusting the weight of the loss function, a balance between the losses is achieved, and the quality of the fused image is improved.

[0076] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A multispectral and hyperspectral image fusion method, characterized in that: include: The original multispectral image is superimposed with the original hyperspectral image on the spectral channel to obtain the input data HSI; Input the input data HSI into the preset VGG-based generator to generate a fused image.

2. The multispectral and hyperspectral image fusion method according to claim 1, characterized in that: The preset VGG-based generator includes an improved VGG module, a CBAM attention module, and a feature recovery module; The step of inputting the original multispectral image and the original hyperspectral image into a preset VGG-based generator to generate a fused image specifically includes: Input the original multispectral image and the original hyperspectral image into the improved VGG module to obtain a first feature map; The first feature map is input into the CBAM attention module, and the first feature map is passed through the channel attention unit in the CBAM attention module to obtain a second feature map that pays more attention to the channel; The first feature map is multiplied by the second feature map and then input into the spatial attention unit in the CBAM attention module to obtain the third feature map; The first feature map is multiplied by the third feature map as the output of the CBAM attention module and then input into the feature recovery module for feature recovery processing to obtain a fused image.

3. The multispectral and hyperspectral image fusion method according to claim 2, characterized in that: The improved VGG module is an improvement on the original VGG module, and the improvements include: deleting the maximum pooling layer in the original VGG module to keep the spatial scale of the input feature map unchanged; modifying the input of the first convolutional layer of the original VGG module, and the improved input is B=C+c, where C is the number of channels of the input data and c is the number of channels of the original multi-spectrum.

4. The multispectral and hyperspectral image fusion method according to claim 1, characterized in that: The convolution kernels in the preset VGG-based generator all use 3×3 convolution kernels.

5. The multispectral and hyperspectral image fusion method according to claim 1, characterized in that: The preset VGG-based generator is obtained by training the initial generative adversarial network using multiple pairs of training samples and target loss functions, and the target loss function includes a generator loss function and a discriminator loss function.

6. The multispectral and hyperspectral image fusion method according to claim 5, characterized in that: The method of inputting the original multispectral image and the original hyperspectral image into a preset VGG-based generator also includes: Get multiple pairs of training samples, each pair of training samples includes an original multispectral image and an original hyperspectral image HSI ; Construct an initial generative adversarial network; The initial generative adversarial network includes: a VGG generator, a spatial discriminator D_spa and a spectral discriminator D_spe; The VGG generator is used to generate a corresponding false high spatial resolution hyperspectral image according to the input original multispectral image and the original hyperspectral image; The spatial discriminator D_spa is used to distinguish the spatial information of the fused image from the spatial information of the multispectral image; The spectral discriminator D_spe is used to distinguish the spectral information of the fused image from the spectral information of the hyperspectral image; Using the multiple pairs of training samples and the target loss function, training the initial generative adversarial network to obtain a trained generative adversarial network; The generator loss function is used to train the VGG generator; The discriminator loss function is used to train the spatial discriminator D_spa and the spectral discriminator D_spe; the trained VGG-based generator in the trained generative adversarial network is used as the preset VGG-based generator.

7. The multispectral and hyperspectral image fusion method according to claim 6, characterized in that: The spatial discriminator D_spa and the spectral discriminator D_spe have the same structure except for the input channel of the first convolutional layer; the input channel of the spatial discriminator D_spa is set to the number of multispectral image channels c, and the input channel of the spectral discriminator D_spe is set to the number of hyperspectral image channels C.

8. The multispectral and hyperspectral image fusion method according to claim 5, characterized in that: The generator loss function is a weighted sum of adversarial loss, spatial loss, and spectral loss, expressed as follows: Loss_G=α*L spa +β*L spe +L adv ; Among them, Loss_G is the total loss of the generator; L spa is the space loss; L spe is the spectral loss; L adv is the adversarial loss; α and β are hyperparameters used to control the balance among adversarial loss, spatial loss and spectral loss.

9. The multispectral and hyperspectral image fusion method according to claim 8, characterized in that: The spatial loss is defined by the sum of the MSE loss and the perceptual loss, which are expressed as follows: Where W and H represent the length and width of the image; C and c represent the number of channels of the generated image and the original MSI respectively; G x,y,z (HSI, MSI) is the image generated by the generator; MSI x,y,z (x, y) is the original multispectral image; L mse and L per The initial weights of are 10 and 0.6 respectively to ensure that the two losses have the same order of magnitude; in the formula, Φ x,y represents the pre-trained VGG network, M and N represent the length and width of the feature map after the VGG network respectively; The spectral loss is specifically: Where W and H represent the length and width of the image; C and c represent the number of channels of the generated image and the original MSI respectively; G x,y (HSI, MSI) is the image generated by the generator; MSI x,y (x, y) is the original multispectral image; is the transpose of the original multispectral image; ||2 is the two-norm operation; The adversarial loss is specifically: L adv =(D spa (G(HSI,MSI))-1) 2 +(D spe (G(HSI,MSI))-1) 2 ; Among them, Dspa is the spatial discriminator; Dpse is the spectral discriminator.

10. The multispectral and hyperspectral image fusion method according to claim 5, characterized in that: The discriminator loss function is specifically: Among them, G(X) is the fake image generated by the generator, a and b are soft labels used to give the original image and the generated image a classification label, and the loss function architecture is the least squares loss in LSGAN, which can make the network more stable; when i = spa, j = ms, the loss function represents the spatial discriminator loss. For the spatial discriminator, once it cannot distinguish between G(X) and I j , the model can preserve the spatial information of the multispectral image; when i = spe, j = hs, the loss function represents the spectral discriminator loss. In this part, it is mandatory to have G(X) and I j With the same spectral data distribution, the spectral discriminator cannot distinguish between G(X) and I j , the goal is achieved; D spe (G(X)) represents the output of the generated hyperspectral image after the spectral discriminator, which serves as the input of both the spatial and spectral discriminator losses, thus ensuring that adversarial training without reference images can be successfully completed.