Double-domain multispectral image enhancement method and system based on deep learning

Through a deep learning dual-domain multispectral image enhancement method, combined with variational inference and multi-scale spatial attention modules, the spatial and spectral enhancement problems of multispectral images in real scenes are solved, achieving high-quality image enhancement effects, which is suitable for crop detection and remote sensing image interpretation.

CN120689225APending Publication Date: 2025-09-23WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510658531.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing spatial and spectral enhancement methods for multispectral images lack generalization capabilities in real scenarios and find it difficult to simultaneously retain the color information of the visible light band and the spatial details of the near-infrared band. Traditional methods also consume large amounts of computational resources, and deep learning-based methods lack effective joint enhancement strategies.

Method used

A dual-domain multispectral image enhancement method based on deep learning is adopted. The blind pansharpening architecture of variational inference and the detail enhancement convolution module DEConvBlock are used for spatial enhancement. A dual encoder-single decoder network architecture and a multi-scale spatial attention module MSSA are designed for spectral enhancement. A three-stage training strategy is combined for collaborative training and tuning.

Benefits of technology

High-quality multispectral image enhancement is achieved in real scenarios, which improves spatial resolution and spectral effects. It has strong generalization and robustness and is suitable for crop detection and remote sensing image interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689225A_ABST
    Figure CN120689225A_ABST
Patent Text Reader

Abstract

The invention provides a double-domain multispectral image enhancement method and system based on deep learning. The method comprises a spatial enhancement sub-network and a spectrum enhancement sub-network, and comprises the following steps of: firstly, aiming at spatial enhancement, designing a blind degradation strategy by utilizing remote sensing satellite data to obtain simulation data as input of the spatial enhancement sub-network, and modeling a spatial enhancement process by utilizing a variational inference strategy; and then constructing a spatial enhancement sub-network SpatialNet containing degradation estimation by using a variational inference result, and carrying out accurate feature extraction by using a detail enhancement convolution module DEConvBlock. Aiming at spectrum enhancement, designing a double encoder-single decoder spectrum enhancement sub-network SpectraNet for multi-source feature extraction and fusion, and designing a multi-scale space attention module MSSA for extracting multi-scale and high-frequency details of a multi-source image; and finally, designing an unsupervised loss function, designing a three-stage training strategy, respectively training a space enhancement sub-network and a spectrum enhancement sub-network on a remote sensing satellite GaoFen-2 data set, and finally carrying out double-domain enhancement sub-network joint training optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and deep learning technology, and relates to a deep neural network method and system based on dual-domain multispectral image enhancement, which is suitable for multispectral image analysis scenarios with high spatial resolution and spectral information. Background Art

[0002] Multispectral imagery, due to its combination of visible and near-infrared wavelengths, can produce images of objects with rich spectral information, playing a crucial role in agriculture, ecological environment, and other fields. Thanks to the visible light band's ability to characterize the color and spatial texture of objects, and the near-infrared band's high cloud transmittance and sensitivity to vegetation, research on semantic segmentation, object detection, and crop detection based on multispectral imagery has been extensive. However, due to limitations of imaging sensors, it is difficult to directly obtain multispectral images with high spatial resolution. Furthermore, many applications of multispectral imagery utilize only the visible light band for visual analysis or feature extraction, while underutilizing the rich information in the near-infrared band. Therefore, it is of great significance to spatially enhance multispectral images using high-spatial-resolution panchromatic images and spectrally enhance the RGB bands using their inherent near-infrared wavelengths. The enhanced images retain the imaging advantages of both the visible and near-infrared bands while maintaining high spatial resolution, making them useful for visual interpretation of remote sensing images and other subsequent tasks.

[0003] The existing research on joint spatial and spectral enhancement of multispectral images has the following problems: (1) For spatial enhancement, pansharpening can only improve the spatial resolution of multispectral images. During visual analysis, the information of the infrared band is often ignored. In addition, the existing pansharpening methods often have poor generalization ability in real scenes, which easily leads to spectral distortion and cannot achieve a good balance between spatial details and spectrum. (2) In terms of spectral enhancement, since the near-infrared band has stronger noise than the visible light band, inappropriate fusion strategies can easily lead to distortion of spatial details in the fusion results. This is more obvious in remote sensing images with more complex ground object types. How to simultaneously retain the color information of the visible light band and the prominent information of the infrared band for certain targets remains to be studied. (3) How to properly design the coupling method for spatial and spectral enhancement algorithms is challenging. At the same time, research on multispectral images in the spatial and spectral dimensions is still a blank.

[0004] Multispectral image fusion enhancement methods based on panchromatic images (pansharpening) provide strong technical support for enhancing the spatial resolution of multispectral images. Pansharpening methods can be divided into traditional methods or deep learning-based methods. Traditional algorithms can be divided into component substitution (CS)-based methods, multi resolution analysis (MRA)-based methods, and variational optimization (VO)-based methods based on how they utilize panchromatic images. Both CS-based and MRA-based methods directly or indirectly replace the spatial components of the panchromatic image with the multispectral image, but neither method can achieve a balance between spatial detail and spectral fidelity. VO-based methods model the fusion process and construct a cost function. Although they can achieve good results, they consume computational resources.

[0005] With the advancement of deep learning, many researchers have designed supervised and unsupervised image fusion methods using architectures such as convolutional neural networks (CNNs), transformers, diffusion models, and generative adversarial networks (GANs). These methods are used to generate multispectral images with high spatial resolution and rich ground detail. For example, Hou et al. (Hou J, Cao Q, Ran R, et al. Bidomain modeling paradigm for pansharpening[C] / / Proceedings of the 31st ACM International Conference on Multimedia. 2023: 347-357.) designed BiMPan, which combines local and global aspects to achieve refined fusion results. Thanks to the global modeling capabilities of the Transformer, numerous Transformer-based methods have been proposed. For example, Ke et al. (Ke C, Liang H, Li D, et al. High-frequency transformer network based on window cross-attention for pansharpening[C] / / ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023: 1-5.) used high-pass filtering to extract high-frequency components from the source image and combined it with a Transformer to extract and fuse high-frequency features. Generative models primarily include those based on diffusion models and generative adversarial networks (GANs). Diffusion-based methods iteratively generate fused results from noisy images. GAN-based methods utilize adversarial training between generators and discriminators. These methods often produce results with clearer spatial details, more in line with human visual habits. The above methods all generate simulation data based on the Wald protocol according to fixed degradation, which makes it difficult to achieve generalization in real scenarios. Therefore, many blind pansharpening methods for unknown degradation (Zhang Z, Li H, Ke C, et al. DeepVariational Network for Blind Pansharpening[J]. IEEE Transactions on NeuralNetworks and Learning Systems, 2024.) have also been proposed.These methods use variational inference or kernel estimation methods to estimate the degradation of the source image and adjust the fusion results, achieving good results.

[0006] In the field of multispectral image enhancement, common methods for fusion of infrared and visible light bands can be divided into traditional methods and deep learning-based methods. Traditional methods primarily rely on multi-scale transformations such as wavelet transforms and pyramid transforms, or methods based on subspaces and sparse representations to achieve multimodal image decomposition and fusion. Deep learning-based methods can be divided into CNN-based methods and GAN-based methods. Because this task lacks a true value, CNN-based methods extract image features by designing a network architecture and establish constraints between the fusion result and the source image to guide network training. For example, Li et al. (Li H, Cen Y, Liu Y, et al. Different input resolutions and arbitrary output resolution: A meta-learning-based deep framework for infrared and visible image fusion[J]. IEEE Transactions on Image Processing, 2021, 30: 4070-4083.) proposed an infrared and visible light fusion method applicable to arbitrary resolutions using a meta-learning strategy combined with a CNN model. Ma et al. (Ma J, Yu W, Liang P, et al. FusionGAN: A generative adversarial network for infrared and visible image fusion[J]. Information fusion, 2019, 48: 11-26.) proposed FusionGAN, which applied the GAN architecture to this task. Subsequently, the GAN network DDcGAN (Ma J, Xu H, Jiang J, et al. DDcGAN: A dual-discriminator conditional generative adversarial network for multi-resolution image fusion[J]. IEEE Transactions on Image Processing, 2020, 29: 4980-4995.) based on multiple discriminators was proposed, hoping to achieve better fusion results. However, these methods focus solely on spatial or spectral enhancement of multispectral images, lacking research on joint spatial-spectral enhancement. Furthermore, while cascading these two algorithms can achieve joint enhancement, the differences between the tasks prevent direct coupling between the two algorithms, leaving the enhancement results to be improved. Summary of the Invention

[0007] In response to the shortcomings of the existing technology, the present invention proposes a dual-domain multispectral image enhancement method based on deep learning. The method includes two levels: spatial domain enhancement and spectral domain enhancement. In terms of spatial enhancement, the method uses a blind pansharpening architecture based on variational inference, and uses the detail enhancement convolution module DEConvBlock as the basic network block to improve the feature extraction capability. In terms of spectral enhancement, we adopt a dual encoder-single decoder network architecture and propose a multi-scale spatial attention module MSSA as the basic network block to improve the spectral effect in real scenes. This method innovatively proposes a three-stage training strategy and adopts a collaborative training and tuning approach to produce high-quality enhancement results.

[0008] The technical solution adopted by the present invention is: a dual-domain multispectral image enhancement method based on deep learning, which includes two sub-networks: spatial enhancement and spectral enhancement. First, for spatial enhancement, a blind degradation strategy is designed using remote sensing satellite data to obtain simulated data as the input of the spatial enhancement sub-network, and the spatial enhancement process is modeled using a variational inference strategy. The variational inference results are then used to construct a spatial enhancement sub-network SpatialNet containing degradation estimation, and the detail enhancement convolution module DEConvBlock is used for precise feature extraction. For spectral enhancement, a dual-encoder-single-decoder spectral enhancement sub-network SpectralNet is designed for multi-source feature extraction and fusion, and a multi-scale spatial attention module is designed to extract multi-scale and high-frequency details from multi-source images. Finally, an unsupervised loss function is designed, and a three-stage training strategy is designed to train the spatial enhancement and spectral enhancement sub-networks on the remote sensing satellite GaoFen-2 dataset, respectively, and finally the dual-domain enhancement sub-network is jointly trained and optimized. The method specifically includes the following steps: Step 1: Based on the blind degradation of remote sensing images and the variational inference strategy, the training data required for training the spatial enhancement sub-network is generated. The variational inference results are used to construct the spatial enhancement sub-network, which includes a degradation estimation network group and a fusion enhancement network. The degradation estimation network group is used to estimate the blur and noise degradation of the input source image and adjust the fusion enhancement result. The fusion enhancement network adopts a U-shaped network architecture to extract multi-scale details of the source image and ultimately generate a high-spatial-resolution multispectral image. Step 2: Construct a spectral enhancement subnetwork with a dual encoder-single decoder network architecture. The input multispectral image is split into visible light band and near-infrared band at the channel level as the input of the dual encoder. The dual encoder includes a visible light band encoder and a near-infrared band encoder. A multi-scale attention module is used as the basic module of the encoder to extract the multi-scale spatial features of the source image. The decoder adopts the basic structure of CNN and splices the output features of the dual encoder in the channel dimension as the decoder input to generate the spectral enhancement result of the multispectral image. Step three: construct a loss function and adopt corresponding training strategies to jointly train and optimize the spatial enhancer network and the spectral enhancer network.

[0009] Furthermore, in step 1, the input image of the spatial enhancement subnetwork SpatialNet is the panchromatic image P and the multispectral image M, and the spatial enhancement fusion result is ; Using blind degradation strategy for simulation data generation, for multispectral images:

[0010] in is the multispectral image obtained in the dataset, which is used as the true value for reference during training. is a randomly generated anisotropic Gaussian blur, is randomly generated Gaussian noise; Indicates that the image is downsampled. is the scale of spatial downsampling; For the full-color image, add randomly generated Gaussian noise to it:

[0011] in, represents the full-color image directly obtained from the dataset, Represents anisotropic Gaussian noise added to the full-color image.

[0012] Furthermore, the input full-color image P and multispectral image M are fused with the spatial enhancement result as follows: The correlation relationship between them is used to output the variational inference results. According to the simulation relationship of image degradation, the relationship between the two input images and the output image is constructed as follows:

[0013]

[0014] in, Represents the number of bands of the multispectral image, is the weighted weight of the i-th band of the multispectral image, Indicates the fusion result The i-th band of Indicates the gradient operation of the image. The noise parameter added to the full-color image.

[0015] Furthermore, the degradation estimation network group includes a sub-network K-Net for estimating blur degradation of multispectral images and a sub-network N-Net for estimating noise degradation of multispectral images. Sub-networks K-Net and N-Net take low-resolution multispectral images as input and output related estimation parameters. In addition, it also includes a sub-network S-Net for estimating noise degradation parameters of panchromatic images, which takes panchromatic images as input and outputs noise-related parameters. The above three degradation estimation networks all adopt the existing DnCNN network architecture.

[0016] Furthermore, the spatial fusion enhancement network adopts a U-shaped network architecture, using downsampling and upsampling operations to extract the spatial features of the input image at each resolution stage. By adding skip links between the encoder and decoder at the same scale, the encoder features are added to the input features of the decoding layer at the channel level. The full-color image P and the multispectral image M are spliced ​​in the channel dimension and input into the network. First, the shallow features are obtained through the convolution layer and input into the encoder. The detail enhancement convolution module DEConvBlock is used as the basic module of the encoder and decoder to extract fine-grained features. Assume that the spatial features of the input encoder layer i are , the output features after DEConvBlock processing are , then DEConvBlock is expressed as:

[0017] in, For detail enhancement convolution, Relu represents the activation function; Finally, the multi-scale features pass through the decoder to obtain a multispectral image with high spatial resolution.

[0018] Furthermore, the encoder for extracting image features in the visible light band and the encoder for the near-infrared band use the same network architecture but keep the parameters independent. The input image first passes through the convolution-activation layer to obtain shallow image features and is input into several layers of multi-scale spatial attention modules MSSA. The final deep features are spliced ​​in the channel dimension and then input into the decoder; the processing process of the multi-scale spatial attention module MSSA is: assuming that the features input to the multi-scale spatial attention module are , the output features are ; The features of the input MSSA are first convolved with detail enhancement and Convolution It is then input into an asymmetric convolution block consisting of row and column convolutions of sizes 3, 5, and 7 to extract multi-receptive field features:

[0019] in, and It is the intermediate output feature obtained during the processing of the input feature; The Sobel operator is used to introduce a high-frequency detail branch. The features output by the high-frequency detail branch and the multi-scale attention output features are added together and then convolution Get output features; .

[0020] Furthermore, the decoder is composed of multiple layers of cascaded convolutional layers. Convolution and LeakyRelu activation function, where the mathematical expression of LeakyRelu activation function is:

[0021] in, Represents the parameter that adjusts the slope of the negative semi-axis of the activation function, x is the independent variable; Finally, the features processed by multiple convolution layers are Convolution is used to perform feature integration to obtain spectral enhancement fusion results.

[0022] Furthermore, the loss function used by the spectral enhancement sub-network includes pixel-level loss and visual perception loss; among them, the pixel-level loss includes pixel loss And structural similarity loss calculation , The calculation formula is as follows:

[0023] in, represents the near-infrared band, VIS represents the visible light band, Represents the final output of the spectral enhancement sub-network SpectralNet, Represents the brightness component of the visible light band, Spectral enhancement results The brightness component; max() means taking the maximum value operation pixel by pixel, and Canny() means using the Canny operator to extract edge information; Structural similarity loss The calculation formula is as follows:

[0024] in, Refers to calculating the structural similarity of variables a and b and then taking the difference; Visual perception loss The calculation includes: using some convolutional layers of VGG19 to extract and near-infrared band And the semantic features of the visible light band VIS as well as And use the L1 norm to construct the loss:

[0025] Finally, the overall loss is:

[0026] in , are the weights of the three losses.

[0027] Furthermore, in step three, a three-stage training strategy is designed for the joint training and tuning of the spatial enhancer network and the spectral enhancer network. In the first stage, the multispectral image M and the panchromatic image P are used as input to generate a high-spatial-resolution multispectral image through the spatial enhancer network SpatialNet. In the second stage, the high-resolution multispectral image is split into visible light image and near-infrared image at the channel level as the input of the spectral enhancer network, and the output is a visible light image that takes into account multi-band information. The third stage is the coupled fine-tuning stage, in which the parameters of the trained spatial enhancement model are fixed, the output result is used as the input of the spectral enhancement model, and the spectral enhancement model is fine-tuned.

[0028] The present invention also provides a dual-domain multispectral image enhancement system based on deep learning, comprising the following modules: The spatial enhancement module is used to generate the training data required for training the spatial enhancement sub-network based on blind degradation of remote sensing images and variational inference strategies. The variational inference results are used to construct a spatial enhancement sub-network including a degradation estimation network group and a fusion enhancement network. The degradation estimation network group is used to estimate the blur and noise degradation of the input source image and adjust the fusion enhancement results. The fusion enhancement network adopts a U-shaped network architecture to extract multi-scale details of the source image and ultimately generate a high-spatial-resolution multispectral image. The spectral enhancement module is used to construct a spectral enhancement subnetwork of a dual encoder-single decoder network architecture. The input multispectral image is split into visible light band and near-infrared band at the channel level as the input of the dual encoder. The dual encoder includes a visible light band encoder and a near-infrared band encoder. The multi-scale attention module is used as the basic module of the encoder to extract the multi-scale spatial features of the source image. The decoder adopts the basic structure of CNN and splices the output features of the dual encoder in the channel dimension as the decoder input to generate the spectral enhancement result of the multispectral image. The training and tuning module is used to construct the loss function and adopt the corresponding training strategy to perform joint training and tuning of the spatial enhancer network and the spectral enhancer network.

[0029] Compared with the existing technology, the advantages and beneficial effects of the present invention are as follows: The present invention proposes a dual-domain multispectral image enhancement method based on deep learning. A blind pansharpening architecture based on variational inference is used, and the detail enhancement convolution module DEConvBlock is used as the basic block of the network to improve the feature extraction capability. In terms of spectral enhancement, a dual encoder-single decoder network architecture is adopted, and a multi-scale spatial attention module MSSA is proposed as the basic block of the network to improve the spectral effect in real scenes. An innovative three-stage training strategy is proposed, and a collaborative training and tuning method is adopted to produce high-quality enhancement results. It has strong generalization and robustness, and has important application potential in crop detection and remote sensing image interpretation. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 1 and 2 are input low-resolution multispectral and panchromatic images of the embodiment, where (a) is a panchromatic image and (b) is a multispectral image.

[0031] Figure 2 2 is a diagram of a spatial and spectral enhancement network architecture of an embodiment.

[0032] Figure 3 It is the constraint relationship diagram of the variational loss function of the spatial enhancement sub-network.

[0033] Figure 4 Comparison chart of the joint spatial-spectral enhancement results compared with various advanced methods. DETAILED DESCRIPTION

[0034] In order to facilitate those skilled in the art to understand and implement the present invention, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0035] In response to the shortcomings of the existing technology, the present invention proposes a dual-domain multispectral image enhancement method based on deep learning. The method includes two levels: spatial domain enhancement and spectral domain enhancement. In terms of spatial enhancement, the method uses a blind pansharpening architecture based on variational inference, and uses the detail enhancement convolution module DEConvBlock as the basic network block to improve the feature extraction capability. In terms of spectral enhancement, we adopt a dual encoder-single decoder network architecture and propose a multi-scale spatial attention module MSSA as the basic network block to improve the spectral effect in real scenes. This method innovatively proposes a three-stage training strategy and adopts a collaborative training and tuning approach to produce high-quality enhancement results.

[0036] Figure 1 is the input of the dual-domain enhancement network for multispectral images. Figure 1 (a) and Figure 1 (b) in the figure shows the full-color image and multispectral image after simulation of the blind degradation strategy.

[0037] like Figure 2 As shown, an embodiment of the present invention provides a dual-domain multispectral image enhancement method based on deep learning, comprising the following steps: Step 1: Design a remote sensing image blind degradation and variational inference strategy to produce the training data required for training the spatial enhancement sub-network, and use the variational inference results to construct the spatial enhancement sub-network SpatialNet.

[0038] The spatial enhancement subnetwork, SpatialNet, includes a degradation estimation network group and a fusion enhancement network, R-Net. The degradation estimation network group is used to estimate the blur and noise degradation of the input source image and adjust the fusion enhancement result. The fusion enhancement network, R-Net, adopts a U-shaped network architecture and uses a detail enhancement convolution module, DEConvBlock, to extract fine-grained features of the source image. The fusion enhancement network R-Net adopts an encoder-decoder network architecture with downsampling and upsampling operations to extract multi-scale details of the source image; The detail enhancement convolution module DEConvBlock is the basic module of the encoder and decoder, which uses detail enhancement convolution DEConv and activation function and adopts residual link and other operations to achieve fine-grained extraction of image features; Step 2: Construct a sub-network SpectralNet for spectral enhancement of multispectral images. The sub-network for spectral enhancement of multispectral images is a dual encoder-single decoder network architecture. The input multispectral image is split into visible light band and near-infrared band at the channel level as network input. The dual encoder includes a visible light band encoder and a near infrared band encoder, and uses the multi-scale attention module MSSA proposed in the present invention as the basic module of the encoder to extract the multi-scale spatial features of the source image; The multi-scale attention module includes a multi-scale convolution branch and a high-frequency detail extraction branch, which are used to extract multi-scale spatial details and high-frequency features of the image; The decoder uses the basic convolutional layer and activation layer of CNN as the basic layer, and splices the output features of the dual encoders in the channel dimension as the decoder input to generate the spectral enhancement result of the multispectral image; Step three: design an unsupervised loss function for training the spectral enhancer network, use the variational inference results to constrain the training of the spatial enhancer network, and design a three-stage training strategy for joint training and tuning of the spatial enhancer network and the spectral enhancer network.

[0039] The specific implementation of step one is as follows: The input images of the spatial enhancement subnetwork SpatialNet are the panchromatic image P and the multispectral image M, and the spatial enhancement fusion result is Different from the classic pansharpening method that uses the Wald protocol to add fixed blur to multispectral images to construct training datasets, we use a blind degradation strategy for simulation data generation. For multispectral images:

[0040] in is the multispectral image obtained in the dataset, which is used as the true value for reference during training. is a randomly generated anisotropic Gaussian blur, is randomly generated Gaussian noise. Indicates that the image is downsampled. is the spatial downsampling scale, and in this embodiment the downsampling scale is 4.

[0041] For the full-color image, we add randomly generated Gaussian noise to it:

[0042] in, represents the full-color image directly obtained from the dataset, represents the anisotropic Gaussian noise added to the full-color image. In this embodiment, the Gaussian noise range is [0, 30].

[0043] It is further necessary to construct the input full-color image P and multispectral image M and the spatial enhancement fusion result as follows: The relationship between the two input images and the output image is constructed as follows:

[0044]

[0045] in, represents the number of bands of the multispectral image. In this embodiment, the number of bands is 4. is the weighted weight of the i-th band of the multispectral image, Indicates the fusion result The i-th band of Indicates the gradient operation on the image. is a noise parameter added to the full-color image. In this embodiment, the Gaussian noise range is [0, 30].

[0046] The above degradation process is then modeled using variational inference related rules, where 、 as well as Related parameters and fusion enhancement results is the result that needs to be estimated for variational inference. The final variational inference estimation result is in the form of:

[0047]

[0048] Where E( ) represents the expectation of the probability distribution, q( ) is the posterior probability distribution to be estimated, and p( ) represents the prior probability of each variable. Represents KL divergence, which is used to measure the distance between two probability distributions. is the blur of the multispectral image M Related parameters, is the noise of the multispectral image M Related parameters, is the noise of the full-color image Related parameters. Represents the lower bound of evidence, and when it is maximized, the optimal spatial enhancement result is achieved.

[0049] Furthermore, the spatial enhancement sub-network SpatialNet includes a degradation estimation network group and a spatial fusion enhancement network R-Net, and the network structure is as follows: Figure 2As shown in (a) in the figure. The degradation estimation network group includes the network K-Net for multispectral image blur degradation estimation and the sub-network N-Net for multispectral image noise degradation estimation. The two sub-networks take low-resolution multispectral images as input, and the blur degradation estimation network K-Net outputs blur kernel related parameters m, , , the noise degradation estimation network N-Net outputs the noise distribution related parameters of the multispectral image In addition, it also includes a sub-network S-Net that estimates the noise degradation parameters of the full-color image. It takes the full-color image as input and outputs the noise distribution related parameters a and b of the full-color image. The above three degradation estimation networks all use the existing DnCNN network architecture. The relevant degradation parameters and spatial enhancement results are used together to construct the variational inference loss function, and the degradation estimation network is constrained to be jointly trained with the spatial enhancement fusion network. The mutual constraints are as follows: Figure 3 As shown, , , , Represent the weights of K-Net, S-Net, N-Net, and R-Net networks respectively. It can be seen that the results of each sub-network estimation are transmitted to the corresponding The term calculates the loss of each sub-network and performs backpropagation to adjust the weights of each sub-network. At the same time, the expectation term (E( )) in the figure is used to constrain the collaborative training of the four networks to achieve synchronous optimization.

[0050] Furthermore, the fusion enhancement network R-Net adopts a U-shaped network architecture, using downsampling and upsampling operations to extract spatial features of the input image at various resolution levels. In this embodiment, the number of downsampling operations is two, meaning that the network has three layers. By adding skip links between the encoder and decoder at the same scale, the encoder features are added to the input features of the decoding layer at the channel level to fully utilize the features. The input M and P are spliced ​​in the channel dimension and then input to the network.

[0051] The image input to the network first passes through the convolution layer to obtain shallow features and then inputs them into the encoder. In view of the low resolution of remote sensing multispectral images and the small scale of objects, we use the detail enhancement convolution module DEConvBlock as the basic module of the encoder and decoder to extract fine-grained features. The structure of DEConvBlock is as follows: Figure 2 As shown in the left image of (b). Assume that the spatial feature of the input encoder layer i is , the output features after DEConvBlock processing are , then DEConvBlock can be expressed as:

[0052] in, Convolution is used to enhance details. Relu represents the activation function, and its mathematical form is:

[0053] Thanks to the differential convolution design in DEConv, the convolution kernel has different weights at different positions, which can better preserve spatial detail information. Finally, the multi-scale features pass through the decoder to obtain a high-spatial-resolution multispectral image. In step 2, for the spectral enhancement sub-network SpectralNet used for spectral enhancement of multispectral images, the visible light band and near infrared band of the multispectral image are used as input, and a dual encoder-single decoder network architecture is adopted. The network architecture of the spectral enhancement sub-network SpectralNet for multispectral images is as follows: Figure 2 As shown in (a) in .

[0054] In terms of encoder design, the encoder for image feature extraction in the visible light band and the encoder for the near-infrared band use the same network architecture but keep the parameters independent. Furthermore, in order to extract deep details of the image in different receptive fields, we designed a multi-scale spatial attention module MSSA, whose network structure is as follows: Figure 2 As shown in (b) in the figure. In this embodiment, a 2-layer MSSA cascade is used. Assume that the feature input to the multi-scale spatial attention module is , the output features are The features input to MSSA are first passed through as well as Convolution It is then input into an asymmetric convolution block consisting of row and column convolutions of sizes 3, 5, and 7 to extract multi-receptive field features:

[0055] in, and is the intermediate output feature obtained during the processing of the input feature, The convolution kernel size is The other convolutional layers use the same naming convention.

[0056] In addition, in order to better preserve high-frequency information, we use the Sobel operator to introduce a high-frequency detail branch. The features output by the high-frequency detail branch and the multi-scale attention output features are added together and then convolution Get the output features.

[0057]

[0058] The input image first passes through the convolution-activation layer to obtain shallow image features and is input into the two-layer MSSA. The final deep features are spliced ​​in the channel dimension and input into the decoder.

[0059] Furthermore, the decoder is composed of three cascaded convolutional layers. Convolution and LeakyRelu activation function. The mathematical expression of LeakyRelu activation function is:

[0060] in, Represents the parameter that adjusts the slope of the negative semi-axis of the activation function, x is the independent variable; Finally, the features processed by multiple convolution layers are Convolution is used to perform feature integration to obtain spectral enhancement fusion results; In step 3, we designed an unsupervised loss function for training the spectral image enhancement network. To ensure that the fusion process preserves the color information of the visible light band and the spatial details of the near-infrared band, while also ensuring that the fusion result conforms to the observer's visual habits, we designed a spectral loss function that includes both pixel-level loss and visual perception loss. Among them, pixel-level loss First, in order to ensure that the fusion result contains the spatial information of both visible light images and near-infrared images, and at the same time ensure the high-frequency details of the fusion result, we designed a loss function :

[0061] in, represents the near-infrared band, VIS represents the visible light band, Represents the final output of the spectral enhancement sub-network SpectralNet, Represents the brightness component of the visible light band, Spectral enhancement results The brightness component of the image. max() represents the pixel-by-pixel maximum operation, and Canny() represents the use of the Canny operator to extract edge information.

[0062] In addition, we also introduce a pixel-level structural similarity loss:

[0063] in, Refers to calculating the structural similarity of variables a and b and then taking the difference.

[0064] In order to ensure that the fusion result is more in line with visual perception habits, we also introduced visual perception loss Specifically, we use some convolutional layers of VGG19 to extract and near-infrared band And the semantic features of the visible light band VIS as well as And use the L1 norm to construct the loss:

[0065] Finally, the overall loss is:

[0066] in , are the weights of the three losses. In this embodiment , =0.1.

[0067] Furthermore, a three-stage training strategy was designed for joint training and tuning of the spatial and spectral enhancement sub-networks. This independent training model for spatial and spectral enhancement ensures controllable enhancement effects at each stage. Furthermore, a two-stage coupled tuning training is implemented as the third stage to enhance the effectiveness of the two models. In the first stage, the spatial enhancement sub-network, SpatialNet, uses M and P as inputs to generate high-spatial-resolution multispectral images. In the second stage, the high-resolution multispectral image is split channel-wise into visible and near-infrared images, which serve as inputs to the spectral enhancement network. The output is a visible image that incorporates multi-band information. In the third stage, the coupled fine-tuning stage, the parameters of the spatial enhancement model are fixed, and the output serves as the input to the spectral enhancement model. The spectral enhancement model is then fine-tuned to improve the proposed algorithm's dual-domain enhancement performance and adaptability to multiple resolutions. The network is then simulated and degraded on 5,300 pairs of images from the GaoFen-2 dataset for network training.

[0068] Based on the results of joint spatial and spectral enhancement of multispectral images obtained in the above steps, in order to compare with other methods, we selected the advanced algorithms ADKNet, HyperDSNet, MSDDN in the field of spatial enhancement and the spectral enhancement algorithm LRRNet and the traditional algorithm YCbCr method for comparison. The results are shown in the attached figure. Figure 4 shown.

[0069] To quantitatively evaluate the combined spatial and spectral enhancement performance of various methods, we introduced structural similarity (SSIM) and information entropy (EN) as evaluation metrics. Due to the unsupervised nature of this task, structural similarity was calculated by averaging the fused enhancement results with the visible and near-infrared bands. For both metrics, higher values ​​indicate better performance. The quantitative evaluation results for these metrics on the GaoFen-2 dataset are as follows:

[0070] Quantitative results show that the algorithm proposed in this paper is superior to other comparative methods in terms of spatial and spectral enhancement effects, and can generate high-quality spatial and spectral enhancement results of multispectral images.

[0071] On the other hand, an embodiment of the present invention further provides a dual-domain multispectral image enhancement system based on deep learning, comprising the following modules: The present invention also provides a dual-domain multispectral image enhancement system based on deep learning, comprising the following modules: The spatial enhancement module is used to generate the training data required for training the spatial enhancement sub-network based on blind degradation of remote sensing images and variational inference strategies. The variational inference results are used to construct a spatial enhancement sub-network including a degradation estimation network group and a fusion enhancement network. The degradation estimation network group is used to estimate the blur and noise degradation of the input source image and adjust the fusion enhancement results. The fusion enhancement network adopts a U-shaped network architecture to extract multi-scale details of the source image and ultimately generate a high-spatial-resolution multispectral image. The spectral enhancement module is used to construct a spectral enhancement subnetwork of a dual encoder-single decoder network architecture. The input multispectral image is split into visible light band and near-infrared band at the channel level as the input of the dual encoder. The dual encoder includes a visible light band encoder and a near-infrared band encoder. The multi-scale attention module is used as the basic module of the encoder to extract the multi-scale spatial features of the source image. The decoder adopts the basic structure of CNN and splices the output features of the dual encoder in the channel dimension as the decoder input to generate the spectral enhancement result of the multispectral image. The training and tuning module is used to construct the loss function and adopt the corresponding training strategy to perform joint training and tuning of the spatial enhancer network and the spectral enhancer network.

[0072] The specific implementation method of each module is the same as that of each step and will not be described in detail in the present invention.

[0073] It should be understood that parts not elaborated in detail in this specification belong to the prior art.

[0074] It should be understood that the above description of the embodiments is relatively detailed and cannot be regarded as limiting the scope of protection of the patent of the present invention. Under the guidance of the present invention, ordinary technicians in this field can also make substitutions or modifications without departing from the scope of protection of the claims of the present invention, which all fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be based on the attached claims.

Claims

1. A dual-domain multispectral image enhancement method based on deep learning, characterized in that: The steps include: Step 1: Based on the blind degradation of remote sensing images and the variational inference strategy, the training data required for training the spatial enhancement sub-network is generated. The variational inference results are used to construct the spatial enhancement sub-network, which includes a degradation estimation network group and a fusion enhancement network. The degradation estimation network group is used to estimate the blur and noise degradation of the input source image and adjust the fusion enhancement result. The fusion enhancement network adopts a U-shaped network architecture to extract multi-scale details of the source image and ultimately generate a high-spatial-resolution multispectral image. Step 2: Construct a spectral enhancement subnetwork with a dual encoder-single decoder network architecture. The input multispectral image is split into visible light band and near-infrared band at the channel level as the input of the dual encoder. The dual encoder includes a visible light band encoder and a near-infrared band encoder. A multi-scale attention module is used as the basic module of the encoder to extract the multi-scale spatial features of the source image. The decoder adopts the basic structure of CNN and splices the output features of the dual encoder in the channel dimension as the decoder input to generate the spectral enhancement result of the multispectral image. Step three: construct a loss function and adopt corresponding training strategies to jointly train and optimize the spatial enhancer network and the spectral enhancer network.

2. The dual-domain multispectral image enhancement method based on deep learning according to claim 1, characterized in that: In step 1, the input image of the spatial enhancement sub-network SpatialNet is the panchromatic image P and the multispectral image M, and the spatial enhancement fusion result is ; Using blind degradation strategy for simulation data generation, for multispectral images: in is the multispectral image obtained in the dataset, which is used as the true value for reference during training. is a randomly generated anisotropic Gaussian blur, is randomly generated Gaussian noise; Indicates that the image is downsampled. is the scale of spatial downsampling; For the full-color image, add randomly generated Gaussian noise to it: in, represents the full-color image directly obtained from the dataset, Represents anisotropic Gaussian noise added to the full-color image.

3. The dual-domain multispectral image enhancement method based on deep learning according to claim 2, characterized in that: Construct the input full-color image P and multispectral image M and the spatial enhancement fusion result as follows: The correlation relationship between them is used to output the variational inference results. According to the simulation relationship of image degradation, the relationship between the two input images and the output image is constructed as follows: in, Represents the number of bands of the multispectral image, is the weighted weight of the i-th band of the multispectral image, Indicates the fusion result The i-th band of Indicates the gradient operation of the image. The noise parameter added to the full-color image.

4. The dual-domain multispectral image enhancement method based on deep learning according to claim 1, characterized in that: The degradation estimation network group includes the sub-network K-Net for estimating the blur degradation of multispectral images and the sub-network N-Net for estimating the noise degradation of multispectral images. The sub-networks K-Net and N-Net take low-resolution multispectral images as input and output related estimation parameters. In addition, it also includes the sub-network S-Net for estimating the noise degradation parameters of panchromatic images, which takes panchromatic images as input and outputs noise-related parameters. The above three degradation estimation networks all adopt the existing DnCNN network architecture.

5. The dual-domain multispectral image enhancement method based on deep learning according to claim 1, characterized in that: The spatial fusion enhancement network adopts a U-shaped network architecture, using downsampling and upsampling operations to extract the spatial features of the input image at each resolution stage. By adding skip links between the encoder and decoder at the same scale, the encoder features are added to the input features of the decoding layer at the channel level. The full-color image P and the multispectral image M are spliced ​​in the channel dimension and input into the network. First, the shallow features are obtained through the convolution layer and input into the encoder. The detail enhancement convolution module DEConvBlock is used as the basic module of the encoder and decoder to extract fine-grained features. Assume that the spatial features of the input encoder layer i are , the output features after DEConvBlock processing are , then DEConvBlock is expressed as: in, For detail enhancement convolution, Relu represents the activation function; Finally, the multi-scale features pass through the decoder to obtain a multispectral image with high spatial resolution.

6. The dual-domain multispectral image enhancement method based on deep learning according to claim 1, characterized in that: The encoder for extracting image features in the visible light band and the encoder for the near-infrared band use the same network architecture but keep the parameters independent. The input image first passes through the convolution-activation layer to obtain shallow image features and is input into several layers of multi-scale spatial attention modules MSSA. The final deep features are spliced ​​in the channel dimension and input into the decoder; the processing process of the multi-scale spatial attention module MSSA is: assuming that the features input to the multi-scale spatial attention module are , the output features are ; The input features of MSSA are first convolved with detail enhancement and Convolution It is then input into an asymmetric convolution block consisting of row and column convolutions of sizes 3, 5, and 7 to extract multi-receptive field features: in, and It is the intermediate output feature obtained during the processing of the input feature; The Sobel operator is used to introduce a high-frequency detail branch. The features output by the high-frequency detail branch and the multi-scale attention output features are added together and then convolution Get output features; 。 7. The dual-domain multispectral image enhancement method based on deep learning according to claim 1, characterized in that: The decoder consists of multiple cascaded convolutional layers. Convolution and LeakyRelu activation function, where the mathematical expression of LeakyRelu activation function is: in, Represents the parameter that adjusts the slope of the negative semi-axis of the activation function, x is the independent variable; Finally, the features processed by multiple convolution layers are Convolution is used to perform feature integration to obtain spectral enhancement fusion results.

8. The dual-domain multispectral image enhancement method based on deep learning according to claim 1, characterized in that: The loss function used by the spectral enhancement sub-network includes pixel-level loss and visual perception loss; among them, pixel-level loss includes pixel loss And structural similarity loss calculation , The calculation formula is as follows: in, represents the near-infrared band, VIS represents the visible light band, Represents the final output of the spectral enhancement sub-network SpectralNet, Represents the brightness component of the visible light band, Spectral enhancement results The brightness component; max() means taking the maximum value operation pixel by pixel, and Canny() means using the Canny operator to extract edge information; Structural similarity loss The calculation formula is as follows: in, Refers to calculating the structural similarity of variables a and b and then taking the difference; Visual perception loss The calculation includes: using some convolutional layers of VGG19 to extract and near-infrared band And the semantic features of the visible light band VIS as well as And use the L1 norm to construct the loss: Finally, the overall loss is: in , are the weights of the three losses.

9. The dual-domain multispectral image enhancement method based on deep learning according to claim 1, characterized in that: In step 3, a three-stage training strategy is designed for joint training and tuning of the spatial enhancer network and the spectral enhancer network. In the first stage, a multispectral image M and a panchromatic image P are used as input to generate a high-spatial-resolution multispectral image through the spatial enhancement sub-network SpatialNet. In the second stage, the high-resolution multispectral image is split into visible light image and near-infrared image at the channel level as input to the spectral enhancement sub-network, which outputs a visible light image that takes into account multi-band information. The third stage is the coupling fine-tuning stage, in which the parameters of the trained spatial enhancement model are fixed, the output results are used as the input of the spectral enhancement model, and the spectral enhancement model is fine-tuned.

10. A dual-domain multispectral image enhancement system based on deep learning, characterized in that: Includes the following modules: The spatial enhancement module is used to generate the training data required for training the spatial enhancement sub-network based on blind degradation of remote sensing images and variational inference strategies. The variational inference results are used to construct a spatial enhancement sub-network including a degradation estimation network group and a fusion enhancement network. The degradation estimation network group is used to estimate the blur and noise degradation of the input source image and adjust the fusion enhancement results. The fusion enhancement network adopts a U-shaped network architecture to extract multi-scale details of the source image and ultimately generate a high-spatial-resolution multispectral image. The spectral enhancement module is used to construct a spectral enhancement subnetwork of a dual encoder-single decoder network architecture. The input multispectral image is split into visible light band and near-infrared band at the channel level as the input of the dual encoder. The dual encoder includes a visible light band encoder and a near-infrared band encoder. The multi-scale attention module is used as the basic module of the encoder to extract the multi-scale spatial features of the source image. The decoder adopts the basic structure of CNN and splices the output features of the dual encoder in the channel dimension as the decoder input to generate the spectral enhancement result of the multispectral image. The training and tuning module is used to construct the loss function and adopt the corresponding training strategy to perform joint training and tuning of the spatial enhancer network and the spectral enhancer network.