A Panchromatic Sharpening Method for Remote Sensing Images Based on Domain Prior Depth Unfolding Networks

By using a domain-prior depth unfolding network-based panchromatic sharpening method, HRMS images are reconstructed using features of PAN and LRMS images. This solves the problem of insufficient model interpretability in existing methods and improves panchromatic sharpening performance.

CN119067881BActive Publication Date: 2025-10-28ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410943055.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2025-10-28
Estimated Expiration
2044-07-15

AI Technical Summary

Technical Problem

Existing panchromatic sharpening methods fail to fully utilize domain-specific prior knowledge between images, resulting in poor performance in spatial and spectral reconstruction and a lack of model interpretability.

Method used

A full-color sharpening method based on domain prior deep unfolding network is adopted. Through a multi-stage processing network, combined with data projection module and prior module, the spatial information of PAN image and spectral information of LRMS image are used to reconstruct HRMS image. Spatial reconstruction prior and spectral modulation prior are introduced to enhance the interpretability of the model.

Benefits of technology

The performance of the panchromatic sharpening model has been improved, the spatial and spectral quality of HRMS images has been enhanced, and the problem of insufficient model interpretability in existing methods has been solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119067881B_ABST
    Figure CN119067881B_ABST
Patent Text Reader

Abstract

This invention discloses a panchromatic sharpening method for remote sensing images based on a domain-prior deep unfolded network. It addresses the shortcomings of existing deep learning-based panchromatic sharpening methods, such as insufficient interpretability, and model-based methods, which fail to adequately consider domain-specific prior knowledge. This invention models the panchromatic sharpening problem as a variational model (i.e., a panchromatic sharpening model) with spatial reconstruction and spectral modulation priors. These priors are extended into networked modules, upon which spatial information reconstruction and spectral information modulation modules are constructed. This invention solves the problems of low interpretability and insufficient consideration of domain prior knowledge in current panchromatic sharpening methods, thus improving performance. Extensive experiments on the GaoFen-2 and WorldView-2 satellite datasets demonstrate the effectiveness and superiority of the panchromatic sharpening model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of deep learning and computer vision technology, and particularly relates to a method for panchromatic sharpening of remote sensing images based on domain prior deep unfolding networks. Background Technology

[0002] Recent advancements in remote sensing technology have led to a growing demand for high-resolution multispectral (HRMS) imagery, which is now widely used in fields such as environmental monitoring and agricultural development. However, physical limitations in existing satellite systems make it challenging to simultaneously acquire images with high spatial and spectral resolution. Therefore, panchromatic sharpening has emerged as an alternative, aiming to fuse high-resolution panchromatic (PAN) images with corresponding low-resolution multispectral (LRMS) images to generate HRMS images with high spatial and spectral resolution.

[0003] Panchromatic sharpening methods are mainly divided into traditional methods and deep learning-based methods. Traditional methods include component substitution (CS), multi-resolution analysis (MRA), and variational optimization (VO). However, these methods mostly rely on manually defined priors and constraints, limiting their ability to represent image features and thus restricting their performance. Inspired by the success of deep learning in visual tasks in recent years, various deep learning-based panchromatic sharpening methods have emerged. These methods can be divided into single-branch network and two-branch network types based on the spatial-spectral information reconstruction method. The single-branch network method directly concatenates the LRMS image and the PAN image, and then applies a CNN network to extract spatial and spectral information to reconstruct the HRMS image. The two-branch network method uses a CNN network to extract spectral information from the LRMS image and spatial information from the PAN image, and then fuses them to generate the HRMS image. Although deep learning-based methods have shown superiority in spatial-spectral information reconstruction, they mostly construct the network topology in a black-box manner, without considering the interpretability of the model.

[0004] Therefore, to improve the interpretability of models, researchers have proposed model-based panchromatic sharpening methods, among which deep unfolding networks (DUNs) are widely used as a framework. However, current model-based panchromatic sharpening methods still have shortcomings because they either only simulate the degradation process of HRMS images without considering the complex spatial and spectral relationships between images (e.g., GPPNN), or only consider the macroscopic perspective without detailed analysis from the perspective of spatial or spectral reconstruction, resulting in poor performance (e.g., MMNet). In summary, although existing model-based methods enhance the interpretability of the panchromatic sharpening task, they do not fully consider domain-specific prior knowledge, i.e., the complex spatial and spectral relationships between images, nor do they perform cross-modal interactions, thus limiting their performance.

[0005] Therefore, how to design a full-color sharpening model that can fully utilize the discovered domain knowledge and improve the interpretability of the designed model to improve full-color sharpening performance is a technical problem that urgently needs to be solved. Summary of the Invention

[0006] The technical problem to be solved by this invention is how to make full use of the discovered domain knowledge in the panchromatic sharpening model, that is, to use PAN images to provide the required spatial information, and to use the spectral information in PAN images and LRMS images to jointly reconstruct the spectral information of HRMS images, while improving the interpretability of the panchromatic sharpening model.

[0007] The purpose of this invention is to solve the above-mentioned problems existing in the prior art and to provide a method for panchromatic sharpening of remote sensing images based on a domain prior depth unfolding network.

[0008] To achieve the above-mentioned objectives, the present invention specifically adopts the following technical solution:

[0009] A remote sensing image panchromatic sharpening method based on a domain prior deep unfolding network includes: upsampling the multispectral image to be panchromatic sharpened to obtain an upsampled multispectral image; inputting the multispectral image to be panchromatic sharpened, the upsampled multispectral image, and the panchromatic image to be panchromatic sharpened into a trained panchromatic sharpening model for multi-stage processing; the panchromatic sharpening model contains multiple processing networks, each stage corresponding to a processing network; each stage's processing network consists of a data projection module and a prior module; the data projection module is used to perform feature projection processing on the input, and the prior module is used to reconstruct spatial information and modulate spectral information; after multi-stage processing, the output of the last stage processing network is used as the final reconstructed high-resolution multispectral image; wherein, the prior module adopts a U-shaped architecture with an information perception module; the information perception module includes a perception enhancement module and a feedforward network; the perception enhancement module includes an adaptive fusion module, a spatial information reconstruction module, and a spectral information modulation module.

[0010] In the first-stage processing network, the upsampled multispectral image, the multispectral image to be panchromatic sharpened, and the panchromatic image to be panchromatic sharpened are input into the first-stage data projection module to obtain the first projection feature. The first projection feature and the panchromatic image to be panchromatic sharpened are input into the first-stage prior module to obtain the first spatial spectral feature.

[0011] Apart from the first stage, in the processing networks of the remaining stages, the spatial spectral features obtained from the corresponding previous stage processing network are used as initialization inputs. The initialization inputs, the multispectral image to be panchromatic sharpened, and the panchromatic image to be panchromatic sharpened are input into the data projection modules of the remaining stages to obtain new projection features. The new projection features and the panchromatic image to be panchromatic sharpened are input into the prior modules of the remaining stages to obtain new spatial spectral features.

[0012] Based on the above scheme, each step can be implemented in the following preferred manner.

[0013] Preferably, the data projection module includes an LRMS image processing branch and a PAN image processing branch. In the LRMS image processing branch, a first deep convolutional layer is used to downsample the output of the previous stage processing network to obtain a first feature map. The first feature map is subtracted from the multispectral image to be panchromatic sharpened to obtain a second feature map. Then, a second deep convolutional layer is used to upsample the second feature map to obtain the LRMS feature map output by the LRMS image processing branch. In the PAN image processing branch, a first point convolutional layer is used to reduce the number of output channels of the previous stage processing network to obtain a third feature map. The first feature map is subtracted from the panchromatic image to be panchromatic sharpened to obtain a fourth feature map. Then, a second point convolutional layer is used to increase the number of channels of the fourth feature map to obtain the PAN feature map output by the PAN image processing branch. The LRMS feature map and the PAN feature map are added to obtain a fifth feature map. The fifth feature map is multiplied by the learnable parameter σ to obtain a weighted fifth feature map. The weighted fifth feature map is subtracted from the output of the previous stage processing network to obtain the projection feature output by the data projection module.

[0014] Preferably, in the data projection module, the kernel size of the first depth convolutional layer is 3×3, the kernel size of the second depth convolutional layer is 3×3, the kernel size of the first point convolutional layer is 1×1, and the kernel size of the second point convolutional layer is 1×1.

[0015] Preferably, the prior module takes the projection features output by the data projection module and the panchromatic image to be sharpened as input. First, it uses an embedding layer to map the projection features output by the data mapping module into a projection embedding feature. The embedding layer then maps the panchromatic image to be sharpened into a panchromatic image embedding feature. The projection embedding feature and the panchromatic image embedding feature are input together into the encoder. After passing through the first information perception module and the second information perception module, a sixth feature map corresponding to the projection embedding feature and a seventh feature map corresponding to the panchromatic image embedding feature are obtained. The sixth feature map is then passed through a downsampling layer in the encoder to obtain a downsampled sixth feature map, and the seventh feature map is also passed through a downsampling layer in the encoder to obtain a downsampled seventh feature map. Finally, the downsampled sixth feature map and the downsampled seventh feature map are... Figure 1 The input is fed into the third information perception module, which obtains the eighth feature map corresponding to the sixth feature map and the ninth feature map corresponding to the seventh feature map; the eighth feature map and the ninth feature map are then processed. Figure 1The input is fed into the decoder. First, the eighth feature map is passed through an upsampling layer in the decoder to obtain an upsampled eighth feature map. Then, the ninth feature map is passed through an upsampling layer in the decoder to obtain an upsampled ninth feature map. The upsampled ninth feature map is then concatenated with the seventh feature map along the channel dimension to obtain a first concatenated feature map. This first concatenated feature map is then processed through a third convolutional layer to obtain a processed first concatenated feature map. Finally, the processed first concatenated feature map is concatenated with the upsampled ninth feature map to obtain a second concatenated feature map. The image and the upsampled eighth feature map are sequentially passed through the fourth and fifth information perception modules in the decoder to obtain the first deep feature map corresponding to the panchromatic image embedding features and the second deep feature map corresponding to the projection embedding features. The first deep feature map is then mapped by the fourth convolutional layer to obtain the first mapped deep feature map. The second deep feature map is then mapped by the fourth convolutional layer to obtain the second mapped deep feature map. The second mapped deep feature map is then added to the projection features output by the data projection module to obtain the spatial spectral features output by the prior module.

[0016] Preferably, in the prior module, the embedding layer is formed by cascading a first convolutional layer and a second convolutional layer, with the kernel size of the first and second convolutional layers both being 1×1, the kernel size of the third convolutional layer being 1×1, and the kernel size of the fourth convolutional layer being 3×3.

[0017] Preferably, in the information perception module, the feature map corresponding to the projection embedding feature is used as the first input, and the feature map corresponding to the panchromatic image embedding feature is used as the second input. The first input is first normalized by a first normalization layer to obtain a normalized first input. The normalized first input and the second input are processed together by the perception enhancement module to obtain a first intermediate output corresponding to the first input and a second output corresponding to the second input. The first input and the first intermediate output are residually connected and added element-wise to obtain the second intermediate output. The second intermediate output is sequentially passed through a second normalization layer and a feedforward network to obtain a third intermediate output. The second intermediate output and the third intermediate output are residually connected and added element-wise to obtain the first output corresponding to the first input.

[0018] Preferably, in the perception enhancement module, the normalized first input is processed by the spectral information modulation module, the second input is processed by the spatial information reconstruction module, the feature map output by the spectral information modulation module and the feature map output by the spatial information reconstruction module are spliced ​​along the channel dimension as the input of the adaptive fusion module, and finally the feature map output by the adaptive fusion module is used as the first intermediate output, and the feature map output by the spatial information reconstruction module is used as the second output.

[0019] In the spectral information modulation module, the normalized first input is passed through a third deep convolutional layer and a discrete Fourier transform to obtain a query feature map in the frequency domain. The normalized first input is then passed through a fourth deep convolutional layer and a discrete Fourier transform to obtain a key feature map in the frequency domain. The normalized first input is then passed through a fifth deep convolutional layer to obtain a value feature map in the spatial domain. In the frequency domain, the query feature map and the key feature map are multiplied element-wise to obtain the feature weights in the frequency domain. The feature weights in the frequency domain are then transformed back to the spatial domain through an inverse discrete Fourier transform to obtain the feature weights in the spatial domain. The feature weights in the spatial domain are then normalized through a third normalization layer to obtain normalized feature weights. The normalized feature weights are then multiplied by the value feature map in the spatial domain to obtain a weighted value feature map. The weighted value feature map is then processed through a fifth convolutional layer to obtain a fourth intermediate output. The fourth intermediate output is then residually concatenated with the first input to obtain the feature map output by the spectral information modulation module.

[0020] In the spatial information reconstruction module, the second input is segmented into non-overlapping window features of fixed size, and their shapes are reshaped. Each non-overlapping window feature is processed by W-MSA to obtain the attention map of each head. Finally, the attention maps of the heads are stitched together and the windows are merged along the channel dimension to obtain the feature map output by the spatial information reconstruction module.

[0021] In the adaptive fusion module, the input to the adaptive fusion module is fed into the sixth convolutional layer for dimensionality reduction to obtain a low-dimensional embedding feature map. The low-dimensional embedding feature map is then fed into the seventh convolutional layer to obtain the fifth intermediate output. The low-dimensional embedding feature map is then fed into the eighth convolutional layer to obtain the sixth intermediate output. The fifth and sixth intermediate outputs are added element-wise to obtain the seventh intermediate output. The seventh intermediate output is then processed by the sigmoid activation function to obtain a gating signal. The gating signal is multiplied by the input to the adaptive fusion module to obtain the eighth intermediate output. The eighth intermediate output is then residually concatenated with the input to the adaptive fusion module to obtain the feature map output by the adaptive fusion module.

[0022] Preferably, in the perception enhancement module, the kernel size of the third deep convolutional layer is 3×3, the kernel size of the fourth deep convolutional layer is 3×3, the kernel size of the fifth deep convolutional layer is 3×3, the kernel size of the fifth convolutional layer is 1×1, the kernel size of the sixth convolutional layer is 3×3, the kernel size of the seventh convolutional layer is 1×5, and the kernel size of the eighth convolutional layer is 5×1.

[0023] As a preferred embodiment, the feedforward network is composed of a ninth convolutional layer, a GELU activation function, a sixth deep convolutional layer, a GELU activation function, and a tenth convolutional layer cascaded in sequence. The kernel size of the ninth convolutional layer is 1×1, the kernel size of the tenth convolutional layer is 1×1, and the kernel size of the sixth deep convolutional layer is 1×1.

[0024] Preferably, the pancolor sharpening model is trained in advance using training data generated after the Wald protocol before being used for actual pancolor sharpening; the loss function used in the training of the pancolor sharpening model is the mean absolute error between the output image of the pancolor sharpening model and the real image.

[0025] Compared with the prior art, the present invention has the following advantages:

[0026] This invention models the panchromatic sharpening problem as a minimization problem of a variational model, and models the spatial and spectral relationships among PAN, LRMS, and HRMS images based on actual observation and analysis, introducing spatial reconstruction and spectral modulation priors. In this way, the invention enhances the spatial and spectral quality of the reconstructed HRMS image. Based on the observed and analyzed spatial and spectral relationships among PAN, LRMS, and HRMS images, this invention designs the panchromatic sharpening model SSPEDUN, which fully utilizes domain prior knowledge and improves the interpretability of the panchromatic sharpening model, solving the problems existing in current panchromatic sharpening methods. The information perception enhancement module in the SSPEDUN panchromatic sharpening model of this invention networks the spatial reconstruction and spectral modulation priors, enabling the extraction of high-quality spatial information of the PAN image in the spatial domain and fine modulation of the spectral relationships between the inputs in the frequency domain, thereby improving the performance of panchromatic sharpening. Experiments on the GaoFen-2 and WorldView-2 satellite datasets have verified that the pancolor sharpening model proposed in this invention can effectively solve the problems existing in current pancolor sharpening methods for remote sensing images and improve the performance of pancolor sharpening for remote sensing images. Attached Figure Description

[0027] Figure 1 This is a model structure diagram of the full-color sharpening model of the present invention;

[0028] Figure 2 This is a schematic diagram of the data projection module of the present invention;

[0029] Figure 3 This is a schematic diagram of the prior art module of the present invention;

[0030] Figure 4 This is a schematic diagram of the information sensing module of the present invention;

[0031] Figure 5This is a schematic diagram of the feedforward network (FFN) of the present invention;

[0032] Figure 6 This is a schematic diagram of the perception enhancement module of the present invention;

[0033] Figure 7 This is a flowchart illustrating the training and testing process of a full-color sharpening model in an embodiment of the present invention. Detailed Implementation

[0034] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.

[0035] In the description of this invention, it should be understood that the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include at least one of those features.

[0036] Recent advancements in remote sensing technology have led to a growing demand for high-resolution multispectral (HRMS) images, which are widely used in environmental monitoring and agricultural development. However, physical limitations in existing satellite systems make it challenging to simultaneously acquire images with high spatial and spectral resolution. Therefore, panchromatic sharpening has emerged as an alternative, aiming to fuse high-resolution panchromatic (PAN) images with corresponding low-resolution multispectral (LRMS) images to generate HRMS images with high spatial and spectral resolution. While existing deep learning-based panchromatic sharpening methods demonstrate superior performance in spatial-spectral information reconstruction, they largely construct network topologies in a black-box manner, neglecting model interpretability. Model-based panchromatic sharpening methods enhance interpretability, but they either only simulate the degradation process of HRMS images without fully considering domain-specific prior knowledge—the complex spatial and spectral relationships between images (e.g., GPPNN)—or only consider a macroscopic perspective without detailed analysis from a spatial or spectral reconstruction perspective, resulting in poor performance (e.g., MMNet). Therefore, to address these issues, it is necessary to design better panchromatic sharpening models.

[0037] This invention provides a detailed analysis of the spatial and spectral relationships between PAN images, LRMS images, and HRMS images. It reveals that the spatial information in PAN images is highly similar to that in HRMS images; therefore, spatial details in HRMS images can be recovered solely from the spatial information extracted from PAN images. Furthermore, it avoids introducing noise from LRMS images, and the spectral information distribution trend of each band in the LRMS image is inconsistent with that in its corresponding HRMS image. This indicates the necessity of modulating the spectral information of the LRMS image with the spectral information of the PAN image to compensate for the differences in spectral information between the LRMS and HRMS images, thereby reconstructing the spectral information of the target HRMS image.

[0038] The core of this invention is based on the analysis of spatial and spectral relationships between PAN, LRMS, and HRMS images, proposing a panchromatic sharpening method for remote sensing images based on a domain prior deep unfolded network. The proposed method, based on the analysis of spatial and spectral relationships between images, introduces spatial reconstruction and spectral modulation priors. This fully utilizes the discovered domain knowledge while improving the interpretability of the designed panchromatic sharpening model, SSPEDUN, and enhancing panchromatic sharpening performance.

[0039] like Figure 1 As shown, in a preferred implementation of the present invention, the aforementioned remote sensing image panchromatic sharpening method based on a domain prior depth unfolding network specifically involves: upsampling the multispectral image to be panchromatic sharpened (low-resolution multispectral image LRMS) to obtain an upsampled multispectral image; inputting the multispectral image to be panchromatic sharpened (low-resolution multispectral image LRMS), the upsampled multispectral image, and the panchromatic image to be panchromatic sharpened (high-resolution panchromatic image PAN) into a trained panchromatic sharpening model SSPEDUN for multi-stage processing. The panchromatic sharpening model SSPEDUN contains multiple processing networks, with each stage corresponding to a processing network, and each stage's processing network consisting of a data projection module. and a prior module Composition, data projection module Used for feature projection processing of the input, prior module The panchromatic sharpening model, used for reconstructing spatial and modulating spectral information, modulates the spectral relationships between images after multi-stage processing, achieving high-quality reconstruction of spatial and spectral information. The output of the final stage processing network is used as the final reconstructed high-resolution multispectral image. The prior module... A U-shaped architecture with an information sensing module (IPM) is adopted. The IPM includes a perception enhancement module (PEB) and a feedforward network (FFN). The PEB includes an adaptive fusion module (AFB), a spatial information reconstruction module (SpaRB), and a spectral information modulation module (SpeMB).

[0040] In the first-stage processing network, the upsampled multispectral image, the multispectral image to be panchromatic sharpened, and the panchromatic image to be panchromatic sharpened are input into the first-stage data projection module to obtain the first projection feature. The first projection feature and the panchromatic image to be panchromatic sharpened are input into the first-stage prior module to obtain the first spatial spectral feature.

[0041] Apart from the first stage, in the processing networks of the remaining stages, the spatial spectral features obtained from the corresponding previous stage processing network are used as initialization inputs. The initialization inputs, the multispectral image to be panchromatic sharpened, and the panchromatic image to be panchromatic sharpened are input into the data projection modules of the remaining stages to obtain new projection features. The new projection features and the panchromatic image to be panchromatic sharpened are input into the prior modules of the remaining stages to obtain new spatial spectral features.

[0042] The specific structure of the full-color sharpening model in this invention is described in detail below. Figure 1 This is a diagram showing the overall structure of the SSPEDUN panchromatic sharpening model. The SSPEDUN backbone network is a deep unfolded network, further enhanced by incorporating networked modules consisting of spatial reconstruction and spectral modulation priors. The panchromatic sharpening model comprises multiple stages, each corresponding to a processing network. Each stage's processing network has the same structure, consisting of a data projection module. and a prior module Composition. Prior module A U-shaped architecture is adopted, including an encoder, bottleneck, and decoder, with the basic component being the Information Perception Module (IPM). The IPM comprises two normalization layers (LN), a Perception Enhancement Module (PEB), and a Feedforward Network (FFN). The PEB includes a Spatial Information Reconstruction Module (SpaRB), a Spectral Information Modulation Module (SpeMB), and an Adaptive Fusion Module (AFB). The first-stage processing network takes the upsampled multispectral image, the multispectral image to be panchromatic sharpened, and the panchromatic image to be panchromatic sharpened as input. Subsequent processing stages take the output of the previous stage's processing network, the multispectral image to be panchromatic sharpened, and the panchromatic image to be panchromatic sharpened as input. After obtaining the input, the three types of input images first undergo feature projection processing through a data projection module. The projection features generated by the data projection module and the panchromatic image to be panchromatic sharpened are used as input to a priori module. The priori module performs targeted processing on the spatial and spectral information, obtaining the output of each stage and inputting it into the next stage's processing network. After multiple stages of processing, the output of the last stage's processing network is finally used as the reconstructed high-resolution multispectral image.

[0043] In this embodiment, the SSPEDUN full-color sharpening model comprises K stages, each stage containing a data projection module and a priori module. Specifically, as... Figure 1 As shown, for the first stage, the multispectral image L to be panchromatic sharpened is first upsampled by four times to obtain the multispectral image after four times upsampling. Then, the multispectral image was upsampled four times. The multispectral image L to be panchromatic sharpened and the panchromatic image P to be panchromatic sharpened are used as inputs. In subsequent processing, taking the k-th stage processing network as an example, it uses the output of the (k-1)-th stage processing network. The multispectral image L to be panchromatic sharpened and the panchromatic image P to be panchromatic sharpened are used as inputs. First, they undergo feature projection processing through the data mapping module to obtain the k-th projected feature. Then project the k-th feature The panchromatic image P to be sharpened is input into the prior module of the k-th stage. The prior module performs targeted processing on the spatial and spectral information to obtain the k-th spatial spectral features. This is then fed into the (k+1)th stage, where the (k-1)th stage processes the network's output. The dimensions of the multispectral image L to be panchromatic sharpened are (H×W×C), the dimensions of the panchromatic image P to be panchromatic sharpened are (H×W×1), and the dimensions of the k-th spatial spectral features are (H×W×1). The dimensions are (H×W×C). H, W, and C represent the height, width, and number of bands of the image, respectively, while h and w represent the height and width of the unsampled image, respectively. In this embodiment, the full-color sharpening model SSPEDUN includes two stages: K=2, h=w=32, and H=W=128.

[0044] It should be noted that in the data projection module of the present invention, the data projection module includes an LRMS image processing branch and a PAN image processing branch. In the LRMS image processing branch, a first deep convolutional layer is used to downsample the output of the previous stage processing network to obtain a first feature map. The first feature map is subtracted from the multispectral image to be panchromatic sharpened to obtain a second feature map. Then, a second deep convolutional layer is used to upsample the second feature map to obtain the LRMS feature map output by the LRMS image processing branch. In the PAN image processing branch, a first point convolutional layer is used to reduce the number of output channels of the previous stage processing network to obtain a third feature map. The first feature map is subtracted from the panchromatic image to be panchromatic sharpened to obtain a fourth feature map. Then, a second point convolutional layer is used to increase the number of channels of the fourth feature map to obtain the PAN feature map output by the PAN image processing branch. The LRMS feature map and the PAN feature map are added to obtain a fifth feature map. The fifth feature map is multiplied by the learnable parameter σ to obtain a weighted fifth feature map. The weighted fifth feature map is subtracted from the output of the previous stage processing network to obtain the projection feature output by the data projection module.

[0045] It should be noted that in the data projection module of the present invention, the kernel size of the first depth convolutional layer is 3×3, the kernel size of the second depth convolutional layer is 3×3, the kernel size of the first point convolutional layer is 1×1, and the kernel size of the second point convolutional layer is 1×1.

[0046] In this embodiment, the data projection module The data projection module takes the spatial spectral features output from the previous stage processing network, the multispectral image LRMS to be panchromatic sharpened, and the panchromatic image PAN to be panchromatic sharpened as input. The data projection module includes one LRMS image processing branch and one PAN image processing branch. The LRMS image processing branch contains two 3×3 depthwise convolutional layers, and the PAN image processing branch contains two 1×1 pointwise convolutional layers, as follows... Figure 2 As shown, a two-branch approach is used for processing. Taking the k-th stage as an example, the data projection module processes the (k-1)-th spatial spectral features of the network output in the (k-1)-th stage. The multispectral image L to be panchromatic sharpened and the panchromatic image P to be panchromatic sharpened are used as inputs. In the LRMS image processing branch, the data projection module uses the first depth convolutional layer to process the (k-1)th spatial spectral features of the network output in the (k-1)th stage. Downsampling is performed to reduce its size to half, resulting in the first feature map. This first feature map is then subtracted from the multispectral image to be panchromatic sharpened, yielding the second feature map. A second depthwise convolutional layer is then used to upsample the second feature map, restoring its original size, resulting in the LRMS feature map output by the LRMS image processing branch. In the PAN image processing branch, the first point convolutional layer is used to process the (k-1)th spatial spectral feature map output by the (k-1)th stage processing network. The number of channels is reduced (from C to 1) to obtain the third feature map. The first feature map is subtracted from the panchromatic image to be sharpened to obtain the fourth feature map. Then, the number of channels in the fourth feature map is increased (from 1 to C) using the second point convolutional layer to obtain the PAN feature map output by the PAN image processing branch. Finally, the data projection module adds the LRMS feature map and the PAN feature map to obtain the fifth feature map. The fifth feature map is multiplied by the learnable parameter σ to obtain the weighted fifth feature map. The weighted fifth feature map is then combined with the (k-1)th spatial spectral feature map output by the (k-1)th stage processing network. Subtracting the two yields the k-th projection feature output by the data projection module. The dimension is (H×W×C). In this embodiment, the feature dimension output by the data projection module is 128×128×4.

[0047] It should be noted that in the prior art module of this invention, the projection features output by the data projection module and the panchromatic image to be sharpened are used as input. First, an embedding layer is used to map the projection features output by the data mapping module into a projection embedding feature. The embedding layer maps the panchromatic image to be sharpened into a panchromatic image embedding feature. The projection embedding feature and the panchromatic image embedding feature are input together into the encoder. After passing through the first information perception module and the second information perception module in sequence, a sixth feature map corresponding to the projection embedding feature and a seventh feature map corresponding to the panchromatic image embedding feature are obtained. The sixth feature map is passed through a downsampling layer in the encoder to obtain a downsampled sixth feature map. The seventh feature map is passed through a downsampling layer in the encoder to obtain a downsampled seventh feature map. The downsampled sixth feature map and the downsampled seventh feature map are then processed together. Figure 1 The input is fed into the third information perception module, which obtains the eighth feature map corresponding to the sixth feature map and the ninth feature map corresponding to the seventh feature map; the eighth feature map and the ninth feature map are then processed. Figure 1The input is fed into the decoder. First, the eighth feature map is passed through an upsampling layer in the decoder to obtain an upsampled eighth feature map. Then, the ninth feature map is passed through an upsampling layer in the decoder to obtain an upsampled ninth feature map. The upsampled ninth feature map is then concatenated with the seventh feature map along the channel dimension to obtain a first concatenated feature map. This first concatenated feature map is then processed through a third convolutional layer to obtain a processed first concatenated feature map. Finally, the processed first concatenated feature map is concatenated with the upsampled ninth feature map to obtain a second concatenated feature map. The image and the upsampled eighth feature map are sequentially passed through the fourth and fifth information perception modules in the decoder to obtain the first deep feature map corresponding to the panchromatic image embedding features and the second deep feature map corresponding to the projection embedding features. The first deep feature map is then mapped by the fourth convolutional layer to obtain the first mapped deep feature map. The second deep feature map is then mapped by the fourth convolutional layer to obtain the second mapped deep feature map. The second mapped deep feature map is then added to the projection features output by the data projection module to obtain the spatial spectral features output by the prior module.

[0048] It should be noted that in the prior module of the present invention, the embedding layer is formed by cascading a first convolutional layer and a second convolutional layer in sequence. The kernel size of the first convolutional layer and the second convolutional layer is 1×1, the kernel size of the third convolutional layer is 1×1, and the kernel size of the fourth convolutional layer is 3×3.

[0049] In this embodiment, the prior module With data projection module Output projection features The panchromatic image P to be sharpened is used as input, and a U-shaped structure containing an encoder, bottleneck, and decoder is employed, such as... Figure 3As shown in the diagram, the encoder and decoder each contain two Information Perception Modules (IPMs) and a resized 3×3 depthwise convolutional layer for upsampling or downsampling. The bottleneck section contains only one IPM. The prior module first uses an embedding layer to map the projection features output by the data projection module to a projection embedding feature X0. This embedding layer maps the panchromatic image to be sharpened to a panchromatic image embedding feature P0. This embedding layer is composed of a first convolutional layer and a second convolutional layer cascaded sequentially, with both convolutional kernels being 1×1. The deep features are then processed sequentially through the encoder, bottleneck, and decoder. Specifically, the projection embedding feature X0 and the panchromatic image embedding feature P0 are first input into the encoder. After passing through the first and second information perception modules, a sixth feature map corresponding to the projection embedding feature and a seventh feature map corresponding to the panchromatic image embedding feature are obtained. The sixth feature map is then passed through a downsampling layer in the encoder to obtain a downsampled sixth feature map, and the seventh feature map is then passed through a downsampling layer in the encoder to obtain a downsampled seventh feature map. After encoder processing, the downsampled sixth and seventh feature maps are processed by the information perception module of the bottleneck section, respectively. Figure 1 The input is fed into the third information perception module, which obtains the eighth feature map corresponding to the sixth feature map and the ninth feature map corresponding to the seventh feature map. After processing by the third information perception module, the decoder further processes the input, converting the eighth and ninth feature maps into their corresponding data. Figure 1 The input is fed into the decoder. First, the eighth feature map is passed through an upsampling layer in the decoder to obtain an upsampled eighth feature map. Then, the ninth feature map is passed through an upsampling layer in the decoder to obtain an upsampled ninth feature map. Furthermore, to facilitate information interaction, the output of the upsampling layer in the decoder is concatenated along the channel dimension with the output of the second information perception module in the encoder. This concatenation is then processed by a 1×1 convolutional layer before being input into the fourth information perception module in the decoder. Specifically, the upsampled ninth feature map is concatenated with the seventh feature map along the channel dimension to obtain a first concatenated feature map. This first concatenated feature map is processed by a 1×1 third convolutional layer to obtain a processed first concatenated feature map. This processed first concatenated feature map is then concatenated with the upsampled ninth feature map to obtain a second concatenated feature map. The second concatenated feature map and the upsampled eighth feature map are then passed sequentially through the fourth and fifth information perception modules in the decoder to obtain the first deep feature map P corresponding to the embedded features of the panchromatic image. d And the second deep feature map X corresponding to the projection embedding features. d Finally, the first deep feature map P... dThe first deep feature map P′ is obtained by feature mapping using a 3×3 fourth convolutional layer, and the second deep feature map X is obtained by mapping. d The fourth convolutional layer performs feature mapping to obtain the second mapped deep feature map R. This second mapped deep feature map R is then added to the projected features output by the data projection module to obtain the spatial spectral features output by the prior module. In this embodiment, the prior module outputs a feature dimension of 128×128×4.

[0050] It should be noted that the information sensing module IPM in this invention includes two normalization layers (LN), a perception enhancement module (PEB), and a feedforward network (FFN), as follows: Figure 4 As shown. Taking the first information perception module as an example, it takes the projection embedding feature X0 and the panchromatic image embedding feature P0 as inputs. In this embodiment, it uses the feature map corresponding to the projection embedding feature as the first input X. in The feature map corresponding to the embedded features of the panchromatic image is used as the second input P. in The first input is X in First, the input is normalized using the first normalization layer to obtain the normalized first input. Then, the normalized first input and the second input P are compared. in After being processed by the Perception Enhancement Module (PEB), the first intermediate output X corresponding to the first input is obtained. mid and the second output P corresponding to the second input out , take the first input X in With the first intermediate output X mid After performing a residual connection and adding element-wise, a second intermediate output is obtained. This second intermediate output is then passed sequentially through a second normalization layer and a feedforward network (FFN) to obtain a third intermediate output. Finally, the second and third intermediate outputs are residually connected and added element-wise to obtain the first output X corresponding to the first input. out .

[0051] It should be noted that in the perception enhancement module of the present invention, the normalized first input is processed by the spectral information modulation module, and the second input is processed by the spatial information reconstruction module. The feature map output by the spectral information modulation module and the feature map output by the spatial information reconstruction module are concatenated along the channel dimension as the input of the adaptive fusion module. Finally, the feature map output by the adaptive fusion module is used as the first intermediate output, and the feature map output by the spatial information reconstruction module is used as the second output.

[0052] In the spectral information modulation module, the normalized first input is passed through a third deep convolutional layer and a discrete Fourier transform to obtain a query feature map in the frequency domain. The normalized first input is then passed through a fourth deep convolutional layer and a discrete Fourier transform to obtain a key feature map in the frequency domain. The normalized first input is then passed through a fifth deep convolutional layer to obtain a value feature map in the spatial domain. In the frequency domain, the query feature map and the key feature map are multiplied element-wise to obtain the feature weights in the frequency domain. The feature weights in the frequency domain are then transformed back to the spatial domain through an inverse discrete Fourier transform to obtain the feature weights in the spatial domain. The feature weights in the spatial domain are then normalized through a third normalization layer to obtain normalized feature weights. The normalized feature weights are then multiplied by the value feature map in the spatial domain to obtain a weighted value feature map. The weighted value feature map is then processed through a fifth convolutional layer to obtain a fourth intermediate output. The fourth intermediate output is then residually concatenated with the first input to obtain the feature map output by the spectral information modulation module.

[0053] In the spatial information reconstruction module, the second input is segmented into non-overlapping window features of fixed size, and their shapes are reshaped. Each non-overlapping window feature is processed by W-MSA to obtain the attention map of each head. Finally, the attention maps of the heads are stitched together and the windows are merged along the channel dimension to obtain the feature map output by the spatial information reconstruction module.

[0054] In the adaptive fusion module, the input to the adaptive fusion module is fed into the sixth convolutional layer for dimensionality reduction to obtain a low-dimensional embedding feature map. The low-dimensional embedding feature map is then fed into the seventh convolutional layer to obtain the fifth intermediate output. The low-dimensional embedding feature map is then fed into the eighth convolutional layer to obtain the sixth intermediate output. The fifth and sixth intermediate outputs are added element-wise to obtain the seventh intermediate output. The seventh intermediate output is then processed by the sigmoid activation function to obtain a gating signal. The gating signal is multiplied by the input to the adaptive fusion module to obtain the eighth intermediate output. The eighth intermediate output is then residually concatenated with the input to the adaptive fusion module to obtain the feature map output by the adaptive fusion module.

[0055] The intermediate processes and parameters of the perception enhancement module are further described below to help those skilled in the art better understand the principles behind its use. The perception enhancement module uses the Spatial Information Reconstruction (SpaRB) module to process the embedded features of the PAN image to extract high-quality spatial features, and the Spectral Information Modulation (SpeMB) module to process the embedded features output by the data projection module to accurately modulate the complex spectral relationships between the inputs. The SpaRB module utilizes a custom self-attention operation based on spatial domain local window processing to extract high-quality spatial details, while the SpeMB module uses element-wise multiplication in the frequency domain to modulate the spectral information between the inputs. The outputs of the SpaRB and SpeMB modules are concatenated along the channel dimension and then adaptively fused by the adaptive fusion module to obtain the output features.

[0056] In this embodiment, the perception enhancement module PEB includes a spatial information reconstruction module (SpaRB), a spectral information modulation module (SpeMB), and an adaptive fusion module (AFB), such as... Figure 6 As shown. In the perception enhancement module, the normalized first input is processed by the spectral information modulation module SpeMB, and the second input is processed by the spatial information reconstruction module SpaRB. The feature map output by the spectral information modulation module SpeMB and the feature map output by the spatial information reconstruction module SpaRB are concatenated along the channel dimension as the input to the adaptive fusion module. Finally, the feature map F output by the adaptive fusion module is... out As the first intermediate output X mid The feature map X output by the SpaRB spatial information reconstruction module spa As the second output P out .

[0057] The following further describes the intermediate processes and parameters of the spectral information modulation module to help those skilled in the art better understand the principles of using the spectral information modulation module. The spectral information modulation module (SpMB) is a networked module for spectral modulation priors, modulating the spectral relationships between inputs to obtain high-quality spectral information. Specifically, the spectral information modulation module SpMB takes the normalized first input as input and processes it using three 3×3 deep convolutional layers in a three-branch configuration. To better modulate the relationships between spectra, the spectral information modulation module SpMB transforms the feature maps output from the first two branches' deep convolutional layers into the frequency domain using Discrete Fourier Transform (DFT). Specifically, the first branch processes the normalized first input through a 3×3 third deep convolutional layer and a DFT to obtain the query feature map X in the frequency domain. qThe second branch processes the normalized first input through a 3×3 fourth-depth convolutional layer and a discrete Fourier transform to obtain the key feature map X in the frequency domain. k The third branch processes the normalized first input through a 3×3 fifth-depth convolutional layer to obtain the spatial feature map X. v According to the convolution theorem, the convolution of two features in the spatial domain is equivalent to their element-wise multiplication in the frequency domain. Therefore, the spectral information modulation module SpeMB modulates the query feature map X in the frequency domain. q Bond feature map X k Element-wise multiplication is performed to obtain the feature weights W in the frequency domain with low computational cost.

[0058] W=X q ⊙X k

[0059] The feature weights W in the frequency domain are transformed back to the spatial domain through the inverse discrete Fourier transform (IDFT) to obtain the feature weights in the spatial domain. These spatial feature weights are then normalized through a third normalization layer to obtain normalized feature weights. Finally, these normalized feature weights are compared with the value feature map X in the spatial domain. v Multiply to obtain a weighted feature map. After processing the weighted feature map through the fifth convolutional layer, the fourth intermediate output is obtained. The fourth intermediate output is then residually concatenated with the first input to obtain the feature map output by the spectral information modulation module SpeMB.

[0060] The following further describes the intermediate processes and parameters of the spatial information reconstruction module to help those skilled in the art better understand the principles behind its use. The SpaRB spatial information reconstruction module is a networked module for spatial reconstruction priors, extracting high-quality spatial information from the embedded features of the PAN image. Specifically, the spatial information reconstruction module uses the second input P... in As input, it is divided into non-overlapping windows of size M×M, and its shape is reshaped as follows: Subsequently, each non-overlapping window feature is processed by W-MSA, and a query feature P is generated by linear projection operation. q Key features P k Value characteristics P v And divide them into h heads along the channel dimension. The dimensions of each head are: These represent the 1st, 2nd, ..., hth headers of the query feature, respectively. These represent the 1st, 2nd, ..., hth headers of the key features, respectively. These represent the 1st, 2nd, ..., hth headers of the value features, respectively. For ease of representation, Figure 6Only the case where h=1 is shown. Attention maps for each head are calculated. Finally, the attention maps of the heads are stitched together along the channel dimension and the windows are merged to obtain the feature map X output by the spatial information reconstruction module. spa .

[0061] The self-attention of the i-th head is calculated as follows:

[0062]

[0063] in, This represents the i headers of the query feature. The i-th header represents the key feature. Represents the i heads of the value feature; d h =C / 2h represents the dimension of each head; Pos i This represents a learnable parameter used to embed positional information; T represents transpose; and i represents the head index.

[0064] In this embodiment, the SpaRB spatial information reconstruction module divides the input into 8×8 non-overlapping windows with a header of 2.

[0065] The intermediate processes and parameters of the adaptive fusion module are further described below to help those skilled in the art better understand the principle of using the adaptive fusion module. The adaptive fusion module AFB contains a 3×3 convolutional layer, a 5×1 convolutional layer, and a 1×5 convolutional layer, which uses the output X of the aforementioned spatial information reconstruction module SpaRB. spa The output X of the SpeMB spectral information modulation module spe As input, they are concatenated along the channel dimension to obtain the feature map F. in The input is fed into a 3×3 sixth convolutional layer, compressing it into a low-dimensional embedded feature map F′. in The dimensions are (2×H×W). Subsequently, the low-dimensional embedding feature maps are input into 1×5 and 5×1 convolutional layers respectively, expanding the receptive field of the embedding space along the vertical and horizontal directions, allowing the adaptive fusion module to focus on learning complementary features at a lower cost: the low-dimensional embedding feature map is input into the 1×5 seventh convolutional layer to obtain the fifth intermediate output, and the low-dimensional embedding feature map is input into the 5×1 eighth convolutional layer to obtain the sixth intermediate output. The embedding spaces in both directions are aggregated using element-wise addition: the fifth and sixth intermediate outputs are added element-wise to obtain the seventh intermediate output F″. in The seventh intermediate output is processed using the Sigmoid activation function to obtain the gated signal. Finally, the gated signal is compared with the feature map F. inAfter multiplication, the eighth intermediate output is obtained. This eighth intermediate output is then residually concatenated with the input of the adaptive fusion module to obtain the high-quality fused feature F. out .

[0066] It should be noted that in this invention, the feedforward network FFN is composed of a ninth convolutional layer, a GELU activation function, a sixth deep convolutional layer, a GELU activation function, and a tenth convolutional layer cascaded sequentially, as follows: Figure 5 As shown in the figure. The kernel size of the ninth convolutional layer is 1×1, the kernel size of the tenth convolutional layer is 1×1, and the kernel size of the sixth depth convolutional layer is 1×1.

[0067] It should be noted that the SSPEDUN pancolor sharpening model is trained in advance using training data generated after the Wald protocol before being used for actual pancolor sharpening; the loss function used in the training of the pancolor sharpening model is the mean absolute error between the output image of the pancolor sharpening model and the real image.

[0068] To better demonstrate the specific implementation and technical effects of the present invention, the remote sensing image panchromatic sharpening method based on a domain prior depth unfolding network shown in the above preferred implementation is applied to a specific example below.

[0069] Example

[0070] The overall process in this embodiment can be divided into three stages: data preprocessing, model training, and image prediction, as detailed below. Figure 7 As shown.

[0071] 1. Data Preprocessing Stage

[0072] Step 1: For the obtained raw remote sensing images LRMS and PAN, preprocessing is performed. First, image cropping and flipping are performed, then data augmentation is performed and the images are processed into the same size (128*128).

[0073] Step 2: Generate the ground truth image GroundTruth using the Wald protocol. Downsample the preprocessed LRMS and PAN images to one-quarter of their original size and use them as the original input images for the pan-color sharpening model. Use the original LRMS image as the ground truth image GroundTruth for the pan-color sharpening model.

[0074] 2. Model Training

[0075] Step 1: Construct the training dataset and divide it into batches of a fixed batch size, with a total of N.

[0076] Step 2: Select a batch of training samples with index i sequentially from the training dataset, where i∈{0,1,…,N}. Train the full-color sharpening model SSPEDUN using each batch of training samples. The specific structure of SSPEDUN is as described above and will not be repeated here. During training, calculate the mean absolute error for each training sample. And based on the total loss of all training samples in the batch The network parameters in the entire model are adjusted until all batches of the training dataset have participated in the training of the pan-color sharpening model. After reaching the specified number of iterations, the pan-color sharpening model converges, and training is complete.

[0077] 3. Pancolor sharpening of remote sensing images

[0078] The test set images are directly input into the trained panchromatic sharpening model SSPEDUN, and the final predicted HRMS image with both high resolution and multispectral resolution is obtained, thus achieving panchromatic sharpening.

[0079] In this embodiment, the test results are shown in Table 1:

[0080] Table 1. Test Results

[0081]

[0082] This embodiment uses Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), Spectral Angle Mapper (SAM), and Relative Dimensionless Comprehensive Error (ERGAS) as evaluation metrics for model performance. Compared with traditional pancolor sharpening methods, deep learning-based pancolor sharpening methods, and model-based pancolor sharpening methods, it can be seen that the pancolor sharpening model SSPEDUN of this invention can achieve pancolor sharpening effect on remote sensing images very well, generating HRMS images with both high resolution and multispectral density. While the reconstructed HRMS images are smoother, it also shows a certain improvement in performance compared with other model-based methods (GPPNN, MDCUN, LGTEUN).

[0083] In Table 1, the methods compared with the pan-sharpening model of this embodiment are as follows: SFIM: Liu J G. Smoothing filter-based intensity modulation: A spectral preserve image fusion technique for improving spatial details[J]. International Journal of remote sensing, 2000, 21(18): 3461-3472. BICUBIC: Huang Z, Cao L. Bicubic interpolation and extrapolation iteration method for high resolution digital holographic reconstruction[J]. Optics and lasers in engineering, 2020, 130: 106090. Wavelet: King R L, Wang J. A wavelet based algorithm for pan sharpening Landsat 7 imagery[C] / / IGARSS 2001. Scanning the Present and Resolving the Future. Proceedings. IEEE 2001 International Geoscience and Remote Sensing Symposium (Cat. No. 01CH37217). IEEE, 2001, 2: 849-851. SRPPNN: Cai J, Huang B. Super-resolution-guided progressive pan sharpening based on a deep convolutional neural network[J]. IEEE Transactions on Geoscience and Remote Sensing, 2020, 59(6): 5206-5220.MSDCNN:Yuan Q,Wei Y,Meng X,et al.Amultiscale and multidepthconvolutional neural network for remote sensing imagery pan-sharpening[J].IEEE Journal of Selected Topics in Applied Earth Observations and RemoteSensing,2018,11(3):978-989。GPPNN:Xu S,Zhang J,Zhao Z,et al.Deep gradientprojection networks for pan-sharpening[C] / / Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition.2021:1366-1375。MutInf:Zhou M,Yan K,Huang J,et al.Mutual information-driven pan-sharpening[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition.2022:1798-1808。HSIT:Bandara W G C,Patel V M.HyperTransformer:Atextural and spectral feature fusion transformer for pansharpening[C] / / Proceedings of the IEEE / CVF conference on computer vision and patternrecognition.2022:1767-1777。MSDDN:He X,Yan K,Zhang J,et al.Multi-scaledual-domain guidance network for pan-sharpening[J].IEEE Transactions on Geoscienceand Remote Sensing,2023。MDCUN: Yang G, Zhou M, Yan K, et al. Memory-augmented deepconditional unfolding network for pan-sharpening[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022:1788-1797. LETEUN: Li M, Liu Y, Xiao T, et al. Local-global transformer enhanced unfolding network for pan-sharpening[J]. arXiv preprint arXiv:2304.14612,2023.

[0084] To further demonstrate the effectiveness of the SpaRB and SpeMB modules designed in the pan-color sharpening model SSPEDUN, this invention conducted ablation experiments on SSPEDUN under four different settings: (I) replacing the SpaRB and SpeMB modules with ordinary convolutional layers respectively; (II) replacing the SpaRB module with ordinary convolutional layers while still using the SpeMB module normally as described above; (III) using the SpaRB module normally as described above, but replacing the SpeMB module with ordinary convolutional layers; (IV) using the SpaRB and SpeMB modules normally as described above, i.e., SSPEDUN (this embodiment). Settings (I) to (III) reduce the degree to which the pan-color sharpening model utilizes prior domain knowledge to varying degrees, thus degrading the performance of the pan-color sharpening model.

[0085] The results of the full-color sharpening model ablation experiment in this embodiment on the GaoFen-2 dataset are shown in Table 2:

[0086] Table 2. Ablation Experiment Results

[0087]

[0088] As can be seen from Table 2, when using convolutional layers to replace the SpaRB or SpeMB modules, the performance of the pancolor sharpening models all decreased to varying degrees. However, the original pancolor sharpening model SSPEDUN achieved the best results on all four metrics of the WorldView-2 dataset. This proves the effectiveness of the SpaRB and SpeMB modules networked in this invention based on spatial reconstruction priors and spectral modulation priors. They utilize prior knowledge of the domain to enable the model to reconstruct better spatial details and spectral information.

[0089] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A method for panchromatic sharpening of remote sensing images based on domain prior deep unfolded networks, characterized in that, include: The multispectral image to be panchromatic sharpened is upsampled to obtain an upsampled multispectral image. The multispectral image to be panchromatic sharpened, the upsampled multispectral image, and the panchromatic image to be panchromatic sharpened are then input into a trained panchromatic sharpening model for multi-stage processing. The panchromatic sharpening model contains multiple processing networks, with each stage corresponding to a processing network. Each stage's processing network consists of a data projection module and a prior module. The data projection module is used to perform feature projection processing on the input, and the prior module is used to reconstruct spatial information and modulate spectral information. After multi-stage processing, the output of the last stage processing network is used as the final reconstructed high-resolution multispectral image. The prior module adopts a U-shaped architecture with an information perception module. The information perception module includes a perception enhancement module and a feedforward network. The perception enhancement module includes an adaptive fusion module, a spatial information reconstruction module, and a spectral information modulation module. In the first-stage processing network, the upsampled multispectral image, the multispectral image to be panchromatic sharpened, and the panchromatic image to be panchromatic sharpened are input into the first-stage data projection module to obtain the first projection feature. The first projection feature and the panchromatic image to be panchromatic sharpened are input into the first-stage prior module to obtain the first spatial spectral feature. Apart from the first stage, in the processing networks of the remaining stages, the spatial spectral features obtained from the corresponding previous stage processing network are used as initialization inputs. The initialization inputs, the multispectral image to be panchromatic sharpened, and the panchromatic image to be panchromatic sharpened are input into the data projection modules of the remaining stages to obtain new projection features. The new projection features and the panchromatic image to be panchromatic sharpened are input into the prior modules of the remaining stages to obtain new spatial spectral features.

2. The remote sensing image panchromatic sharpening method based on a domain prior depth unfolding network as described in claim 1, characterized in that, The data projection module includes an LRMS image processing branch and a PAN image processing branch. In the LRMS image processing branch, a first deep convolutional layer is used to downsample the output of the previous stage processing network to obtain a first feature map. The first feature map is subtracted from the multispectral image to be panchromatic sharpened to obtain a second feature map. Then, a second deep convolutional layer is used to upsample the second feature map to obtain the LRMS feature map output by the LRMS image processing branch. In the PAN image processing branch, a first point convolutional layer is used to reduce the number of output channels of the previous stage processing network to obtain a third feature map. The first feature map is subtracted from the panchromatic image to be panchromatic sharpened to obtain a fourth feature map. Then, a second point convolutional layer is used to increase the number of channels of the fourth feature map to obtain the PAN feature map output by the PAN image processing branch. The LRMS feature map and the PAN feature map are added to obtain a fifth feature map. The fifth feature map is multiplied by the learnable parameter σ to obtain a weighted fifth feature map. The weighted fifth feature map is subtracted from the output of the previous stage processing network to obtain the projection feature output by the data projection module.

3. The remote sensing image panchromatic sharpening method based on a domain prior depth unfolding network as described in claim 2, characterized in that, In the data projection module, the kernel size of the first depth convolutional layer is 3×3, the kernel size of the second depth convolutional layer is 3×3, the kernel size of the first point convolutional layer is 1×1, and the kernel size of the second point convolutional layer is 1×1.

4. The remote sensing image panchromatic sharpening method based on a domain prior depth unfolding network as described in claim 1, characterized in that, The prior module takes the projection features output by the data projection module and the panchromatic image to be sharpened as input. First, it uses an embedding layer to map the projection features output by the data mapping module into a projection embedding feature. The embedding layer then maps the panchromatic image to be sharpened into a panchromatic image embedding feature. The projection embedding feature and the panchromatic image embedding feature are input together into the encoder. After passing through the first and second information perception modules, a sixth feature map corresponding to the projection embedding feature and a seventh feature map corresponding to the panchromatic image embedding feature are obtained. The sixth feature map is then passed through a downsampling layer in the encoder to obtain a downsampled sixth feature map. Similarly, the seventh feature map is passed through a downsampling layer in the encoder to obtain a downsampled seventh feature map. The downsampled sixth and seventh feature maps are then input together into the third information perception module to obtain an eighth feature map corresponding to the sixth feature map and a ninth feature map corresponding to the seventh feature map. Finally, the eighth and ninth feature maps are input together into the decoder. The eighth feature map is first upsampled in the decoder. After sampling, the eighth feature map is obtained after upsampling. The ninth feature map is then passed through an upsampling layer in the decoder to obtain another upsampling feature map. The upsampling ninth feature map is then concatenated with the seventh feature map along the channel dimension to obtain the first concatenated feature map. The first concatenated feature map is then processed through a third convolutional layer to obtain a processed first concatenated feature map. The processed first concatenated feature map is then concatenated with the upsampling ninth feature map to obtain the second concatenated feature map. The second concatenated feature map and the upsampling eighth feature map are then passed sequentially through the fourth and fifth information perception modules in the decoder to obtain the first deep feature map corresponding to the panchromatic image embedding feature and the second deep feature map corresponding to the projection embedding feature. The first deep feature map is then feature-mapped by the fourth convolutional layer to obtain the first mapped deep feature map. The second deep feature map is then feature-mapped by the fourth convolutional layer to obtain the second mapped deep feature map. The second mapped deep feature map is then added to the projection feature output by the data projection module to obtain the spatial spectral feature output by the prior module.

5. The remote sensing image panchromatic sharpening method based on a domain prior depth unfolding network as described in claim 4, characterized in that, In the prior module, the embedding layer is formed by cascading a first convolutional layer and a second convolutional layer. The kernel size of the first and second convolutional layers is 1×1, the kernel size of the third convolutional layer is 1×1, and the kernel size of the fourth convolutional layer is 3×3.

6. The remote sensing image panchromatic sharpening method based on a domain prior depth unfolding network as described in claim 1, characterized in that, In the information perception module, the feature map corresponding to the projection embedding feature is used as the first input, and the feature map corresponding to the panchromatic image embedding feature is used as the second input. The first input is first normalized by a first normalization layer to obtain a normalized first input. The normalized first input and the second input are processed together by the perception enhancement module to obtain a first intermediate output corresponding to the first input and a second output corresponding to the second input. The first input and the first intermediate output are residually connected and added element by element to obtain the second intermediate output. The second intermediate output is then passed through a second normalization layer and a feedforward network to obtain a third intermediate output. The second intermediate output and the third intermediate output are residually connected and added element by element to obtain the first output corresponding to the first input.

7. The remote sensing image panchromatic sharpening method based on a domain prior depth unfolding network as described in claim 6, characterized in that, In the perception enhancement module, the normalized first input is processed by the spectral information modulation module, and the second input is processed by the spatial information reconstruction module. The feature map output by the spectral information modulation module and the feature map output by the spatial information reconstruction module are concatenated along the channel dimension as the input of the adaptive fusion module. Finally, the feature map output by the adaptive fusion module is used as the first intermediate output, and the feature map output by the spatial information reconstruction module is used as the second output. In the spectral information modulation module, the normalized first input is passed through a third deep convolutional layer and a discrete Fourier transform to obtain a query feature map in the frequency domain. The normalized first input is then passed through a fourth deep convolutional layer and a discrete Fourier transform to obtain a key feature map in the frequency domain. The normalized first input is then passed through a fifth deep convolutional layer to obtain a value feature map in the spatial domain. In the frequency domain, the query feature map and the key feature map are multiplied element-wise to obtain the feature weights in the frequency domain. The feature weights in the frequency domain are then transformed back to the spatial domain through an inverse discrete Fourier transform to obtain the feature weights in the spatial domain. The feature weights in the spatial domain are then normalized through a third normalization layer to obtain normalized feature weights. The normalized feature weights are then multiplied by the value feature map in the spatial domain to obtain a weighted value feature map. The weighted value feature map is then processed through a fifth convolutional layer to obtain a fourth intermediate output. The fourth intermediate output is then residually concatenated with the first input to obtain the feature map output by the spectral information modulation module. In the spatial information reconstruction module, the second input is segmented into non-overlapping window features of fixed size, and their shapes are reshaped. Each non-overlapping window feature is processed by W-MSA to obtain the attention map of each head. Finally, the attention maps of the heads are stitched together and the windows are merged along the channel dimension to obtain the feature map output by the spatial information reconstruction module. In the adaptive fusion module, the input to the adaptive fusion module is fed into the sixth convolutional layer for dimensionality reduction to obtain a low-dimensional embedding feature map. The low-dimensional embedding feature map is then fed into the seventh convolutional layer to obtain the fifth intermediate output. The low-dimensional embedding feature map is then fed into the eighth convolutional layer to obtain the sixth intermediate output. The fifth and sixth intermediate outputs are added element-wise to obtain the seventh intermediate output. The seventh intermediate output is then processed by the sigmoid activation function to obtain a gating signal. The gating signal is multiplied by the input to the adaptive fusion module to obtain the eighth intermediate output. The eighth intermediate output is then residually concatenated with the input to the adaptive fusion module to obtain the feature map output by the adaptive fusion module.

8. The remote sensing image panchromatic sharpening method based on a domain prior depth unfolding network as described in claim 7, characterized in that, In the perception enhancement module, the kernel size of the third deep convolutional layer is 3×3, the kernel size of the fourth deep convolutional layer is 3×3, the kernel size of the fifth deep convolutional layer is 3×3, the kernel size of the fifth convolutional layer is 1×1, the kernel size of the sixth convolutional layer is 3×3, the kernel size of the seventh convolutional layer is 1×5, and the kernel size of the eighth convolutional layer is 5×1.

9. The remote sensing image panchromatic sharpening method based on a domain prior deep unfolded network as described in claim 1, characterized in that, The feedforward network consists of a ninth convolutional layer, a GELU activation function, a sixth deep convolutional layer, a GELU activation function, and a tenth convolutional layer, which are cascaded together. The kernel size of the ninth convolutional layer is 1×1, the kernel size of the tenth convolutional layer is 1×1, and the kernel size of the sixth deep convolutional layer is 1×1.

10. The remote sensing image panchromatic sharpening method based on a domain prior depth unfolding network as described in claim 1, characterized in that, Before being used for actual full-color sharpening, the full-color sharpening model is trained in advance using training data generated after the Wald protocol; the loss function used in the training of the full-color sharpening model is the mean absolute error between the output image of the full-color sharpening model and the real image.