Panchromatic sharpening method based on multi-resolution panchromatic feature guidance

By adopting multi-resolution full-color feature guidance technology in the full-color sharpening method, using the spatial frequency Transformer module and self-attention mechanism, multi-scale features are extracted and fused, the problem of insufficient retention of texture and spectral details in the existing technology is solved, and high-quality full-color sharpening results are achieved.

CN120013808APending Publication Date: 2025-05-16NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 16 Cited by

Patent Information

Application Number
CN202510103622.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing full-color sharpening methods have shortcomings in the retention of texture details and spectral details, resulting in spectral distortion or loss of spatial details in the sharpening results, especially in poor results in high-frequency texture retention.

Method used

The full-color sharpening method based on multi-resolution full-color feature guidance is adopted. By designing a multi-resolution feature extraction network based on spatial frequency Transformer module, multi-scale features of full-color and multi-spectral images are extracted, and the texture injection module based on multi-head self-attention and spectral fusion module based on adaptive channel attention is optimized to ensure that the spectral information is highly consistent with the original multi-spectral image.

Benefits of technology

The spectral fidelity and spatial resolution of the fusion image are significantly improved, the problem of improper mixing of texture details and spectral details is solved, and the quality of the full-color sharpening results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013808A_ABST
    Figure CN120013808A_ABST
Patent Text Reader

Abstract

The invention discloses a panchromatic sharpening method based on multi-resolution panchromatic feature guidance, which comprises the following steps: firstly, designing a multi-resolution feature extraction network based on a spatial frequency Transform module, and respectively extracting multi-scale features from panchromatic and up-sampled multispectral images; secondly, in order to restrain and optimize the extracted detail features, a multi-head self-attention-based texture injection module is adopted to extract a dependency relationship between panchromatic and multispectral features, and more similar detail features are guided to be injected; thirdly, in order to solve information loss caused by multispectral image up-sampling, multi-level detail features are injected into a multispectral reconstruction network, the detail features are learned step by step and reconstructed to a panchromatic image size, and a spectrum fusion module based on adaptive channel attention is utilized to reconstruct a fused image in a high-quality mode; and finally, optimizing the model performance by combining reconstruction loss based on L1 norm, structural similarity constraint and transmission perception loss in training. According to the method, the spectral fidelity and the spatial resolution of the fused image are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and in particular relates to a method for panchromatic sharpening of remote sensing images. Background Art

[0002] Pansharpening is a remote sensing image processing method that fuses a high spatial resolution panchromatic image with a low spatial resolution multispectral image to generate an image with both high spatial resolution and rich spectral information. Panchromatic images are usually acquired by single-band sensors, with high spatial resolution but limited spectral information; multispectral images are acquired by multiple-band sensors, providing rich spectral information such as red, green, blue, and near-infrared, but with low spatial resolution. The core goal of pansharpening is to combine the high resolution of panchromatic images with the spectral characteristics of multispectral images to generate images with high spatial resolution and rich spectral information. This technology is widely used in remote sensing image processing, environmental monitoring, land use change analysis, agricultural monitoring and other fields, and can obtain more accurate ground information and improve the usability and analysis accuracy of images.

[0003] Traditional full-color sharpening methods mainly rely on fusion algorithms in the spatial domain or transform domain. The spatial domain method usually achieves image fusion by directly operating at the pixel level, while the transform domain method extracts image features from the frequency domain by processing the image through Fourier transform, wavelet transform, etc. However, these traditional methods often fail to effectively extract detailed information in the image, especially in terms of retaining texture details and spectral details. This may cause spectral distortion or loss of spatial details in the sharpening results, especially in terms of high-frequency texture retention. With the rapid development of deep learning technology, full-color sharpening methods based on deep learning have gradually become a hot topic of research. With the powerful feature extraction ability of deep learning, this type of method can better explore the potential features of the image and achieve more accurate image fusion.

[0004] In this context, a large number of deep learning-based panchromatic sharpening frameworks have begun to emerge. These methods enhance the spatial resolution of images by directly injecting detail information from panchromatic images into low-resolution multispectral images. However, although these methods have made some progress, they have ignored the effective constraints and guidance of texture details and spectral details. Usually, these methods only inject the extracted detail information into low-resolution multispectral images for reconstruction by superposition or concatenation. This simple processing method may cause irrelevant features to be incorrectly embedded in the final multispectral image, which in turn causes improper mixing of texture and spectral details, and ultimately causes obvious spectral distortion and spatial distortion in the panchromatic sharpened multispectral image.

[0005] In addition, panchromatic images and multispectral images often have different spatial resolutions. Panchromatic images have high spatial resolution but lack rich spectral information; multispectral images have rich spectral information but low spatial resolution. In the field of remote sensing image panchromatic sharpening, how to effectively utilize the different resolution features of these two types of images to improve the quality of the final image has become a key issue that needs to be solved. However, most panchromatic sharpening methods only use simple upsampling to match the spatial resolution of multispectral images with panchromatic images. Although this method can improve the spatial resolution of the image, it fails to fully utilize the advantages of different resolution features, ignores the complementary relationship between high-resolution and low-resolution features, and often leads to the loss of spatial details, especially in the reconstruction of high-frequency textures. Summary of the invention

[0006] In order to overcome the shortcomings of the prior art, the present invention provides a panchromatic sharpening method based on multi-resolution panchromatic feature guidance. First, a multi-resolution feature extraction network based on the spatial frequency Transformer module is designed to extract multi-scale features from panchromatic and upsampled multispectral images respectively; secondly, in order to constrain and optimize the extracted detail features, a texture injection module based on multi-head self-attention is used to extract the dependency between panchromatic and multispectral features, and guide the injection of more similar detail features; thirdly, in order to solve the information loss caused by upsampling of multispectral images, multi-level detail features are injected into the multispectral reconstruction network, and the detail features are gradually learned and reconstructed to the size of the panchromatic image, and the spectral fusion module based on adaptive channel attention is used to reconstruct the fused image with high quality; finally, the reconstruction loss based on the L1 norm, the structural similarity constraint and the transmission perception loss are combined in the training to optimize the model performance. The present invention significantly improves the spectral fidelity and spatial resolution of the fused image.

[0007] The technical solution adopted by the present invention to solve the technical problem comprises the following steps:

[0008] Step 1: Dataset preparation;

[0009] The data used in the present invention comes from two satellite sensors: Gaofen-2 (GF2) and WorldView-3 (WV3); according to the Wald protocol, the data collected by the two satellites are preprocessed to construct a down-resolution data set; then, a training set, a validation set and a down-resolution test set are divided from the generated panchromatic image blocks and multispectral image blocks; in addition, the test set also includes full-resolution images, which are obtained from the original data through data cutting;

[0010] Step 2: Construct a panchromatic sharpening model guided by multi-resolution panchromatic features;

[0011] The network model proposed in the present invention adopts a bidirectional input network structure, and its specific construction process is as follows:

[0012] Step 2-1: Construct a feature extraction module based on spatial frequency domain Transformer;

[0013] Based on the Spatial Frequency Transformer Module (SFTM), different types of features are extracted through two complementary branches. One branch extracts local detail features based on the Window-based Multi-head Self-Attention (W-MSA), and the other branch obtains global information based on the Frequency Domain Feature Extraction Module (FDFEM). The aggregated features are then input into the Multi-Scale Feedforward Gated Network (MSFGN). MSFGN expands the feature channel through two 1×1 convolutions and divides the input features into two branches for processing. At the same time, a gating mechanism is introduced to dynamically adjust the weights of different features by calculating the point-by-point product of the feature elements in the two branches, thereby enhancing the nonlinear transformation capability.

[0014] Step 2-2: Construct a texture injection module based on multi-head self-attention;

[0015] The Multi-head Self-Attention Texture Injection Module (MSATIM) is designed to identify spectrally similar and clearer texture detail features from the features extracted from the panchromatic image, and inject these features into the multispectral reconstruction network to further guide the sharpening of the multispectral image. First, the obtained multi-resolution spectral features and panchromatic features are input into the module, where the features extracted from the multispectral image are used as Q values, and the features extracted from the panchromatic image are used as K values ​​and V values. The fully connected layer is used to perform feature mapping on Q, K, and V respectively to generate N vector descriptors q i , k i and ν i ; Secondly, calculate q by dot product i and k i The feature cross-correlation matrix C between i , and with ν iMultiply them to get a single-headed attention, which is used to measure the similarity of two vectors in direction and intensity. Again, multiple single-headed self-attentions are connected in series in the channel dimension, and the dimension is restored through a fully connected layer and added to the input V to prevent information loss during the learning process, and the high-level texture detail feature T is obtained. Finally, the spectral feature F extracted by the multi-spectral reconstruction network is m It is concatenated with T in the channel dimension and fused with the spectral texture features through a 3×3 convolution layer to generate the residual component that needs to be injected into the multi-spectral reconstruction network.

[0016] Step 2-3: Construct a spectral fusion module based on adaptive channel attention;

[0017] The spectral fusion module based on adaptive channel attention (ACSA) performs correlation modeling on the channel dimension of the output feature map and dynamically adjusts the weights of each channel feature, thereby achieving more complete feature fusion on the spectral dimension. In the ACSA module, the input feature is firstly subjected to a 3×3 convolution layer to perform a preliminary linear transformation on its channel dimension to obtain the feature. Secondly, the feature is respectively input into the channel self-attention branch (Channel Self-Attention, CSA) and the spatial convolution branch (Spatial Convolution, SConv) to extract the spectral channel feature and spatial feature. Thirdly, the spatial interaction module (Spatial-Interaction, SI) and the channel interaction module (Channel-Interaction, CI) are used to realize the interaction between the spectral channel feature and the spatial feature, and dynamically adjust the feature weights of the two branches. Finally, the feature results of the two branches are added and fused through a 3×3 convolution layer to obtain the final fusion feature.

[0018] Step 2-4: Build the overall network structure;

[0019] First, the multispectral image MS is upsampled by 4 times to obtain M ↑4 , to achieve the same resolution (h, w, c) as the full-color image; then the full-color image and M ↑4 Input into the multi-resolution feature extraction network based on the spatial frequency Transformer module to extract multi-resolution texture features and spectral features;

[0020] Secondly, the features of various resolutions are input into the texture injection module based on multi-head self-attention respectively. By calculating the similarity between the texture features extracted from the panchromatic image and the spectral features extracted from the multispectral image, the texture features that best match the spectral features of the multispectral image are identified. At the same time, the features obtained from the multispectral reconstruction network are input into the texture injection module and fused with the generated texture features to obtain finer detail information and ensure that the spectral information is highly consistent with the original multispectral image.

[0021] Finally, the obtained detail features of multiple resolution scales are injected into the reconstruction network in the form of residual connection, and the detail features are fused step by step through the convolution residual module; the size of the feature map at each level is adjusted through PixelShuffle, and spliced ​​in the channel dimension through jump connection; the spectral fusion module based on adaptive channel attention is used to model the dependency relationship between feature channels, further enhance the expression of spectral and detail information of the features, and thus generate a high-quality full-color sharpened result image;

[0022] Step 3: Design loss function;

[0023] The loss function consists of three parts: reconstruction loss function based on L1 norm, structural similarity constraint between fused image and reference image, and transmission perception loss;

[0024] The reconstruction loss function based on the L1 norm is as follows:

[0025]

[0026] In the formula, x ref represents the reference image, x represents the generated full-color sharpened image, and N represents the total number of pixels in the image;

[0027] The loss function of the structural similarity constraint between the fused image and the reference image is as follows:

[0028] L SSIM =1-SSIM(x ref ,x) (2)

[0029] In the formula, SSIM(x ref , x) represents the structural similarity between the reference image and the fused image calculated using the structural similarity (SSIM) constraint;

[0030] The transmission perception loss function is shown as follows:

[0031]

[0032] Where, T srepresents the detail features generated by the texture injection module based on multi-head self-attention; f T (x) s Indicates that x is input into a multi-resolution feature network based on spatial frequency Transformer;

[0033] Finally, the overall loss function used to train the network is defined as follows:

[0034] L=λ rec L rec +λ SSIM L SSIM +λ transfer L transfer (4)

[0035] In the formula, λ rec , SSIM and λ transfer Corresponding to L rec , L SSIM and L transfer The weight coefficient of the loss function is used to balance the contribution among the three loss functions;

[0036] Step 4: Train the network model;

[0037] Set the training parameters: batch size, learning rate, weight decay, maximum number of iterations, and optimizer parameters; use the dataset preprocessed in step 1 and train using the loss function in step 3 to obtain the final panchromatic sharpening model guided by multi-resolution panchromatic features;

[0038] Step 5: Test the network model;

[0039] The panchromatic image and multispectral image to be tested are input into the final panchromatic sharpening model guided by multi-resolution panchromatic features, and a fused image is output.

[0040] Preferably, the processed data sets are all panchromatic image blocks with a resolution of 256×256 and multispectral image blocks with a resolution of 64×64; wherein, the training set, validation set and reduced-resolution test set images are obtained by Wald protocol processing, and the full-resolution test set images are obtained by data segmentation; the test set data volume is 20 groups of reduced-resolution and full-resolution data sets, and the data volume ratio of the training set and validation set except the reduced-resolution test set is 9:1.

[0041] Preferably, the weight coefficient of the setting loss is: rec is 1, λ SSIM is 0.5, λ transfer is 0.05.

[0042] Preferably, the training parameters are set as follows: the model is trained for 200 cycles using the Adam optimizer with a batch size of 2; the initial learning rate of the optimizer is 0.0001, and the learning rate is decayed to 0.5 times the original after every 20 training cycles until the minimum learning rate reaches 0.000001 and stops decaying.

[0043] The beneficial effects of the present invention are as follows:

[0044] 1. The present invention proposes a multi-resolution feature extraction network based on the spatial frequency domain Transformer module (SFTM). SFTM can capture feature information of different levels in the spatial domain and frequency domain simultaneously by combining the local detail features extracted based on the window self-attention mechanism and the global information extracted based on the frequency domain feature extraction module, thus achieving more accurate and rich feature representation.

[0045] 2. The present invention proposes a multi-head self-attention-based texture injection module (MSATIM) to extract the cross-feature spatial dependencies between panchromatic features and multispectral features, and guide the network model to inject detail features that are more similar to the spectral information of the multispectral image, thereby solving the problem of improper mixing of texture details and spectral details caused by the lack of constraint guidance in the panchromatic sharpening method based on detail injection.

[0046] 3. The present invention proposes a spectral fusion module based on adaptive channel attention (ACSA), which performs correlation modeling in the channel dimension and adaptively adjusts the fusion weights through a set of spatial information generated by convolution, ensuring high-quality reconstruction of the fused image.

[0047] 4. In addition to using the reconstruction loss function based on the L1 norm and the structural similarity constraint, the training of the model in the present invention also designs a transfer-aware loss function, so that high-resolution texture features can be more accurately injected into the multispectral image, thereby enhancing the overall effect of panchromatic sharpening. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 The figure is a flow chart of the method of the present invention.

[0049] Figure 2 FIG. 1 is a schematic diagram of the structure of a feature extraction module (SFTM) based on a spatial frequency domain Transformer in the method of the present invention, wherein Figure 2 (a) is a schematic diagram of the overall structure of SFTM. Figure 2 (b) is a schematic diagram of the structure of the multi-scale feedforward gating network (MSFGN) in SFTM. Figure 2 (c) is the structural diagram of the frequency domain feature extraction module (FDFEM) in SFTM.

[0050] Figure 3Schematic diagram of the structure of the texture injection module based on multi-head self-attention in the method of the present invention.

[0051] Figure 4 Schematic diagram of the structure of the spectral fusion module (ACSA) based on adaptive channel attention in the method of the present invention, wherein Figure 4 (a) is a schematic diagram of the overall structure of ACSA. Figure 4 (b) is a schematic diagram of the structure of the spatial interaction module (SI) in ACSA. Figure 4 (c) is a schematic diagram of the structure of the channel interaction module (CI) in ACSA.

[0052] Figure 5 It is a schematic diagram of the overall structure of the panchromatic sharpening model guided by multi-resolution panchromatic features in the method of the present invention.

[0053] Figure 6 : is a comparison diagram of the reduced-resolution WV3 test image and the fusion results of all algorithms in the embodiment of the present invention, where Figure 6 (a) is the WV3 multispectral test image, Figure 6 (b) is the WV3 full-color test image. Figure 6 (c) is the reference image, Figure 6 (d) to Figure 6 (o) The fusion results of BT-H, BDSD-PC, MTF-GLP-FS, TV, PNN, DiCNN, BDPN, GPPNN, DR-Net, DCPNet, PAPS and the method of the present invention on the reduced-resolution WV3 test image are shown in sequence.

[0054] Figure 7 : is a comparison diagram of the reduced-resolution GF2 test image and the fusion results of all algorithms in the embodiment of the present invention, wherein Figure 7 (a) is the GF2 multispectral test image, Figure 7 (b) is the GF2 full-color test image, Figure 7 (c) is the reference image, Figure 7 (d) to Figure 7 (o) The fusion results of BT-H, BDSD-PC, MTF-GLP-FS, TV, PNN, DiCNN, BDPN, GPPNN, DR-Net, DCPNet, PAPS and the method of the present invention on the reduced-resolution GF2 test image are shown in sequence.

[0055] Figure 8 : is a comparison diagram of the full-resolution WV3 test image and the fusion results of all algorithms in the embodiment of the present invention, where Figure 8 (a) is the WV3 multispectral test image, Figure 8 (b) is the WV3 full-color test image, from Figure 8(c) to Figure 8 (n) The fusion results of BT-H, BDSD-PC, MTF-GLP-FS, TV, PNN, DiCNN, BDPN, GPPNN, DR-Net, DCPNet, PAPS and the method of the present invention on the full-resolution WV3 test image.

[0056] Fig. 9 : is a comparison diagram of the full-resolution GF2 test image and the fusion results of all algorithms in the embodiment of the present invention, wherein Fig. 9 (a) is the GF2 multispectral test image, Fig. 9 (b) is the GF2 full-color test image, from Fig. 9 (c) to Fig. 9 (n) The fusion results of BT-H, BDSD-PC, MTF-GLP-FS, TV, PNN, DiCNN, BDPN, GPPNN, DR-Net, DCPNet, PAPS and the method of the present invention on the full-resolution GF2 test image are shown in sequence. DETAILED DESCRIPTION

[0057] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0058] like Figure 1 As shown, a panchromatic sharpening method based on multi-resolution panchromatic feature guidance includes the following steps:

[0059] Step 1: Dataset preparation;

[0060] The data comes from two satellite sensors, Gaofen-2 (GF2) and WorldView-3 (WV3). Specifically, the spatial resolution of the panchromatic image in the GF2 dataset is 0.8 meters, and the resolution of the multispectral image is 3.2 meters, including four bands, covering red, green, blue and near-infrared. The WV3 dataset has higher resolution and spectral information. The spatial resolution of its panchromatic image reaches 0.31 meters, and the resolution of its multispectral image is 1.24 meters. The multispectral image is expanded to 8 bands, covering richer spectral information from visible light to near-infrared.

[0061] Since the real high-resolution multispectral image (i.e., the reference image of the fusion result) does not exist, the present invention constructs a training set and a test set under reduced resolution according to the Wald protocol. According to the Wald protocol, the data collected by the two satellites are preprocessed to generate panchromatic (PAN) image blocks with a resolution of 256×256 and multispectral (MS) image blocks with a resolution of 64×64, and a reduced resolution data set is constructed. Subsequently, 20 groups are randomly selected from the generated PAN and MS image blocks as the reduced resolution test set, and the remaining data are divided into a training set and a validation set at a ratio of 9:1. The training set and the validation set only contain PAN and MS image blocks with reduced resolution. In addition, the test set also includes full-resolution images, which are obtained from the original data through data cutting, specifically consisting of 20 groups of PAN images with a resolution of 256×256 and MS images with a resolution of 64×64.

[0062] Step 2: Construct a panchromatic sharpening model guided by multi-resolution panchromatic features;

[0063] The network model proposed in the present invention adopts a bidirectional input network structure, and its specific construction process is as follows:

[0064] Step 2-1: Construct a feature extraction module based on spatial frequency domain Transformer;

[0065] The overall structure of the feature extraction module (SFTM) based on the spatial frequency domain Transformer is as follows Figure 2 (a). The SFTM module extracts different types of features through two complementary branches, one of which extracts local detail features based on the window self-attention mechanism (W-MSA), and the other branch obtains global information based on the frequency domain feature extraction module (FDFEM). Among them, the local feature extraction branch is inspired by the Swin Transformer and introduces a window-based self-attention mechanism to efficiently capture local detail features. In this process, the input feature map is first divided into multiple small windows, and self-attention calculation is performed in each window to extract local features. However, due to the limitation of the local receptive field, the single window self-attention mechanism is difficult to capture global information. To this end, the present invention proposes an FDFEM module to further extract global features and enhance the global perception ability of the model.

[0066] The structure of FDFEM is as follows Figure 2(c) is shown. First, the input features are preliminarily processed through a 3×3 convolutional layer to enhance shallow local feature information. Secondly, a two-dimensional Fast Fourier Transformation (FFT) is used to transform the features from the spatial domain to the frequency domain, so that the model can focus on the global frequency distribution and periodic patterns of the features. In the frequency domain, low-frequency components are usually related to global structures or overall contours, while high-frequency components correspond to detailed features or texture changes. By combining frequency domain information, the model can enhance the perception of the overall structure, texture patterns, and global context of the image by performing convolution in the frequency domain, thereby effectively capturing global features. Thirdly, the two-dimensional Inverse Fourier Transform (IFFT) is used to transform the frequency domain features back to the spatial domain, thereby restoring the spatial feature representation that is easy to intuitively process. Through this operation, the model effectively maps the global features extracted in the frequency domain back to the spatial domain, which is consistent with the original spatial representation, facilitating further feature fusion operations. Finally, the features extracted after Fourier transform are residually connected with the input features of the FDFEM module, and the features are transformed through a 1×1 convolutional layer to output the final global spatial domain features.

[0067] In order to enhance the expressive power of the model and the nonlinear transformation of features, a multi-scale feedforward gating network (MSFGN) is proposed, whose structure is as follows Figure 2 As shown in (b), MSFGN expands the feature channel through two 1×1 convolutions and divides the input features into two branches for processing. At the same time, a gating mechanism is introduced to dynamically adjust the weights of different features by calculating the point-by-point product of the feature elements in the two branches, thereby enhancing the nonlinear transformation capability. The introduction of the gating mechanism can not only flexibly capture the dependencies between multi-scale features, but also suppress the interference of invalid features. In these two branches, one branch uses depthwise separable convolution for feature extraction, and the other branch uses parallel 3×3 dilated convolutions (with dilation rates of 2 and 3, respectively) for multi-scale feature extraction.

[0068] Step 2-2: Construct a texture injection module based on multi-head self-attention;

[0069] The multi-head self-attention-based texture injection module (MSATIM) aims to identify spectrally similar and clearer texture detail features from the features extracted from the full-color image, and inject these features into the multispectral reconstruction network to further guide the sharpening of the multispectral image. Its structure is as follows Figure 3 The obtained multi-resolution spectral features and panchromatic features are input into the module, wherein the features extracted from the multispectral image are used as Q values, and the features extracted from the panchromatic image are used as K values ​​and V values.

[0070] First, Transformed into a two-dimensional tensor in represents the set of real numbers, c q 、c k and c v represents the number of channels of Q, K and V respectively. In the present invention, their values ​​are equal. h and w represent the height and width of the feature map respectively. Then, the fully connected layer is used to transform Mapped to N subspaces, the generated N vector descriptors are represented as Where i∈{1, 2, ..., N}, β represents the dimensionality reduction rate.

[0071] Secondly, q is calculated by dot product i and k i The feature cross-correlation matrix C between i , and with v i Multiply them together to get the single-head attention. The calculated C i The larger the value is, the more similar the features extracted from the panchromatic image are to those extracted from the multispectral image, indicating that the texture details of the current panchromatic features are highly consistent with the multispectral features, and it is suitable to inject these details into the multispectral features to improve the fusion effect. The specific formula is expressed as follows:

[0072]

[0073] A i =MatMul(C i ,v i ) (7)

[0074] In the formula, MatMul(·) is matrix multiplication, and SoftMax(·) represents the SoftMax activation function. i Represented as q i or k i The dimension of represents the obtained standardized feature cross-correlation matrix, Represents the obtained single-head attention.

[0075] Again, multiple single-head self-attentions are connected in series in the channel dimension, the dimension is restored through a fully connected layer, and added to the input V to prevent information loss during the learning process, and the required detail features T are obtained. The following formula:

[0076] T=f Linear (Concat(A1,A2,…A N ))+V (8)

[0077] In the formula, fLinear (·) represents a fully connected layer, and Concat(·) is used to concatenate feature maps in the channel dimension.

[0078] Finally, the spectral features F extracted by the multi-spectral reconstruction network are m It is concatenated with T in the channel dimension and fused with the spectral texture features through a 3×3 convolution layer to generate the residual component that needs to be injected into the multi-spectral reconstruction network. The specific calculation process is as follows:

[0079]

[0080] In the formula, Conv 3×3 (·) represents a 3×3 convolutional layer with a stride of 1. Through the texture injection module based on multi-head self-attention, multi-scale texture detail information is extracted and injected into the multispectral reconstruction network step by step, thereby effectively guiding the fusion and reconstruction process of multispectral images.

[0081] Step 2-3: Construct a spectral fusion module based on adaptive channel attention;

[0082] The spectral fusion module based on adaptive channel attention (ACSA) dynamically adjusts the weight of each channel feature by modeling the dependency between channels, thereby achieving more accurate feature fusion in the spectral dimension. Its structure is as follows: Figure 4 (a) shown.

[0083] In the ACSA module, the input feature D is first passed through a 3×3 convolutional layer to perform a preliminary linear transformation on its channel dimension to obtain the feature Secondly, the characteristics The spectral channel features and spatial features are respectively input into the channel self-attention branch (CSA) and the spatial convolution branch (SConv) to extract the spectral channel features and spatial features. Again, the spectral channel features and spatial features are interacted through the spatial interaction module (SI) and the channel interaction module (CI), and the feature weights of the two branches are dynamically adjusted. Finally, the feature results of the two branches are added and fused through a 3×3 convolution layer to obtain the final fusion feature.

[0084] In the SConv branch, two depth-wise convolutions are used to map the feature maps. Process, extract spatial detail features, and generate spatial feature map D S In the CSA branch, the feature map is firstly transformed through three convolutional layers with a stride of 2. Transform and compress the spatial information of the feature map. Secondly, flatten each channel of the compressed feature map into a vector to form a channel descriptor sequence Each vector represents The global information of the corresponding channel in . Again, the channel descriptor sequence is mapped to the query matrix Q through a linear layer C and key matrix K C In order to preserve the fine-grained detail information, the original feature map Flatten the channels and directly generate the value matrix V C Finally, Q is calculated through the dot product attention mechanism C , K C and V C The correlation between the channel features and the C .

[0085] In order to better interact the information of the two branches and realize the effective fusion of channel features and spatial features, the ACSA module uses the SI module and the CI module for interactive feature learning. Figure 4 As shown in (b), the CI module is Figure 4 (c) In the SI module, the spatial feature map D S Input it, aggregate the channel dimensions, effectively capture the global correlation of spatial detail features, and generate a spatial attention feature map The SI module uses two 1×1 convolutions to reduce the channel dimension and compress the multi-channel features into a single-channel attention map feature map S MAP In the CI module, the channel feature map D C Input it, refine the spatial information of each channel dimension, and generate a channel attention feature map The CI module first compresses the spatial features of each channel into a scalar through global average pooling to extract the global semantic information of the channel. Secondly, two 1×1 convolutions are used for processing. The first convolution layer reduces the channel dimension from c to a lower dimension c / r (where r is a hyperparameter, which is set to 8 in the present invention) to reduce computational costs and achieve feature compression. The second convolution layer restores the reduced features to the original channel dimension c. The GELU activation function is combined between the two convolutional layers to enhance the ability to express nonlinear features. Finally, the channel weights generated by the CI module are mapped to the (0, 1) range through the Sigmoid activation function, which is used to dynamically adjust the importance of each channel. The specific calculation process is as follows:

[0086] S MAP =φ(Conv 1×1 (σ(Conv 1×1 (D S )))) (10)

[0087] C MAP =φ(Conv 1×1 (σ(Conv 1×1 (HGP (D C ))))) (11)

[0088] In the formula, Conv 1×1 (·) represents a 1×1 convolutional layer, σ(·) represents the GELU activation function, φ(·) represents the Sigmoid activation function, and H GP (·) represents the global average pooling operation.

[0089] By taking the spatial attention feature map S MAP With channel feature map D C , dynamically adjust the weights of features in the spatial dimension to reflect their importance. At the same time, the channel attention feature map C MAP With spatial feature map D S Multiply them to perform weighted optimization on the features in the channel dimension. After weighted interaction, the features of these two branches are fused by element-by-element addition to generate the final output feature map Y. The specific process is shown in the following formula:

[0090]

[0091] Where ⊙ represents element-wise multiplication, Represents element-wise addition.

[0092] Step 2-4: Build the overall network structure;

[0093] The overall network model of the present invention is a bidirectional input network structure, which mainly consists of three parts: a multi-spectral reconstruction network, a multi-resolution feature extraction network based on a spatial frequency Transformer module, and a texture injection module based on multi-head self-attention. Figure 5 For the convenience of explanation, let the size of the panchromatic image PAN be (h, w, 1), and the size of the multispectral image MS be (h / r, w / r, c), where r represents the difference in spatial resolution between the panchromatic image and the multispectral image.

[0094] First, the multispectral image MS is upsampled by 4 times to obtain M ↑4 , to achieve the same resolution as the full-color image PAN. Then, PAN and M ↑4 The input is sent to a multi-resolution feature extraction network based on the spatial frequency Transformer module to extract multi-resolution texture features and spectral features. The network mainly consists of stacking multiple SFTMs and PatchMerging. Among them, PatchMerging is used to reduce the spatial resolution of features. Compared with conventional downsampling methods (such as maximum pooling), PatchMerging retains contextual information and long-range dependencies by merging image block features, reducing information loss. The specific process is shown in the following formula:

[0095] [F M1 ,F M2 ,F M3 ] = f SFTM (M ↑4 ) (13)

[0096] [F P1 ,F P2 ,F P3 ] = f SFTM (P) (14)

[0097] Where M ↑4 represents the upsampled multispectral image, P represents the panchromatic image, and f SFTM (·) represents a multi-resolution feature extraction network based on the spatial frequency Transformer module. M1 、F M1 and F M3 represent the spectral features of size (h, w, c), (h / 2, w / 2, c), and (h / 4, w / 4, c) respectively. Similarly, F P1 、F P2 and F P3 They represent full-color features of size (h, w, c), (h / 2, w / 2, c), and (h / 4, w / 4, c), respectively.

[0098] Secondly, the features of various resolutions are input into the texture injection module based on multi-head self-attention. Among them, the features extracted from the multispectral image are used as Query (Q), and the features extracted from the panchromatic image are used as Key (K) and Value (V). By calculating the similarity between the texture features extracted from the panchromatic image and the spectral features extracted from the multispectral image, the texture features that best match the spectral features of the multispectral image are identified. At the same time, the features obtained in the multispectral reconstruction network are input into the module and fused with the generated texture features to obtain finer detail information and ensure that the spectral information is highly consistent with the original multispectral image. The specific workflow can be described by the following formula:

[0099]

[0100] In the formula, f MSATIM (·) represents the texture injection module based on multi-head self-attention, F M1 、F M2 and F M3 is the spectral characteristics of different resolutions obtained above, F P1 、F P2 and F P3 is the full color feature of different resolutions, F m1 、F m2 and F m3Features extracted step by step in the multi-spectral reconstruction network. It is the fusion feature obtained through the texture injection module based on multi-head self-attention.

[0101] Finally, the panchromatic sharpened image from coarse resolution to fine resolution is gradually reconstructed in the multispectral reconstruction network. Feedback is made to the multi-spectral reconstruction network in the form of residual connection, and the detail information is further extracted through the convolution residual module. Between feature maps of different resolutions, the spatial resolution is gradually improved through PixelShuffle upsampling to restore more high-frequency information. At the end of the network, the feature maps of each scale are adjusted to the same size through jump connections, spliced ​​in the channel dimension, and the inter-channel relationship is modeled through the spectral fusion module based on adaptive channel attention to further enhance the detail information. Finally, the full-color sharpening result image is generated through the convolution residual module. The specific workflow is shown in the following formula:

[0102]

[0103] X=f RBs (f ACSA (Concat(D3,Up2(D2),Up4(D1)))) (21)

[0104] In the formula, f RBs (·) represents the convolution residual module, D1, D2 and D3 are feature maps of different resolutions generated by the multi-spectral reconstruction network according to the guidance of the injected features. Up2(·) and Up4(·) represent upsampling by 2 and 4 times using PixelShuffle, respectively. Concat(·) is the operation of concatenating vectors in the channel dimension. ACSA (·) represents the spectral fusion module of adaptive channel attention, and X is the final pan-sharpened result image.

[0105] Step 3: Design loss function;

[0106] The loss function consists of three parts: the reconstruction loss function based on the L1 norm, the structural similarity constraint between the fused image and the reference image, and the transmission perception loss. The reconstruction loss function based on the L1 norm is shown as follows:

[0107]

[0108] In the formula, x ref represents the reference image, x represents the fused image, and N represents the total number of pixels in the image.

[0109] The loss function of the structural similarity constraint between the fused image and the reference image is as follows:

[0110] L SSIM =1-SSIM(x ref ,x) (2)

[0111] In the formula, SSIM(x ref , x) represents the structural similarity between the reference image and the fused image calculated using the structural similarity (SSIM) constraint. The higher the similarity, the corresponding L SSIM The smaller the value.

[0112] The transmission perception loss function is shown as follows:

[0113]

[0114] Where, T s represents the detail features generated by the texture injection module based on multi-head self-attention, where s represents the spatial scale, which is 1 times, 2 times, and 4 times the original multispectral image size. T (x) s Indicates that the fusion result x is input into the multi-resolution feature network based on spatial frequency Transformer and the texture injection module based on multi-head self-attention, and the generated multiple resolution feature maps are respectively s Calculate the L1 norm and get the final transmission perception loss L transfer Finally, the overall loss function for training the panchromatic sharpening model guided by multi-resolution panchromatic features is defined as follows:

[0115] L=λ rec L rec +λ SSIM L SSIM +λ transfer L transfer (4)

[0116] In the formula, λ rec , SSIM and λ transfer Corresponding to L rec , L SSIM and L transfer The weight coefficient of the loss function is used to balance the contribution of the three loss functions. The present invention sets λ rec is 1, λ SSIM is 0.5, λ transfer is 0.05.

[0117] Step 4: Train the network model;

[0118] Set the training parameters: batch size, learning rate, weight decay, maximum number of iterations, and optimizer parameters; use the dataset preprocessed in step 1 and train using the loss function in step 3 to obtain the final panchromatic sharpening model guided by multi-resolution panchromatic features.

[0119] The model was trained with the Adam optimizer for 200 cycles, and the batch size was set to 2. The initial learning rate of the optimizer was set to 0.0001, and the learning rate was decayed to 0.5 times the original value after every 20 cycles of training until the minimum learning rate reached 0.000001 and stopped decaying.

[0120] Step 5: Test the network model;

[0121] The panchromatic image and multispectral image to be tested are input into the final panchromatic sharpening model guided by multi-resolution panchromatic features, and a fused image is output. Specific embodiment:

[0123] 1. Experimental conditions

[0124] The experimental environment is Intel(R) Core(TM) i5-13600KF CPU@5.10GHz, with 32GB of memory, the GPU processor is NVIDIA GEFORCE RTX 3090, the Python version used is 3.9.18, and it is programmed based on the PyTorch deep learning framework.

[0125] 2. Experimental content

[0126] In order to verify the effectiveness of the present invention, the method of the present invention is compared with several representative algorithms in the field of full color sharpening. These algorithms include traditional algorithms and algorithms based on deep learning. Traditional algorithms include BT-H and BDSD-PC based on component replacement, MTF-GLP-FS based on multi-resolution analysis, and TV based on variational optimization. Algorithms based on deep learning include PNN, DiCNN, BDPN, GPPNN, DR-Net, DCPNet and PAPS. After the network is trained on the reduced resolution training set, it is tested and compared on the reduced resolution and full resolution test sets respectively.

[0127] 3. Evaluation indicators

[0128] For the resolution reduction experiment, the present invention uses multiple quantitative indicators to evaluate the fusion results, including SAM, ERGAS, SCC, Q2n and PSNR, to evaluate the algorithm performance from multiple aspects such as spectral similarity, spatial similarity and comprehensive quality. Among them, SAM evaluates the spectral similarity by measuring the spectral angle between the fusion result and the reference image. The smaller the value, the higher the spectral similarity. ERGAS is used to measure the overall similarity between the fusion result and the reference image. The smaller the value, the closer the fusion result is to the reference image. SCC quantifies the spatial dependence or similarity between the fusion result and the reference image, and evaluates the spatial quality of the image. The higher the value, the better the spatial quality. Q2n (Q4 is used when MS is 4 bands, and Q8 is used when MS is 8 bands) calculates the quality score of multi-band or multi-channel images, evaluates the spectral and spatial fidelity between the fusion result and the reference image, and the higher the value, the higher the spectral and spatial fidelity of the result. PSNR measures the signal-to-noise ratio. The higher the value, the better the reconstruction quality of the image.

[0129] For full-resolution experiments, the spectral distortion factor D is used. λ , spatial distortion coefficient D S and no-reference quality rating QNR as the evaluation index. λ Indicates the degree of spectral distortion between the experimental results and the input low-resolution multispectral image. The smaller the value, the higher the spectral similarity. S It indicates the degree of spatial distortion between the experimental results and the input full-color image. The smaller the value, the higher the spatial similarity. QNR measures the spectral and spatial distortion between the experimental results and the input image. The higher the value, the better the generated image performs in retaining spectral and spatial information.

[0130] 4. Simulation test

[0131] Figure 6The fusion results of all algorithms on the down-resolution WV3 dataset are shown. In order to compare the performance of these algorithms more intuitively, the box-marked area is enlarged and placed in the lower right corner of the image to provide a more detailed visual comparison. It is obvious from the visual comparison that the high-resolution multispectral images generated by the traditional algorithms are not as good as the deep learning-based methods in retaining spatial and spectral details. Especially in the vegetation area in the figure, the BDSD-PC, MTF-GLP-FS and TV methods show excessive enhancement of spatial details. At the same time, these methods also have different degrees of spectral distortion, and their overall hue deviates greatly from the reference image. The deep learning-based method more effectively retains spectral and spatial information, but it can be seen from the enlarged area that PNN, BDPN and GPPNN have insufficient ability to retain spectral detail information and lose the color information on the building. Although DiCNN, DR-Net, DCPNet and PAPS retain part of the spectral information, they are still insufficient compared with the reference image. The method of the present invention performs well in retaining local spectral details, with clear edges, showing superior spatial detail quality.

[0132] The quantitative performance of each fusion algorithm on the down-resolution WV3 dataset is shown in Table 1. Each algorithm is tested using 20 pairs of multispectral and panchromatic test images, and the average value of each indicator is recorded. The optimal result of each indicator is indicated in bold, and the suboptimal result is indicated by underline. Since the deep learning-based methods have stronger nonlinear fitting capabilities and can better learn image features, they achieve better results in the evaluation indicators. The method of the present invention is slightly better than other comparison algorithms in all indicators on the down-resolution WV3 dataset, which further illustrates the effectiveness of the method of the present invention on the down-resolution WV3 dataset.

[0133] Table 1 Quantitative comparison of all algorithms on the down-resolution WV3 dataset

[0134]

[0135]

[0136] The fusion results of all algorithms on the reduced-resolution GF2 dataset are shown in Figure 7As shown. Similarly, the local area of ​​interest is marked and enlarged for display. From the visual result analysis, it can be seen that the TV algorithm has obvious spectral distortion, especially the vegetation area is brighter overall. MTF-GLP-FS also has slight spectral distortion, and the other algorithms all maintain spectral information close to the reference image. Observing the enlarged area, PNN and TV are obviously inferior to other algorithms in terms of spatial details, and there is blurred edges. DR-Net, DCPNet and PAPS all retain spatial and spectral information better than other algorithms, but the method of the present invention is more prominent in detail enhancement capabilities, especially in the street in the lower left corner of the enlarged area, where the similarity with the reference image is higher.

[0137] Table 2 shows the quantitative evaluation results of all algorithms on the down-resolution GF2 dataset. Each algorithm is verified using 20 pairs of multispectral and panchromatic test images, and the average value of each indicator is recorded. The analysis results show that the traditional fusion methods generally lag behind the deep learning-based algorithms in terms of evaluation indicators. The PAPS algorithm performs better on the GF2 dataset than on the WV3 dataset, and achieves suboptimal results in all evaluation indicators. The method of the present invention has excellent performance in all indicators of the down-resolution GF2 dataset. SAM, ERGAS, Q4 and PSNR are significantly better than other algorithms, which verifies its superiority in maintaining the spectral and spatial information of the enhanced image.

[0138] Table 2 Quantitative comparison of all algorithms on the down-resolution GF2 dataset

[0139]

[0140] In order to verify the performance of the algorithm on the full-resolution real dataset, all algorithms were tested on the full-resolution WV3 and GF2 datasets in the same way as the analysis method of the down-resolution dataset, and qualitative and quantitative analysis were performed. The fusion results of all algorithms on the full-resolution WV3 dataset are shown in Figure 2. Figure 8As shown, the area marked by the square is enlarged and placed in the lower right corner of the image. It can be seen that MTF-GLP-FS and TV perform poorly in processing spatial details, which is manifested as obvious artifacts and blurred edges in the fused image, reducing the overall quality of the image. The fused images of PNN, BDPN and GPPNN have certain problems in spectral fidelity, especially in the presentation of spectral information on the roof of the building, which deviates greatly from the actual situation, indicating the shortcomings of these algorithms in spectral reconstruction. It can be seen from the enlarged area that although the detail enhancement network of PAPS brings better spatial details, some spectral effects appear in the fused image, and the spectral information is over-enhanced. The method of the present invention can well preserve spatial detail information and spectral information. As shown in the enlarged area, the edges of the building are clearer and more precise than other algorithms, indicating that the method of the present invention has significant advantages in maintaining image quality and authenticity.

[0141] The quantitative evaluation results of all algorithms on the full-resolution WV3 dataset are shown in Table 3. The data are the average values ​​obtained by testing all algorithms on 20 pairs of multispectral and panchromatic test images, where the best results are in bold and the suboptimal results are underlined. λ The PAPS algorithm shows good results in spatial detail processing through its detail enhancement network and achieves the second best D S The method of the present invention is λ The indicators did not achieve the best results, but in D S The performance of the two indicators, QNR, is relatively outstanding, which shows its strong ability in comprehensively processing spectral and spatial information and proves its effectiveness and practicability in hyperspectral image reconstruction.

[0142] Table 3 Quantitative comparison of all algorithms on the full-resolution WV3 dataset

[0143]

[0144] Fig. 9 This is the fusion result of all algorithms on the full-resolution GF2 dataset, and the area marked by the box is also enlarged for display. Fig. 9 (c) and Fig. 9(f) It can be clearly seen from the results that in the images generated by the BT-H and TV fusion algorithms, the colors of the vegetation areas are more saturated than those of other algorithms. This excessive color saturation is often a clear sign of spectral distortion. In particular, in the results of the TV algorithm, the vegetation areas are too dense and do not conform to the actual situation of the natural scene. The challenges of PNN, BDPN, GPPNN and DCPNet in terms of spectral fidelity can be found in the enlarged image area, among which the GPPNN algorithm is particularly obvious, and the tones of the buildings in its fused image are obviously inconsistent with its original image. Artifacts appear at the edges of the buildings in the enlarged area of ​​the DR-Net fusion result image, and the spatial details are not clear enough. Although PAPS retains the spectral details very well, it sacrifices the clarity of some spatial details. The method of the present invention demonstrates its excellent comprehensive performance, and performs well in both detail clarity and spectral fidelity, proving its effectiveness and superiority in panchromatic sharpening on the full-resolution GF2 dataset.

[0145] Table 4 shows the quantitative evaluation results of all algorithms on the full-resolution GF2 dataset. The data is obtained by testing all algorithms on 20 pairs of multispectral and panchromatic images. The table shows the average performance of each algorithm on different evaluation indicators. As can be seen from the table, the spectral distortion indicator D of the TV algorithm is λ The worst, which is consistent with the visual analysis. λ The performance is good in terms of indicators, but its ability to reconstruct spatial details is insufficient. S The index results are relatively poor. Although the DCPNet algorithm has the problem of spectral distortion, it is S The proposed method achieves the best results in all quantitative evaluation indicators on the full-resolution GF2 dataset.

[0146] Table 4 Quantitative comparison of all algorithms on the full-resolution GF2 dataset

[0147]

Claims

1. A panchromatic sharpening method based on multi-resolution panchromatic feature guidance, characterized in that: The following steps are involved: Step 1: Dataset preparation; The data used in the present invention comes from two satellite sensors: Gaofen-2 (GF2) and WorldView-3 (WV3); according to the Wald protocol, the data collected by the two satellites are preprocessed to construct a down-resolution data set; then, a training set, a validation set and a down-resolution test set are divided from the generated panchromatic image blocks and multispectral image blocks; in addition, the test set also includes full-resolution images, which are obtained from the original data through data cutting; Step 2: Construct a panchromatic sharpening model guided by multi-resolution panchromatic features; The network model proposed in the present invention adopts a bidirectional input network structure, and its specific construction process is as follows: Step 2-1: Construct a feature extraction module based on spatial frequency domain Transformer; Based on the Spatial Frequency Transformer Module (SFTM), different types of features are extracted through two complementary branches. One branch extracts local detail features based on the Window-based Multi-head Self-Attention (W-MSA) mechanism, and the other branch obtains global information based on the Frequency Domain Feature Extraction Module (FDFEM). The aggregated features are then input into the Multi-Scale Feedforward Gated Network (MSFGN). The MSFGN expands the feature channel through two 1×1 convolutions and divides the input features into two branches for processing. At the same time, the gating mechanism is introduced to dynamically adjust the weights of different features by calculating the point-by-point product of the feature elements in the two branches, thereby enhancing the nonlinear transformation capability. Step 2-2: Construct a texture injection module based on multi-head self-attention; The Multi-head Self-Attention Texture Injection Module (MSATIM) is designed to identify spectrally similar and clearer texture detail features from the features extracted from the panchromatic image, and inject these features into the multispectral reconstruction network to further guide the sharpening of the multispectral image. First, the obtained multi-resolution spectral features and panchromatic features are input into the module, where the features extracted from the multispectral image are used as Q values, and the features extracted from the panchromatic image are used as K values ​​and V values. The fully connected layer is used to perform feature mapping on Q, K, and V respectively to generate N vector descriptors q i , k i and v i ; Secondly, calculate q by dot product i and k i The feature cross-correlation matrix C between i , and with v i Multiply them to get a single-headed attention, which is used to measure the similarity of two vectors in direction and intensity. Again, multiple single-headed self-attentions are connected in series in the channel dimension, and the dimension is restored through a fully connected layer and added to the input V to prevent information loss during the learning process, and the high-level texture detail feature T is obtained. Finally, the spectral feature F extracted by the multi-spectral reconstruction network is m It is concatenated with T in the channel dimension and fused with the spectral texture features through a 3×3 convolution layer to generate the residual component that needs to be injected into the multi-spectral reconstruction network. Step 2-3: Construct a spectral fusion module based on adaptive channel attention; The spectral fusion module based on adaptive channel attention (ACSA) performs correlation modeling on the channel dimension of the output feature map and dynamically adjusts the weights of each channel feature, thereby achieving more complete feature fusion on the spectral dimension. In the ACSA module, the input feature is first subjected to a 3×3 convolution layer to perform a preliminary linear transformation on its channel dimension to obtain the feature. Secondly, the feature is input into the channel self-attention branch (CSA) and the spatial convolution branch (SConv) to extract the spectral channel feature and spatial feature. Thirdly, the spatial interaction module (SI) and the channel interaction module (CI) are used to realize the interaction between the spectral channel feature and the spatial feature, and the feature weights of the two branches are dynamically adjusted. Finally, the feature results of the two branches are added and fused through a 3×3 convolution layer to obtain the final fusion feature. Step 2-4: Build the overall network structure; First, the multispectral image MS is upsampled by 4 times to obtain M ↑4 , to achieve the same resolution (h, w, c) as the full-color image; then the full-color image and M ↑4 Input into the multi-resolution feature extraction network based on the spatial frequency Transformer module to extract multi-resolution texture features and spectral features; Secondly, the features of various resolutions are input into the texture injection module based on multi-head self-attention respectively. By calculating the similarity between the texture features extracted from the panchromatic image and the spectral features extracted from the multispectral image, the texture features that best match the spectral features of the multispectral image are identified. At the same time, the features obtained from the multispectral reconstruction network are input into the texture injection module and fused with the generated texture features to obtain finer detail information and ensure that the spectral information is highly consistent with the original multispectral image. Finally, the obtained detail features of multiple resolution scales are injected into the reconstruction network in the form of residual connection, and the detail features are fused step by step through the convolution residual module; the size of the feature map at each level is adjusted through PixelShuffle, and spliced ​​in the channel dimension through jump connection; the spectral fusion module based on adaptive channel attention is used to model the dependency between feature channels, further enhance the expression of the feature on the spectrum and detail information, and thus generate a high-quality full-color sharpened result image; Step 3: Design loss function; The loss function consists of three parts: the reconstruction loss function based on the L1 norm, the structural similarity constraint between the fused image and the reference image, and the transmission perception loss; the reconstruction loss function based on the L1 norm is shown as follows: In the formula, x ref represents the reference image, x represents the generated full-color sharpened image, and N represents the total number of pixels in the image; The loss function of the structural similarity constraint between the fused image and the reference image is as follows: L SSIM =1-SSIM(x ref ,x) (2) In the formula, SSIM(x ref , x) represents the structural similarity between the reference image and the fused image calculated using the structural similarity (SSIM) constraint; The transmission perception loss function is shown as follows: Where, T s represents the detail features generated by the texture injection module based on multi-head self-attention; f T (x) s Indicates inputting x into a multi-resolution feature network based on spatial frequency Transformer; Finally, the overall loss function used to train the network is defined as follows: L=λ rec L rec +λ SSIM L SSIM +λ transfer L transfer (4) In the formula, λ rec , SSIM and λ transfer Corresponding to L rec , L SSIM and L transfer The weight coefficient of the loss function is used to balance the contribution among the three loss functions; Step 4: Train the network model; Set the training parameters: batch size, learning rate, weight decay, maximum number of iterations, and optimizer parameters; use the dataset preprocessed in step 1 and train using the loss function in step 3 to obtain the final panchromatic sharpening model guided by multi-resolution panchromatic features; Step 5: Test the network model; The panchromatic image and multispectral image to be tested are input into the final panchromatic sharpening model guided by multi-resolution panchromatic features, and a fused image is output.

2. The method for panchromatic sharpening based on multi-resolution panchromatic feature guidance according to claim 1, characterized in that: The processed data sets are all panchromatic image blocks with a resolution of 256×256 and multispectral image blocks with a resolution of 64×64; among them, the training set, validation set and reduced-resolution test set images are obtained by Wald protocol processing, and the full-resolution test set images are obtained by data segmentation; the test set data volume is 20 groups of reduced resolution and full resolution, and the training set and validation set are taken from the data set excluding the reduced-resolution test set, and the data volume ratio is 9:

1.

3. The method for panchromatic sharpening based on multi-resolution panchromatic feature guidance according to claim 1, characterized in that: The weight coefficient of the loss in step 3: λ rec is 1, λ SSIM is 0.5, λ transfer is 0.

05.

4. The method for panchromatic sharpening based on multi-resolution panchromatic feature guidance according to claim 1, characterized in that: The training parameters are set as follows: the model is trained for 200 cycles using the Adam optimizer, with a batch size of 2, an initial learning rate of the optimizer of 0.0001, and the learning rate is decayed to 0.5 times the original value after every 20 cycles of training until the minimum learning rate reaches 0.000001 and stops decaying.

Citation Information

Cited By

  • Frequency domain difference image reconstruction method and system based on SwinIR model

    CN120198295A

  • Panchromatic sharpening method based on progressive expansion frame

    CN120339123A

  • Remote sensing image panchromatic sharpening method and system based on spatial spectrum transformer

    CN120451005A

  • Panchromatic sharpening image fusion method and device based on double-domain flexible converter

    CN120807318A

  • Multispectral image panchromatic sharpening method based on plug-and-play gradient feature guidance fusion

    CN120976059A