A remote sensing image pan-sharpening method based on SwinTransformer and CNN fusion

By combining SwinTransformer and CNN methods, deep and shallow features of remote sensing images are extracted and fusion processing is performed, which solves the problem that full-color sharpening is difficult to extract global and local features in the prior art, and achieves a better full-color sharpening effect for remote sensing images.

CN116739911BActive Publication Date: 2025-05-06ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310422287.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2025-05-06
Estimated Expiration
2043-04-19

AI Technical Summary

Technical Problem

The prior art is difficult to extract the global and local features of remote sensing images at the same time, resulting in poor full-color sharpening effect.

Method used

Using a fusion method based on SwinTransformer and CNN, the deep and shallow features of the image are extracted respectively through two feature extraction paths, and the features are fused through residual connection and channel stitching, and image reconstruction is finally carried out in the image reconstruction module.

Benefits of technology

While retaining global features, the focus on local features is enhanced, the performance of full-color sharpening of remote sensing images is improved, and the generated HRMS images have higher spatial and spectral resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116739911B_ABST
    Figure CN116739911B_ABST
Patent Text Reader

Abstract

The present invention discloses a remote sensing image panchromatic sharpening method based on the fusion of SwinTransformer and CNN. The present invention introduces the SwinTransformer method in the panchromatic sharpening task and combines it with the CNN method to extract the deep features and shallow features of the image respectively, and reconstructs the image after splicing the obtained features. Compared with the traditional Transformer-based method, it strengthens the focus on local features while retaining the model's ability to extract global features of the image. The local attention and shift window mechanism of SwinTransformer bring better nonlinear texture features, further improving the ability to extract local features. The innovative process of combining SwinTransformer and CNN solves the embarrassment of the current panchromatic sharpening of remote sensing images that focuses on global features but ignores local features. Experiments on the WorldView‑3 and GaoFen‑2 datasets verify that our model can improve the performance of panchromatic sharpening of remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning and computer vision technology, and in particular to a remote sensing image full color sharpening method based on the fusion of SwinTransformer and Convolutional Neural Network (CNN). Background Art

[0002] Satellite images with high spatial resolution and high spectral resolution have wide application value in land cover classification, change detection, environmental monitoring and other fields. Most remote sensing satellites can only capture multispectral (MS) images with high spectral resolution and low spatial resolution and panchromatic (PAN) images with high spatial resolution and low spectral resolution. Therefore, the purpose of the panchromatic sharpening task is to fuse MS and PAN images to generate images with both high spatial resolution and high spectral resolution (HRMS).

[0003] Traditional pan-sharpening can be divided into three categories: component substitution (CS), multi-resolution analysis (MRA), and variational optimization (VO). These methods require hand-crafted priors to regularize the solution space of HRMS images, which usually leads to spatial or spectral distortion.

[0004] In recent years, with the rapid development of deep learning technology and computer vision, deep neural networks have gradually replaced traditional methods in the application of panchromatic sharpening tasks. Existing deep neural network panchromatic sharpening methods can be roughly divided into two categories: CNN-based methods and Transformer-based methods.

[0005] As the first CNN-based method, PNN extracts the features of MS and PAN images by splicing them and inputting them into a three-layer convolutional network. Since then, most CNN-based methods have improved the feature extraction capability by stacking network layers. Although CNN-based methods can effectively extract local features of images, they always keep the same weight matrix for different image regions, while the relationship between pixels in local regions in the full color sharpening task is different, which often affects the final prediction image generation.

[0006] In recent years, Transformer-based models have used self-attention mechanisms to capture global interactions between contexts, have the ability to learn global information from images, and have achieved better results in full-color sharpening tasks. However, it is worth noting that the Transformer itself divides the input image into fixed-size blocks and processes them independently, which makes it difficult for Transformer-based methods to learn pixel-level attention and thus difficult to obtain local fine features of the image.

[0007] In the task of full-color sharpening, the global and local features of remote sensing images are equally important for image reconstruction. The CNN-based model focuses on the extraction of local features of remote sensing images, while the Transformer-based model focuses on the extraction of global features of remote sensing images.

[0008] Therefore, how to design a full-color sharpening model that can simultaneously utilize the advantages of CNN and Transformer models to extract global and local features of MS images and PAN images is a technical problem that needs to be solved urgently. Summary of the invention

[0009] The technical problem to be solved by the present invention is how to fully combine the advantages of CNN and Transformer models, reasonably extract the global information and local information of MS images and PAN images, and provide a remote sensing image full color sharpening method based on the fusion of SwinTransformer and CNN.

[0010] The specific technical solutions adopted by the present invention are as follows:

[0011] A remote sensing image panchromatic sharpening method based on the fusion of SwinTransformer and CNN, the specific method of which is as follows: a panchromatic image and a multispectral image to be panchromatic sharpened are input into a panchromatic sharpening model, and the panchromatic sharpening model includes a first feature extraction path, a second feature extraction path and an image reconstruction module, the panchromatic image and the multispectral image are respectively used as input images of the first feature extraction path and the second feature extraction path, the two input images are respectively subjected to shallow feature extraction and deep feature extraction by a shallow feature extraction module based on CNN and a deep feature extraction module based on SwinTransformer in their respective feature extraction paths, finally the deep features and shallow features extracted in the second feature extraction path are fused through residual connection and then spliced ​​with the deep features extracted in the first feature extraction path in the channel dimension, and the spliced ​​features are input into the image reconstruction module for image reconstruction;

[0012] The shallow feature extraction modules in the two feature extraction paths are both CNN convolutional networks formed by cascading a group of convolutional modules; the deep feature extraction module in the first feature extraction path is formed by cascading two groups of STB (Swin Transformer Block) modules, and the front end of each group of STB modules is equipped with a PM (PatchMerging) module, and the PM module is used to perform PatchMerging operation on the input of each group of STB modules; the deep feature extraction module in the second feature extraction path is formed by cascading two groups of STB modules; each group of STB modules in the two feature extraction paths is formed by cascading a first STB module and a second STB module; the first STB module uses the shallow feature map output by the front-end cascade module as input, and each input shallow feature map is divided into non-overlapping local window blocks according to a fixed-size partition window after layer normalization, and each local window block is encoded through a linear projection matrix shared across windows to obtain the feature vector of each local window. Then, the feature vectors of each local window are used as the query (Query), value (Value) and key (Key) of the multi-head attention mechanism in the multi-head attention layer, and the attention map is obtained through attention fusion. The attention map is residually connected with the input shallow feature map to obtain the intermediate feature map, and then the intermediate feature map is residually connected with the result of layer normalization and linear classifier to obtain the output feature of the first STB module; the second STB module uses the output feature of the first STB module of the front-end cascade as the input feature, and the second STB module adds a window shift operation before the multi-head attention layer of the first STB module; in both feature extraction paths, the output feature of the second STB module in the second group of STB modules is used as the deep feature extracted by the deep feature extraction module;

[0013] In the image reconstruction module, multiple convolutions and up-sampling are performed on the input splicing features to obtain the final panchromatic sharpening result.

[0014] Preferably, the deep feature extraction module uses the Panformer model as a baseline model, and is obtained by replacing the self-attention modules in the Panformer model with the STB module.

[0015] Preferably, the CNN convolutional network is composed of five 3*3 convolutional layers and one 1*1 convolutional layer cascaded in sequence.

[0016] Preferably, in the second STB module, when the shift operation is performed on the window divided in the first STB module, the moving distances in both the horizontal and vertical directions are half of the window size rounded down.

[0017] Preferably, the image reconstruction module is formed by sequentially cascading a first 3*3 convolutional layer, a first pixel reorganization layer, a second 3*3 convolutional layer, a second pixel reorganization layer, a third 3*3 convolutional layer and a fourth 3*3 convolutional layer.

[0018] Preferably, the pan-sharpening model is pre-trained using training data generated after the Wald protocol before being used for actual pan-sharpening, and the loss function used in the pan-sharpening model training is the mean absolute error.

[0019] Preferably, when the PM module performs the PatchMerging operation, elements are selected at intervals of two positions in the row and column directions of the input features to form new window image blocks, and then all the window image blocks are spliced ​​together as a whole tensor, and finally expanded and the channel dimension is adjusted to twice the original through a fully connected layer to form output features passed to the rear.

[0020] Preferably, in the first STB module, the size of the divided window for forming the local window block is fixed to 4×4.

[0021] Preferably, in the deep feature extraction module, the number of deep feature channels finally extracted is 64.

[0022] Preferably, in the second STB module, the multi-head attention mechanism uses a single input feature as Query, Value and Key to perform attention fusion, thereby obtaining an attention map.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] The present invention introduces the SwinTransformer method in the panchromatic sharpening task and combines it with the CNN method to extract the deep features and shallow features of the image respectively, and then reconstruct the image after splicing the obtained features. Compared with the traditional Transformer-based method, the model's ability to extract global features of the image is retained while strengthening the focus on local features. The local attention and shift window mechanism of SwinTransformer brings better nonlinear texture features, further improving the ability to extract local features. The innovative process of combining SwinTransformer and CNN solves the current defect of focusing on global features while ignoring local features in panchromatic sharpening of remote sensing images. Experiments on the WorldView-3 and GaoFen-2 datasets have verified that the model provided by the present invention can improve the performance of panchromatic sharpening of remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is the SWCPAN model structure diagram;

[0026] Figure 2 It is a schematic diagram of the STB module;

[0027] Figure 3 It is a schematic diagram of the image reconstruction module;

[0028] Figure 4 The figure is a training and testing flow chart of the SWCPAN model in an embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to make the above-mentioned purpose, features and advantages of the present invention more obvious and easy to understand, the specific implementation mode of the present invention is described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined accordingly without conflicting with each other.

[0030] In the description of the present invention, it should be understood that the terms "first" and "second" are only used to distinguish the description purpose, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features.

[0031] In the field of remote sensing, remote sensing images with both high spatial resolution and high spectral resolution have wide application value in land cover classification, change detection, environmental detection and other fields. However, remote sensing satellites can usually only obtain multispectral images (MS) with high spectral resolution and low spatial resolution, and panchromatic images (PAN) with high spatial resolution and low spectral resolution. At this time, panchromatic sharpening is needed to synthesize MS and PAN into HRMS images with both high spatial resolution and high spectral resolution. Panchromatic sharpening algorithms are generally based on deep learning methods. The current mainstream methods can be divided into CNN-based models and Transformer-based models. Among them, the CNN-based model focuses on the local features of the image, but always maintains the same weight matrix for different image regions. The relationship between pixels in the local area in the panchromatic sharpening task is different, which often affects the final predicted image generation; the Transformer-based model focuses on the global features of the image, but divides the input image into fixed-size blocks and processes them independently, which makes it difficult for the Transformer-based method to learn pixel-level attention, and thus it is difficult to obtain local fine features of the image. Therefore, it is necessary to design a better model that pays attention to both the global and local features of the image. The core of the present invention is to propose a special panchromatic sharpening method for remote sensing images that combines SwinTransformer and CNN in the panchromatic sharpening model. On the one hand, the local attention and shift window mechanism of SwinTransformer are introduced to enhance the Transformer's attention to local features. On the other hand, the local and global features of the image are extracted respectively through CNN and Transformer, and the HRMS image is reconstructed after the features are fused, thereby achieving a better panchromatic sharpening effect.

[0032] In a preferred embodiment of the present invention, a remote sensing image full color sharpening method based on SwinTransformer and CNN fusion is provided as follows:

[0033] The MS image and PAN image to be pan-sharpened are input into the pan-sharpening model in a dual-path manner according to the modality, and the pan-sharpening model includes a first feature extraction path, a second feature extraction path and an image reconstruction module. The panchromatic image and the multispectral image are respectively used as input images of the first feature extraction path and the second feature extraction path. The first feature extraction path and the second feature extraction path respectively have a shallow feature extraction module based on CNN and a deep feature extraction module based on SwinTransformer. The two input images are respectively subjected to shallow feature extraction and deep feature extraction by the shallow feature extraction module and the deep feature extraction module in their respective feature extraction paths; finally, the deep features and shallow features extracted in the second feature extraction path are fused through residual connection and then spliced ​​with the deep features extracted in the first feature extraction path in the channel dimension, and the obtained spliced ​​features are then input into the image reconstruction module for image reconstruction.

[0034] The above remote sensing image pan-sharpening method based on SwinTransformer and CNN fusion can essentially be described as a data processing process in a remote sensing image pan-sharpening model (denoted as SWCPAN). The specific structure of the SWCPAN model of the present invention is described in detail below. Figure 1 The overall structure diagram of the SWCPAN model is shown in Figure 2. The SWCPAN model consists of three modules: shallow feature extraction module, deep feature extraction module and image reconstruction module. Since MS and PAN images are different modalities, the SWCPAN model constructs a dual-path encoder to extract features, that is, the first feature extraction path and the second feature extraction path mentioned above both have shallow feature extraction modules and deep feature extraction modules, which extract features from images of the two modalities respectively.

[0035] Specifically, in the two feature extraction paths of the SWCPAN model, the shallow feature extraction modules are all CNN convolutional networks formed by cascading a group of convolutional modules. In an embodiment of the present invention, the shallow feature extraction module of the SWCPAN model is formed by cascading five groups of 3×3 convolutional layers and a group of 1×1 convolutional layers in sequence. The input of the shallow feature extraction module is mainly the original MS image M, with dimensions (B×C0×H0×W0), and the PAN image P, with dimensions (B×1×4H0×4W0). B is the input Batchsize. B depends on the sample size of each batch in the training stage, and can be set to 1 in the prediction stage. C0, H0, and W0 are the number of bands, height, and width of the image, respectively. Here, C0 can be set to 4. The MS image M and the PAN image P are respectively extracted through shallow features to obtain the corresponding shallow feature maps M0 and P0, with dimensions of (B×C1×H1×W1) and (B×C1×4H1×4W1), respectively. C1, H1, and W1 are the number of feature channels, height, and width of the shallow feature map M0. The extracted shallow features will be input into the deep feature extraction module. The deep feature extraction module is based on the Panformer model as the baseline model, and the self-attention modules in the Panformer model are replaced by STB modules.

[0036] Specifically, see Figure 1 As shown, the deep feature extraction modules in the two paths are composed of a total of 4 groups of SwinTransformer Block (STB) modules and two groups of PatchMerging modules (PM modules), each group of STB modules includes a first STB module and a second STB module, which are used to extract deep features of the image. The 4 groups of STB modules and the two groups of PM modules are respectively located in the two feature extraction paths, wherein the deep feature extraction module in the first feature extraction path is composed of two groups of STB modules and two groups of PM modules, wherein the two groups of STB modules are cascaded, and each group of STB modules has a PM module at the front end, and the PM module is used to perform PatchMerging operations on the input of each group of STB modules; and the deep feature extraction module in the second feature extraction path is directly formed by the cascade of two groups of STB modules, and each group of STB modules has no PM module at the front end.

[0037] The basic structure of the two STB modules in each group of STB modules mentioned above is the same, and the only difference is the pre-processing of the multi-head self-attention mechanism in the STB module. The basic structure of each STB module consists of two groups of layer normalization layers (LayerNorm layer, LN), a group of multi-head self-attention layers (Multi-headSelf-Attention layer, MSA) and a group of linear classifiers (Multi-Layer Perceptron, MLP), such as Figure 2As shown in the figure. Since the input shallow features M0 and P0 have different heights and widths, the model adds a PM module before the first STB module and the second STB module on the first feature extraction path. The PM module is used to perform the PatchMerging operation. Each PatchMerging module reduces the height and width of the input feature to half of the original. The first STB module uses the shallow feature map output by the front-end cascade module as input. Each input shallow feature map is divided into non-overlapping local window blocks according to a fixed-size partition window after layer normalization (in an embodiment of the present invention, the partition window size for forming local window blocks can be fixed to 4×4). Each local window block is encoded through a linear projection matrix shared across windows to obtain a feature vector for each local window. The feature vectors of each local window are then used as the query (Query), value (Value) and key (Key) of the multi-head attention mechanism to perform attention fusion to obtain an attention map. The attention map is residually connected with the input shallow feature map to obtain an intermediate feature map. The intermediate feature map is then residually connected with the intermediate feature map after layer normalization and linear classifier to obtain the final output features of the first STB module. Specifically, the first STB module first normalizes the input features to X, whose dimension is (B×C×H×W), and then divides the features into non-overlapping windows, each window size is M 2 , so the overall feature dimension is modified to In this way, each feature size is (B×C×M 2 ) window X i Apply the self-attention mechanism to reduce the amount of computation. Specifically, for each X i , through the linear projection matrix P shared across windows Q , P K and P V Encode to obtain the feature vector of each window, and use it as the Query, Value and Key of the multi-head attention mechanism. In this way, the self-attention mechanism is executed separately for each window. For each window feature, the self-attention mechanism is executed in parallel in the actual process, and the results are connected in series to the results of MSA. Then the first STB module normalizes the output features of the residual connection through the LN layer again, and encodes the results of the multi-head self-attention through an MLP. The purpose of this encoding is to further extract the features and compress the features with too high dimensions. Finally, the output result of the first STB module is obtained by superimposing the residuals. The processing process in the first STB module is expressed by the following formula:

[0038] X←MSA(LN(X))+X

[0039] X←MLP(LN(X))+X

[0040] The entire first STB module can extract and fuse the local detail information inside the window at the corresponding position, so as to make up for the missing local detail information of the deep information during the feature extraction process, but lacks feature extraction of cross-window information. Therefore, the second STB module adds a shift window mechanism to the input features.

[0041] Specifically, the data processing flow in each second STB module is basically similar to that in the first STB module, and the only difference is that the second STB module adds a window shift operation before the multi-head self-attention module compared to the first STB module. And in this embodiment, when the shift operation is performed on the windows divided in the first STB module, the moving distances in both the horizontal and vertical directions are the half of the window size M rounded down, that is, the windows originally divided in the second STB module need to be moved After the position of , multi-head self-attention operation is performed, and the first STB module directly performs multi-head self-attention operation. Therefore, in the second STB module, the single feature of the input is successively passed through layer normalization, window shift, multi-head attention mechanism, residual connection, layer normalization, linear classifier and residual connection, to form an output feature transmitted to the rear, wherein except for window shift, the remaining layer normalization, multi-head attention mechanism, residual connection, layer normalization, linear classifier and residual connection are the same as the first STB module. Due to the increase of window shift before the multi-head attention mechanism, the multi-head attention mechanism MSA in the second STB module uses the feature of the input module after window shift as Query, Value and Key for attention fusion, thereby obtaining an attention graph.

[0042] Continue to see Figure 1 As shown in the figure, in the second feature extraction path of the MS image, the features output by the previous second STB module are used as the input of the next first STB module and are passed step by step until they reach the last second STB module. The feature vector output by the last second STB module is M1, and its dimension is (B×C×H2×W2); in the first feature extraction path of the PAN image, the features output by the previous second STB module are input to the PatchMerging module. The PM module selects elements of the input features at intervals of one position in the row and column directions through the PatchMerging operation to form a new window image block, and then all the window image blocks are spliced ​​together as a whole tensor. Finally, they are expanded and the channel dimension is adjusted to twice the original through a fully connected layer to form the output features passed to the next first STB module, and are passed step by step until they reach the last PM module. The feature vector output by the last PM module is P1, and its dimension is (B×C×H2×W2).

[0043] In the above SWCPAN model, the output features of the second STB module in the second group of STB modules are used as the deep features extracted by the deep feature extraction module in both feature extraction paths. The number of deep feature channels finally extracted by the deep feature extraction module can be set to 64.

[0044] Finally, the shallow features and deep features extracted from the two features need to be fused and input into the image reconstruction module. In the image reconstruction module, the input spliced ​​features are convolved and upsampled multiple times to obtain the final full-color sharpening result. Specifically, the deep feature M1 extracted from the MS image by the deep feature extraction module and the shallow feature M0 extracted from the MS image by the shallow feature extraction module are processed by residual to obtain feature M2, that is: M2 = M1 + M0. M2 is spliced ​​with the deep feature P1 extracted from the PAN image by the deep feature extraction module to obtain the splicing result, that is: M′ = concat(P1, M2), and the splicing result M′ is input into the image reconstruction module.

[0045] Specifically, the image reconstruction module consists of four groups of 3×3 convolutional layers and two groups of pixel reorganization layers (Pixel-Shufflelayer), such as Figure 3 As shown, the image reconstruction module is specifically composed of the first 3*3 convolution layer, the first pixel reorganization layer, the second 3*3 convolution layer, the second pixel reorganization layer, the third 3*3 convolution layer and the fourth 3*3 convolution layer, which are cascaded in sequence. The feature M′ output from the deep feature extraction module first enters the first convolution layer to extract the feature F1, whose dimension is (4C×H0×W0), and then adjusts the feature size to (C×2H0×2W0) through the first pixel reorganization layer, and then inputs it to the second convolution layer to extract the feature F2, whose dimension is (4C×2H0×2W0), and then adjusts the feature size to (C×4H0×4W0) through the second pixel convolution layer, and transmits the output feature to the third convolution layer to extract the feature F3, whose dimension is (C×4H0×4W0), and finally sends F3 to the last convolution layer to output the final feature map vector F, whose dimension is (4×4H0×4W0)

[0046] It should be noted that the pan-sharpening model SWCPAN is pre-trained using the training data after the Wald protocol before being used for actual pan-sharpening, and the loss function used in the pan-sharpening model training may use the mean absolute error.

[0047] The above remote sensing image full color sharpening method combining SwinTransformer and CNN is applied to a specific embodiment to demonstrate the technical effect that can be achieved.

[0048] Example

[0049] In this embodiment, the above remote sensing image pan-sharpening method based on SwinTransformer and CNN fusion is applied to a specific data set. The overall training and testing process of the SWCPAN model can be divided into three stages: data preprocessing, model training, and image prediction. Figure 4 shown.

[0050] 1. Data preprocessing stage

[0051] Step 1: For the original remote sensing images MS and PAN, image preprocessing is performed. First, image cutting, flipping and other operations are performed, and then data enhancement is performed and processed into images of the same size (128*128).

[0052] Step 2: Generate GroundTruth through the Wald protocol, downsample the preprocessed MS and PAN images to one-fourth of their original size as the original input image of the model, and use the original MS image as the GroundTruth of the model.

[0053] 2. Model training

[0054] Step 1: Build a training data set and divide the training data set into batches according to a fixed batch size, with a total of N.

[0055] Step 2: Sequentially select a batch of training samples with index i from the training data set, where i∈{0,1,…,N}. Use each batch of training samples to train the full color sharpening model SWCPAN. The specific structure of SWCPAN is as described above and will not be repeated here. During the training process, the mean absolute error of each training sample is calculated. , and based on the total loss of all training samples in the batch , adjust the network parameters in the entire model until all batches of the training data set participate in the model training. After reaching the specified number of iterations, the model converges and the training is completed.

[0056] 3. Panchromatic sharpening of remote sensing images

[0057] The images in the test set are directly used as input into the trained panchromatic sharpening model SWCPAN, and the final prediction is a HRMS image with both high spatial resolution and high spectral resolution, thereby achieving panchromatic sharpening.

[0058] In this embodiment, the test results are as follows:

[0059]

[0060] In this embodiment, peak signal-to-noise ratio (PSNR), structural similarity (SSIM), spectral angle mapper (SAM) and relative dimensionless overall error (ERGAS) are used as evaluation indicators of the model effect. It can be seen that the full-color sharpening model SWCPAN that combines SwinTransformer and CNN can well achieve full-color sharpening effect for remote sensing images, and generate HRMS images with both high spatial resolution and spectral resolution. While the fused image is smoother, it has a certain improvement in effect compared to the traditional Transformer method.

[0061] In order to further demonstrate the influence of the extraction of local features in the SwinTransformer and CNN modules in SWCPAN on the overall effect, this example conducts an ablation experiment on SWCPAN under two different settings: (I) removing the shallow feature extraction module and (II) replacing the STB module in the deep feature extraction module with the same number of conventional Transformer modules. Both settings reduce the model's ability to extract local features of the image to a certain extent.

[0062] In this ablation experiment, the test results are as follows:

[0063]

[0064] From the table, we can see that compared with the experimental settings of Ⅰ and Ⅱ, the original SWCPAN shows better results in all four indicators on the Gaofen-2 dataset, which shows that the design of SwinTransformer and CNN modules in the SWCPAN model of the present invention is conducive to improving the full-color sharpening effect. The shallow feature module relying on CNN focuses on the local features of the image, and the innovative process relying on SwinTransformer local attention and shift window brings better nonlinear texture features, further improving the ability to extract local features.

[0065] In summary, the innovative process of combining SwinTransformer and CNN solves the defect that the current panchromatic sharpening of remote sensing images focuses on global features while ignoring local features, and provides the possibility for further development of the combination of CNN and Transformer in the panchromatic sharpening task.

[0066] The above-described embodiment is only a preferred solution of the present invention, but it is not intended to limit the present invention. A person skilled in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.

Claims

1. A remote sensing image full color sharpening method based on SwinTransformer and CNN fusion, characterized by: The panchromatic image and the multispectral image to be panchromatically sharpened are input into the panchromatic sharpening model, and the panchromatic sharpening model includes a first feature extraction path, a second feature extraction path and an image reconstruction module. The panchromatic image and the multispectral image are respectively used as input images of the first feature extraction path and the second feature extraction path. The two input images are respectively subjected to shallow feature extraction and deep feature extraction by a shallow feature extraction module based on CNN and a deep feature extraction module based on SwinTransformer in their respective feature extraction paths. Finally, the deep features and the shallow features extracted in the second feature extraction path are fused through residual connection and then spliced ​​with the deep features extracted in the first feature extraction path in the channel dimension. The spliced ​​features are input into the image reconstruction module for image reconstruction. The shallow feature extraction modules in the two feature extraction paths are both CNN convolutional networks formed by cascading a group of convolutional modules; the deep feature extraction module in the first feature extraction path is formed by cascading two groups of STB modules, and the front end of each group of STB modules is equipped with a PM module, and the PM module is used to perform PatchMerging operation on the input of each group of STB modules; the deep feature extraction module in the second feature extraction path is formed by cascading two groups of STB modules; each group of STB modules in the two feature extraction paths is formed by cascading a first STB module and a second STB module; the first STB module uses the shallow feature map output by the front-end cascade module as input, and each input shallow feature map is normalized by the layer and then divided into non-overlapping local window blocks according to the fixed-size partitioning window. For each local window block, it is encoded through a linear projection matrix shared across windows to obtain the feature vector of each local window, and then the feature vector of each local window is used as the query of the multi-head attention mechanism in the multi-head attention layer. Query, value Value and key Key An attention map is obtained by attention fusion, and an intermediate feature map is obtained by residually connecting the attention map with the input shallow feature map, and then the intermediate feature map is residually connected with the result after layer normalization and linear classifier to obtain the output features of the first STB module; the second STB module uses the output features of the first STB module of the front-end cascade as input features, and the second STB module adds a window shift operation before the multi-head attention layer of the first STB module; in both feature extraction paths, the output features of the second STB module in the second group of STB modules are used as the deep features extracted by the deep feature extraction module; In the image reconstruction module, multiple convolutions and up-sampling are performed on the input splicing features to obtain the final panchromatic sharpening result.

2. The remote sensing image full color sharpening method based on SwinTransformer and CNN fusion as claimed in claim 1, characterized in that: The deep feature extraction module uses the Panformer model as a baseline model, and is obtained by replacing the self-attention modules in the Panformer model with the STB module.

3. The remote sensing image full color sharpening method based on SwinTransformer and CNN fusion as claimed in claim 1, characterized in that: The CNN convolutional network is composed of five 3*3 convolutional layers and one 1*1 convolutional layer cascaded in sequence.

4. The remote sensing image pan-sharpening method based on SwinTransformer and CNN fusion as claimed in claim 1, characterized in that: In the second STB module, when the shift operation is performed on the window divided in the first STB module, the moving distances in both the horizontal and vertical directions are the lower integer values ​​of half the window size.

5. The remote sensing image pan-sharpening method based on SwinTransformer and CNN fusion as claimed in claim 1, characterized in that: The image reconstruction module is formed by sequentially cascading a first 3*3 convolutional layer, a first pixel reorganization layer, a second 3*3 convolutional layer, a second pixel reorganization layer, a third 3*3 convolutional layer and a fourth 3*3 convolutional layer.

6. The remote sensing image pan-sharpening method based on SwinTransformer and CNN fusion as claimed in claim 1, characterized in that: Before being used for actual pan-sharpening, the pan-sharpening model is pre-trained using training data generated after the Wald protocol, and the loss function used in the pan-sharpening model training is the mean absolute error.

7. The remote sensing image full color sharpening method based on SwinTransformer and CNN fusion as claimed in claim 1, characterized in that: When the PM module performs the PatchMerging operation, it selects elements at intervals of two positions in the row and column directions of the input features to form new window image blocks, and then concatenates all the window image blocks as a whole tensor. Finally, it is expanded and the channel dimension is adjusted to twice the original through a fully connected layer to form the output features transmitted to the rear.

8. The remote sensing image pan-sharpening method based on SwinTransformer and CNN fusion as claimed in claim 1, characterized in that: In the first STB module, the size of the divided window for forming the local window block is fixed to 4×4.

9. The remote sensing image pan-sharpening method based on SwinTransformer and CNN fusion as claimed in claim 1, characterized in that: In the deep feature extraction module, the number of deep feature channels finally extracted is 64.

10. The remote sensing image full color sharpening method based on SwinTransformer and CNN fusion as claimed in claim 1, characterized in that: In the second STB module, the multi-head attention mechanism uses a single feature of the input as Query , Value and Key Perform attention fusion to obtain the attention map.

Citation Information

Patent Citations

  • Panchromatic sharpening method and system based on multi-scale delay channel attention network

    CN114549366A

  • Transform-based optical remote sensing target detection method

    CN114821357A