Multi-temporal Remote Sensing Image Cloud Removal Method Based on Global Perception Selective Fusion
By introducing a generative adversarial network framework of global perceptual selective fusion in the multi-time phase remote sensing image decloud method, combining the triple weight selection module and the ReSwin global modeling module, the problem of spatiotemporal feature fragmentation and limited global context modeling performance is solved, and high-quality decloud image generation is achieved.
Patent Information
- Application Number
- CN202510362711.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The existing multi-time phase remote sensing image decloud method based on neural networks breaks the solid correlation of space-time features and limits the global context modeling efficiency.
A generative adversarial network framework based on global perception selective fusion is adopted. Through the collaborative work of generator and discriminator, the triple weight selection module and ReSwin global modeling module are used to realize the cloud removal of multi-time phase remote sensing images.
It effectively solves the problem of space-time discontinuity in multi-temporal data fusion, realizes high-quality feature reconstruction in cloud-covered areas and retains surface details, and improves the quality of cloud-de-cloud images.
Smart Images

Figure CN119887589B_ABST
Abstract
Description
Technical Field
[0001] The present invention discloses a multi-temporal remote sensing image cloud removal method based on global perception selective fusion, belonging to the technical field of image cloud removal. Background Art
[0002] Affected by atmospheric conditions, the spectral information loss caused by cloud occlusion seriously damages the continuity of the image, greatly restricting its effectiveness in subsequent analysis and applications. With the rapid development of deep learning technology, data-driven cloud removal methods have gradually become a research hotspot, showing great application potential in key fields such as environmental dynamic monitoring and disaster emergency response. Especially in the context of the rapid increase in high-resolution earth observation satellite data, how to break through the performance bottleneck of traditional methods and develop efficient and robust cloud removal algorithms has become an urgent need to improve the usability and timeliness of remote sensing data products.
[0003] The current mainstream remote sensing image cloud removal methods can be classified into four categories: methods based on spatial inpainting, methods based on spectral analysis, hybrid methods, and methods based on multi-temporal data. In recent years, multi-temporal cloud removal technology has received extensive attention due to its unique advantages. This method realizes the repair of cloud-covered areas by integrating data of different time phases in the same geographical area. Compared with the spatial domain cloud removal method that is easily restricted by complete cloud cover and the problem of heterologous data registration in SAR image fusion, the multi-temporal method can not only break through the bottleneck of continuous cloud occlusion but also avoid the complexity of multi-modal data fusion, showing significant application value.
[0004] Early studies mostly adopted traditional multi-temporal image stitching strategies, such as the MODIS product method based on optimal pixel synthesis. However, such methods are sensitive to the radiometric differences and surface changes between time phases, easily resulting in spatio-temporal discontinuity in the repair results. With the development of deep learning, cloud removal methods based on neural networks have shown powerful data modeling capabilities, learning the cloud distribution law through large-scale time series data. However, their feature fusion mechanism lacks effective spatio-temporal correlation modeling. Subsequent studies have gradually improved the feature selection and global modeling capabilities by introducing attention mechanisms and Swin Transformer architectures. However, there are still two key bottlenecks. First, most existing attention modules adopt an independent calculation mode for channel and spatial weights, splitting the inherent correlation of spatio-temporal features. Second, the Transformer architecture is directly transplanted from general vision tasks and has not been adaptively improved for the characteristics of multi-temporal data, restricting the global context modeling efficiency. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-temporal remote sensing image cloud removal method based on global perception selective fusion to solve the problems in the prior art that the cloud removal method based on neural networks splits the inherent correlation of spatio-temporal features and restricts the global context modeling efficiency.
[0006] A multi-temporal remote sensing image cloud removal method based on global perception selective fusion, which includes inputting three-temporal cloudy remote sensing images into a generative adversarial network framework, and after obtaining the output result, re-inputting it into the generative adversarial network framework, repeating the cycle until a cloud-free image that meets the requirements is obtained.
[0007] The generative adversarial network framework includes a generator and a discriminator. The generator includes an encoder, a high-level feature extraction module, and a decoder. The discriminator outputs a probability value indicating the image source. The multi-scale convolutional layer extracts the image texture features, and the adversarial loss function optimizes the network parameters. The adversarial loss includes a pixel-level reconstruction loss and an adversarial training loss.
[0008] The encoder includes a downsampling module, a triple weight selection module, and a feature fusion layer;
[0009] The downsampling module includes three groups of parallel downsampling units. Each group of downsampling units includes a convolutional layer, a ReLU activation function, and a batch normalization layer, and realizes the reduction of the feature map size and the doubling of the number of channels through convolutional kernels with different strides;
[0010] The triple weight selection module includes three parallel processing branches and a dynamic weight generator.
[0011] The input feature of the triple weight selection module has a size of , where is the number of channels, is the height,
[0012] Input into the first processing branch. After passing through the first rotation layer, has a size of . After passing through the spatial attention layer and then through the second rotation layer, has a size of ;
[0013] Input into the second processing branch and pass through the spatial attention layer;
[0014] Input into the third processing branch. After passing through the first rotation layer, has a size of . After passing through the spatial attention layer and then through the second rotation layer, has a size of ;
[0015] The results of the three processing branches are respectively used to generate three weight maps through global pooling, average pooling, and convolutional layers, obtaining the important parts in the feature maps of each processing branch, and then the input features are weighted and restored to the original dimension.
[0016] The dynamic weight generator includes a first feature transformation layer, a second feature compression layer, and a weight prediction layer;
[0017] The first feature transformation layer uses a 3×3 convolutional kernel with a stride of 1, and the number of output channels is expanded to 32 times that of the input features, and a non-linear mapping is performed through the ReLU activation function;
[0018] The second feature compression layer realizes spatial feature aggregation through a 3×3 convolutional kernel, and the number of output channels is reduced to 16 times, and the ReLU activation function is connected to enhance feature sparsity;
[0019] The weight prediction layer uses a 1×1 convolutional kernel to compress the number of channels to 3, and generates three-dimensional weight coefficients of the original feature map after passing through the Softmax activation function.
[0020] The feature fusion layer multiplies the three-dimensional weight coefficients with the corresponding weight maps respectively, and then sums them to obtain the enhanced feature map.
[0021] The advanced feature extraction module adopts a UNet structure, including four residual downsampling blocks and six transposed convolutional upsampling blocks;
[0022] The residual downsampling block extracts high-level semantic features through 3×3 convolution, and each block of the upsampling block contains a ReLU activation function, a transposed convolution layer, and a batch normalization layer, gradually restoring the size of the feature map.
[0023] The decoder includes two ReSwin stages, and each ReSwin stage includes a ReLN normalization layer and a variable module;
[0024] The ReLN normalization layer includes three branches, and the three branches correspond to three outputs. The first branch includes a mean layer and a point convolution layer to form the first output. The second branch includes a variance layer and a point convolution layer to form the second output. The central branch includes layer normalization. The output of the layer normalization is combined with the first output, and then feature fusion is performed with the second output to obtain the central output.
[0025] The input features of the decoder are sequentially input into the ReLN normalization layer to generate three outputs. The central output enters the variable module. The second output is combined with the output of the variable module, and then feature fusion is performed with the first output, and finally output.
[0026] The variable module of the first ReSwin stage is a windowed multi-head self-attention unit W-MSA, and the variable module of the second ReSwin stage is a multi-layer perceptron MLP.
[0027] Compared with the prior art, the present invention has the following beneficial effects: The present invention effectively solves the spatio-temporal discontinuity problem in multi-temporal data fusion, realizes the retention of surface details and the capture of temporal variation rules in the feature reconstruction of cloud-covered areas, and obtains a cloud-removed image with better quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a structural diagram of a triple weight selection module;
[0029] Figure 2 It is a structural diagram of the ReSwin stage;
[0030] Figure 3 It is a structural diagram of the ReLN normalization layer. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] A multi-temporal remote sensing image cloud removal method based on globally-aware selective fusion includes inputting three-temporal cloud-covered remote sensing images into a generative adversarial network framework, re-inputting the output result into the generative adversarial network framework, and repeating the loop until a cloud-free image that meets the requirements is obtained.
[0033] The generative adversarial network framework includes a generator and a discriminator. The generator includes an encoder, a high-level feature extraction module, and a decoder. The discriminator outputs a probability value indicating the image source. The multi-scale convolutional layer extracts the image texture features, and the adversarial loss function optimizes the network parameters. The adversarial loss includes a pixel-level reconstruction loss and an adversarial training loss.
[0034] The encoder includes a downsampling module, a triple weight selection module, and a feature fusion layer;
[0035] The downsampling module includes three groups of parallel downsampling units. Each group of downsampling units includes a convolutional layer, a ReLU activation function, and a batch normalization layer, and realizes the reduction of the size of the feature map and the doubling of the number of channels through convolutional kernels with different strides;
[0036] The triple weight selection module includes three parallel processing branches and a dynamic weight generator.
[0037] The input feature of the triple weight selection module has a size of , where is the height, is the width;
[0038] Input into the first processing branch. After passing through the first rotation layer, the size of is . After passing through the spatial attention layer and then through the second rotation layer, the size of ;
[0039] Input into the second processing branch and pass through the spatial attention layer;
[0040] Input into the third processing branch. After passing through the first rotation layer, the size of is the size of ;
[0041] For the results of the three processing branches, generate three weight maps respectively through global pooling, average pooling, and convolutional layers to obtain the important parts in the feature maps of each processing branch, and then restore the original dimension after weighting the input features.
[0042] The dynamic weight generator includes a first feature transformation layer, a second feature compression layer, and a weight prediction layer;
[0043] The first feature transformation layer uses a 3×3 convolutional kernel with a stride of 1, and the number of output channels is expanded to 32 times that of the input features, and a non-linear mapping is performed through the ReLU activation function;
[0044] The second feature compression layer realizes spatial feature aggregation through a 3×3 convolutional kernel, and the number of output channels is reduced to 16 times, and the ReLU activation function is connected to enhance feature sparsity;
[0045] The weight prediction layer uses a 1×1 convolutional kernel to compress the number of channels to 3, and generates three-dimensional weight coefficients of the original feature map after passing through the Softmax activation function.
[0046] The feature fusion layer multiplies the three-dimensional weight coefficients with the corresponding weight maps respectively, and then sums them up to obtain the enhanced feature map.
[0047] The advanced feature extraction module adopts the UNet structure, including four residual downsampling blocks and six transposed convolutional upsampling blocks;
[0048] The residual downsampling block extracts high-level semantic features through 3×3 convolution. Each block of the upsampling block contains a ReLU activation function, a transposed convolution layer, and a batch normalization layer, and gradually restores the size of the feature map.
[0049] The decoder includes two ReSwin stages, and each ReSwin stage includes a ReLN normalization layer and a variable module;
[0050] The ReLN normalization layer includes three branches, and the three branches correspond to three outputs. The first branch includes a mean layer and a point convolution layer to form a first output. The second branch includes a variance layer and a point convolution layer to form a second output. The central branch includes layer normalization. The output of the layer normalization is combined with the first output and then feature fused with the second output to obtain a central output.
[0051] The input features of the decoder are sequentially input into the ReLN normalization layer to generate three outputs. The central output enters the variable module. The second output is combined with the output of the variable module and then feature fused with the first output, and finally output.
[0052] The variable module of the first ReSwin stage is a windowed multi-head self-attention unit W-MSA, and the variable module of the second ReSwin stage is a multi-layer perceptron MLP.
[0053] The triple weight selection module TWSM of the present invention is as Figure 1 shown, including three parallel processing branches and a dynamic weight generator. The input features of the triple weight selection module have a size of , where is the number of channels, is the height, is input into the first processing branch. After passing through the first rotation layer, has a size of . After passing through the spatial attention layer and then through the second rotation layer, has a size of ; is input into the second processing branch and passes through the spatial attention layer; is input into the third processing branch. After passing through the first rotation layer, has a size of . After passing through the spatial attention layer and then through the second rotation layer, has a size of ; The results of the three processing branches are respectively passed through global pooling, average pooling, and a convolutional layer to generate three weight maps, obtaining the important parts in the feature maps of each processing branch. After weighting the input features, the original dimension is restored. The dynamic weight generator includes a first feature transformation layer, a second feature compression layer, and a weight prediction layer; the first feature transformation layer uses a 3×3 convolutional kernel with a stride of 1, and the number of output channels is expanded to 32 times that of the input features, and a non-linear mapping is performed through the ReLU activation function; the second feature compression layer realizes spatial feature aggregation through a 3×3 convolutional kernel, and the number of output channels is reduced to 16 times, and the ReLU activation function is connected to strengthen feature sparsity; the weight prediction layer uses a 1×1 convolutional kernel to compress the number of channels to 3, and after passing through the Softmax activation function, three-dimensional weight coefficients of the original feature map are generated. The feature fusion layer multiplies the three-dimensional weight coefficients with the corresponding weight maps respectively, and then sums them to obtain the enhanced feature map 。
[0054] The ReSwin stage is as Figure 2 shown. Each ReSwin stage includes a ReLN normalization layer and a variable module; the input features of the decoder are sequentially input into the ReLN normalization layer to generate three outputs. The central output enters the variable module, the second output is combined with the output of the variable module, and then feature fusion is performed with the first output, and finally the output features are generated 。The variable module of the first ReSwin stage is the windowed multi-head self-attention unit W-MSA, and the variable module of the second ReSwin stage is the multi-layer perceptron MLP.
[0055] The ReLN normalization layer is as Figure 3 shown. The ReLN normalization layer includes three branches, and the three branches correspond to three outputs. The first branch includes a mean layer and a point convolutional layer to form the first output ,The second branch includes a variance layer and a point convolutional layer to form the second output ,The central branch includes layer normalization. The output of the layer normalization is combined with the first output, and then feature fusion is performed with the second output to obtain the central output 。
[0056] The present invention adopts two datasets, namely the STGAN dataset and the Sen2_MTC dataset. The image acquisition period of the STGAN dataset covers 2017 - 2021, forming a continuous observation sequence across seasons and years. It is constructed from multi-temporal remote sensing images of the Sentinel-2 satellite, with a new image taken at the same geographical location every 6 days on average, covering various climate zones and terrains globally. This dataset contains 3,130 cloud-free images and the corresponding 9,390 cloud-containing images (each cloud-free image corresponds to three cloud-containing images at different time phases). The image size is 256×256 pixels, including four bands: red (R), green (G), blue (B), and near-infrared (NIR). The present invention only uses the RGB band data.
[0057] The Sen2_MTC dataset is screened from Sentinel-2 images, covering real cloud distribution scenarios. It contains 50 non-overlapping image patches, each patch containing 70 consecutive images, with a size of 256×256 pixels, a time span of 14 months (2019 - 2020), and a time interval maintaining 5-day-level observations. Focusing on high-cloud-occurrence regions, the 50 image patches are distributed in the Southeast Asian monsoon region, the equatorial convective active zone, and the mid-latitude frontal cloud system region. This dataset contains complex light changes and cloud dynamic movements to verify the robustness of the model under dynamic cloud interference.
[0058] The present invention verifies the effectiveness of the proposed method from three parts: dataset description, comparative experiments, and ablation experiments. All experiments are based on the same hardware environment (V100) and hyperparameter settings (initial learning rate 0.0002, batch size 8) to ensure the fairness of comparison. To verify the effectiveness of the proposed network, the following representative methods are selected for comparison, including Pix2Pix, ResNet, SIF-GAN, and MSC-GAN.
[0059] Pix2Pix stitches the RGB bands of three-time-phase images (t-1, t, t+1) along the time dimension into a 9-channel input, and realizes cross-time-phase feature interaction through the skip connections of the U-Net generator. The encoder adopts 5 layers of downsampling (convolution with a stride of 2), and each layer stacks a spatio-temporal attention mechanism to enhance the ability to capture cloud movement trajectories. An improved PatchGAN structure is adopted, expanding the discriminant window into a spatio-temporal cube to simultaneously evaluate the spatial texture continuity and time dynamic consistency.
[0060] ResNet is based on the ResNet generator, with the input being the stitching features of multi-time-phase images, relying on residual learning to alleviate gradient disappearance. A pyramid-style residual block is constructed, including a basic residual path, an expanded residual path, and a feature fusion layer. On the basis of ResNet, the temporal correlation of time-series images is processed, and the gating residual connection fuses the temporal features with the convolutional spatial features.
[0061] STGAN is a multi-temporal cloud removal network that combines U-Net and ResNet, introducing the infrared band to assist cloud detection and feature recovery. In the first stage, the U-Net cloud detection module is pre-trained with synthetic cloud data. In the second stage, the encoder parameters are frozen, and the ResNet feature recovery module is jointly trained to avoid feature conflicts caused by modal differences.
[0062] SIF-GAN solves the problems of cloud residue and distorted ground object recovery caused by redundant feature fusion mechanisms in multi-temporal remote sensing image cloud removal tasks, and realizes the effective screening and fusion of multi-temporal features through dynamic weight allocation. It includes the design of a dynamic feature screening mechanism. An improved channel attention module is introduced in the encoding stage. Through parallel channel statistical analysis and spatial gradient perception, the response weights of cloud-contaminated areas in the RGB band are dynamically calculated. This module can suppress interference features in cloud-covered areas while enhancing texture details in cloud-free areas. At the same time, a dual-branch spatio-temporal feature interaction structure is designed: the main branch uses a residual shrinkage network to extract features of the current time phase; the auxiliary branch models the temporal dependence relationship through a gated recurrent unit (GRU) to generate a cross-temporal feature weight matrix, realizing the dynamic weighted fusion of adjacent time-phase features. A triple-constraint loss function, including adversarial loss, spectral fidelity loss, and edge sharpening loss.
[0063] MSC-GAN uses the proposed multi-stream complementary coding architecture to address the dual challenges of insufficient cross-temporal feature interaction and deep feature semantic degradation in multi-temporal remote sensing image cloud removal tasks. A parallel three-stream coding network (current time phase, previous time phase, subsequent time phase) is constructed, and dynamic calibration of temporal features is achieved through a bidirectional gated attention module. Each coding stream contains 5 levels of downsampling layers. When generating feature maps at each level, a cross-stream feature distillation operation is introduced to retain cloud movement trajectories and stable surface features. Residual dense connections are used to replace traditional skip connections, establishing an inter-layer feature reuse channel in the encoding stage, enabling shallow cloud detection features to directly participate in deep surface reconstruction and alleviating information attenuation caused by the increase in network depth. In addition, a group feature reweighting module is designed. The high- and low-level features are divided into 4 semantic groups (terrain group, texture group, spectral group, context group) according to the channel dimension. After extracting multi-scale features through parallel dilated convolutions (dilation rate = 2, 4, 6, 8), a dynamic weight allocator is used to enhance cross-group features.
[0064] In the present invention, the peak signal-to-noise ratio PSNR, structural similarity SSIM, and cosine similarity are selected as accuracy evaluation indicators, and experiments are carried out on the STGAN dataset as shown in Table 1.
[0065] Table 1. Experiments on the STGAN dataset (unit: dB)
[0066] ;
[0067] In Table 1, GSF represents the method of the present invention, and the experiments on the Sen2_MTC dataset are shown in Table 2.
[0068] Table 2. Experiments on the Sen2_MTC Dataset (Unit: dB)
[0069] ;
[0070] In the STGAN and Sen2_MTC datasets, the present invention achieves the best results in all three metrics of PSNR, SSIM, and cosine similarity (CS). On the STGAN dataset, MSC-GAN achieves sub-optimal results due to its multi-stream complementary architecture, but its performance on Sen2_MTC drops significantly (SSIM decreases by 0.22), indicating its insufficient adaptability to dynamic real clouds. The network proposed by the present invention maintains stable performance in both datasets through the synergistic effect of the temporal weighted selection module (TWSM) and ReSwin global modeling.
[0071] In terms of the ability to recover thick clouds, in the high-cloud coverage area of the STGAN dataset, the proposed network accurately recovers the boundaries of the ground objects under the clouds by capturing multi-temporal global context, while ResNet and Pix2Pix result in ground object confusion, such as mis-recovering bare soil as vegetation, due to the lack of temporal information fusion. In terms of processing dynamic clouds, in the cirrus cloud occlusion scenario of Sen2_MTC, the proposed network effectively suppresses cloud artifacts through the long-range dependence modeling of the ReSwin module, while SIF-GAN results in poor detail recovery of the ground under the clouds due to the deviation in the allocation of local feature weights.
[0072] To verify the contributions of each module, key components were gradually introduced with U-Net as the baseline, and the experimental results are shown in Table 3.
[0073] Table 3. Ablation Experiments (Unit: dB)
[0074] ;
[0075] When using TWSM alone, the PSNR only increases by 0.21 dB. The reason is that although TWSM can enhance the feature representation of key regions (such as cloud edges), it lacks a cross-temporal feature fusion mechanism and cannot fully utilize multi-temporal complementary information. When using the ReSwin module alone, the PSNR increases by 0.95 dB, but due to the lack of high-quality local feature support, the global modeling is limited, resulting in a blurred recovery result. When TWSM and ReSwin are used in combination, the PSNR increases by 1.32 dB. The local key features extracted by TWSM are complementary to the global temporal patterns modeled by ReSwin.
[0076] In the present invention, the triple-weight selection module includes a multi-dimensional attention branch that establishes the correlation between the channel dimension and the spatio-temporal dimension through tensor rotation operations; a dynamic weight generator that generates the attention weights of each branch based on the spatial context of the feature map; and an adaptive fusion layer that performs weighted fusion on the features processed by the spatial attention, channel attention, and spatio-temporal attention branches. During the four-stage encoding process, the resolution of the output feature map of each stage gradually decreases (from ), and the number of channels gradually increases ( ). Each decoding stage includes: a ReLN normalization layer that maintains the relative differences between blocks through learnable scaling parameters and translation parameters; a windowed multi-head self-attention unit that calculates the temporal-spatial joint attention within a local window and dynamically models temporal patterns such as vegetation growth and water body changes; and a multi-layer perceptron that performs a non-linear transformation on the attention output. During the decoding process, the feature map resolution is restored to the original size through two upsampling operations.
[0077] The key components included in TWSM are BatchNorm, ReLU activation layer, convolutional layer, spatial attention mechanism, and dynamic weight generator. First, the input feature map is sent to a three-dimensional feature rotation unit that includes three parallel processing branches: the first branch rotates the input feature map counterclockwise by 90° along the channel-width plane to generate X_out1 with dimensions H×C×W, capturing the interaction features between the channel and width dimensions; the second branch rotates the input feature map counterclockwise by 90° along the channel-height plane to generate X_out2 with dimensions W×C×H, modeling the correlation between the channel and height dimensions; the third branch retains X_out3 with the original dimensions C×H×W and directly processes the local features of the spatial dimension. Then, the three branches respectively generate weight maps through global pooling and average pooling, convolutional layers, obtain the important parts in the feature maps of each branch, and restore the original dimensions after weighting the input feature map.
[0078] The dynamic weight generator is composed of three convolutional neural networks connected in sequence. Its processing flow includes a first feature transformation layer that uses a 3×3 convolutional kernel with a stride of 1 and expands the output number of channels to 32 times that of the input features, and performs a non-linear mapping through the ReLU activation function; a second feature compression layer that realizes spatial feature aggregation through a 3×3 convolutional kernel, reduces the output number of channels to 16 times, and connects the ReLU activation function to enhance the feature sparsity; and a weight prediction layer that compresses the number of channels to 3 using a 1×1 convolutional kernel to generate the three-dimensional weight coefficients of the original feature map.
[0079] The core idea of the ReSwin module design is the necessary improvement made to make the Swin Transformer module more suitable for the de-clouding task. Utilizing the global information capture ability of the Swin Transformer enables the model to learn the temporal patterns in multi-temporal images. Since there are slight differences in multi-temporal images, it is not possible to simply splice multiple images directly to obtain the final result. It is necessary to learn the internal temporal patterns, such as climate or weather changes, to make a more reasonable prediction of the final image. Therefore, the powerful global information capture ability of the Swin Transformer is used to enrich the model. However, some of the components in the Swin Transformer are not suitable for low-level de-clouding tasks, so the module is improved for the de-clouding task. First, ReLN normalization is designed to replace the original layer normalization. Since layer normalization will smooth out the relative differences in brightness (mean) and contrast (variance) between different image patches, making it difficult for the model to distinguish between cloud-covered areas and clear surface areas. Therefore, the statistics mean and variance are reintroduced in ReLN normalization. The introduced parameters are passed through a convolutional layer and then an affine transformation is performed on the original output, and the mean and variance are output.
[0080] The spatio-temporal joint attention unit includes a window partitioning module that divides the feature map into 8×8 non-overlapping windows; a cross-window shifting mechanism that alternately uses window offsets of (4, 4) and (-4, -4) in adjacent layers; and multi-head attention calculation that is performed on the Q, K, V matrices within each window. The windowed multi-head self-attention unit reduces the computational complexity and effectively models the global context information of the image, better understanding the structure and ground object distribution of the entire image, and thus more accurately inferring the content of the area occluded by clouds. Then, after another ReLN normalization, it enters the improved MLP module.
[0081] The MLP includes a convolutional layer and a ReLU activation function. The activation function in the MLP of Swin is the GeLU activation function. The non-monotonicity of GeLU may lead to the problem of gradient reversal, and its non-linearity also results in its irreversibility. However, the essence of image de-clouding is the inverse process of recovering the original scene from the degraded input, and the irreversibility will hinder the image recovery. The simple inverse transformation property of ReLU is more conducive to image recovery. In addition, the number of convolutions is increased (twice before and after ReLU). After passing through the MLP, an affine transformation is performed according to the mean and variance weights output by ReLN, and finally the output is obtained.
[0082] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the same. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-temporal remote sensing image cloud removal method based on global perception selective fusion, characterized in that: The method includes inputting three-temporal cloud remote sensing images into a generative adversarial network framework, re-inputting the output results into the generative adversarial network framework, and repeating the cycle until a cloud-free image that meets the requirements is obtained; The generative adversarial network framework includes a generator and a discriminator. The generator includes an encoder, an advanced feature extraction module and a decoder. The discriminator outputs a probability value to represent the source of the image. The multi-scale convolution layer extracts the image texture features. The adversarial loss function optimizes the network parameters. The adversarial loss includes pixel-level reconstruction loss and adversarial training loss. The encoder includes a downsampling module, a triple weight selection module and a feature fusion layer; The downsampling module includes three groups of parallel downsampling units, each group of downsampling units includes a convolution layer, a ReLU activation function and a batch normalization layer, and achieves size reduction of feature maps and multiplication of the number of channels through convolution kernels of different step sizes; The triple weight selection module includes three parallel processing branches and a dynamic weight generator; Input features of triple weight selection module The size is , is the number of channels, is the height, is the width; Will Enter the first processing branch, after passing through the first rotation layer, The size is , after passing through the spatial attention layer and then the second rotation layer, The size is ; Will Input to the second processing branch, passing through the spatial attention layer; Will Enter the third processing branch, after passing through the first rotation layer, The size is , after passing through the spatial attention layer and then the second rotation layer, The size is ; The results of the three processing branches are respectively processed through global pooling, average pooling, and convolutional layers to generate three weight maps, obtain the important parts of the feature map of each processing branch, and restore the original dimension after weighting the input features; The dynamic weight generator includes a first feature transformation layer, a second feature compression layer and a weight prediction layer; The first feature transformation layer uses a 3×3 convolution kernel with a step size of 1. The number of output channels is expanded to 32 times the input feature, and nonlinear mapping is performed through the ReLU activation function. The second feature compression layer achieves spatial feature aggregation through a 3×3 convolution kernel, reducing the number of output channels to 16 times, and connecting a ReLU activation function to enhance feature sparsity; The weight prediction layer uses a 1×1 convolution kernel to compress the number of channels to 3, and generates the three-dimensional weight coefficients of the original feature map after the Softmax activation function; The decoder consists of two ReSwin stages, each of which includes a ReLN normalization layer and a variable module; The ReLN normalization layer includes three branches, and the three branches correspond to three outputs. The first branch includes a mean layer and a point convolution layer to form a first output. The second branch includes a variance layer and a point convolution layer to form a second output. The central branch includes layer normalization. The output of the layer normalization is combined with the first output, and then feature fused with the second output to obtain the central output. The input features of the decoder are sequentially input to the ReLN normalization layer to generate three outputs. The central output enters the variable module, the second output is combined with the output of the variable module, and then feature fused with the first output and finally output; The variable module of the first ReSwin stage is the windowed multi-head self-attention unit W-MSA, and the variable module of the second ReSwin stage is the multi-layer perceptron MLP.
2. The multi-temporal remote sensing image cloud removal method based on global perception selective fusion according to claim 1 is characterized in that: The feature fusion layer multiplies the three-dimensional weight coefficients with the corresponding weight maps respectively, and then sums them up to obtain the enhanced feature map.
3. The multi-temporal remote sensing image cloud removal method based on global perception selective fusion according to claim 2 is characterized in that: The high-level feature extraction module adopts the UNet structure, which includes four residual downsampling blocks and six transposed convolution upsampling blocks; The residual downsampling block extracts high-level semantic features through 3×3 convolution, and each block of the upsampling block contains a ReLU activation function, a deconvolution layer, and a batch normalization layer to gradually restore the feature map size.
Citation Information
Patent Citations
Method and system for LULC guided SAR visualization
US20230408682A1
Electronic device for colorizing black and white image using GAN based model comprising transformer block and method for operation thereof
US20240265587A1