Method and device for removing thin cloud of remote sensing image
Through the improved generative adversarial network, combined with multi-scale and convolutional block attention module, the problems of insufficient multi-scale feature processing and high computational complexity of the existing thin cloud removal technology are solved, and the cloud removal effect and quality of remote sensing images are improved, and the complex terrain and spectral conditions are adapted to complex terrain and spectral conditions are improved.
Patent Information
- Application Number
- CN202510521682.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-18
AI Technical Summary
The existing thin cloud removal technology has insufficient multi-scale feature processing capabilities, single feature extraction, high computational complexity and weak generalization capabilities, resulting in poor remote sensing image quality, especially under complex terrain and multi-spectral conditions.
The improved generative adversarial network is adopted, combined with the multi-scale attention module and the convolutional block attention module, and the generator and discriminator are trained through the multi-dimensional loss function to improve the processing capability and image quality of cloud features.
The model's adaptability to different cloud thicknesses and spatial structures is enhanced, feature extraction methods are enriched, cloud removal effect is improved, image quality and generalization ability are improved, and adverse effects are reduced due to computational complexity and insufficient generalization.
Smart Images

Figure CN120339119A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a method and device for removing thin clouds from remote sensing images. Background Art
[0002] Cloud cover seriously affects the usability of optical remote sensing data, and thin cloud removal is a key task to improve image quality. According to statistics, about 67% of the Earth's surface is covered by clouds, and the average cloud cover rate on land reaches 35%. Although thin clouds have a certain degree of transmissibility, their presence still leads to a decrease in image contrast and blurring of ground object details, seriously affecting the accuracy of applications such as land use classification, environmental monitoring, and disaster assessment. Traditional thin cloud removal methods are mainly divided into three categories: spectral feature-based, spatio-temporal modeling, and frequency domain filtering, but all have significant limitations:
[0003] Spectral feature-based methods use the spectral differences between clouds and ground objects for correction, but rely on prior assumptions, have poor adaptability to complex terrains (such as water bodies, ice and snow), and are prone to causing spectral distortion of ground objects.
[0004] Spatio-temporal modeling methods use temporal or spatial correlations to reconstruct cloud-free regions, but require high-quality reference images and are computationally complex.
[0005] Frequency domain filtering methods achieve cloud removal by suppressing the low-frequency components of clouds, but will lose the low-frequency information of the image, resulting in blurred ground object contours.
[0006] In recent years, deep learning methods have made significant progress in the task of thin cloud removal. In particular, generative adversarial networks (GANs) have become a research hotspot due to their unsupervised learning and high-fidelity image generation capabilities. However, many existing methods still face the following challenges: insufficient capture of multi-scale features: traditional GAN methods use single-scale convolutional kernels, which are difficult to effectively model the multi-scale spatial distribution characteristics of thin clouds, resulting in incomplete detail recovery (such as blurred textures, edge artifacts); limitations of the attention mechanism: existing attention models are mostly used independently, lacking the ability to synergistically enhance the channel-space dimensions, and unable to adaptively distinguish clouds from complex backgrounds.
[0007] Therefore, there is an urgent need for a solution in the current thin cloud removal technology that can adaptively fuse multi-scale features, dynamically enhance key regions, and balance details and authenticity. Summary of the Invention
[0008] The purpose of the present invention is to provide a method and device for removing thin clouds from remote sensing images to improve efficiency.
[0009] To achieve the above purpose, the present invention provides the following technical solutions:
[0010] In the first aspect, the present invention provides a method for removing thin clouds from remote sensing images, including:
[0011] Obtain remote sensing images and a trained target generative adversarial network; the remote sensing images include cloud-containing images to be processed; the target generative adversarial network includes an improved generator and an improved discriminator; the improved generator at least includes a multi-scale attention module and a convolutional block attention module;
[0012] Use the trained target generative adversarial network to perform image generation processing on the cloud-containing image to be processed to obtain a final cloud-free image.
[0013] Optionally, the improved generator further includes an input convolutional layer, a residual network, and an output convolutional layer;
[0014] The using the trained target generative adversarial network to perform image generation processing on the cloud-containing image to be processed to obtain a final cloud-free image at least includes:
[0015] Input the cloud-containing image to be processed into the input convolutional layer to obtain an image after feature extraction;
[0016] Input the image after feature extraction into the multi-scale attention module to obtain an image containing multi-scale feature information;
[0017] Input the image containing multi-scale feature information into the convolutional block attention module to obtain an image with recalibrated channel and spatial dimensions;
[0018] Input the image with recalibrated channel and spatial dimensions into the residual network to obtain an optimized image;
[0019] Input the optimized image into the output convolutional layer for processing to obtain the final cloud-free image.
[0020] Optionally, the inputting the image after feature extraction into the multi-scale attention module to obtain an image containing multi-scale feature information includes:
[0021] Extract context information at a specific level from the image after feature extraction and output multiple feature vectors;
[0022] Perform element-wise addition on the multiple feature vectors to obtain a comprehensive attention weight map;
[0023] Perform cross-channel information integration on the comprehensive attention weight map to form an attention mask;
[0024] Perform pixel-wise weighting of the attention mask and the image after feature extraction to obtain the image containing multi-scale feature information.
[0025] Optionally, the convolutional block attention module includes at least a channel attention sub-module and a spatial attention sub-module;
[0026] Inputting the image containing multi-scale feature information into the convolutional block attention module to obtain an image with recalibrated channel and spatial dimensions includes:
[0027] Inputting the image output by the multi-scale attention module into the channel attention sub-module to obtain a channel attention feature map;
[0028] Inputting the channel attention feature map into the spatial attention sub-module to obtain the image with recalibrated channel and spatial dimensions.
[0029] Optionally, inputting the image output by the multi-scale attention module into the channel attention sub-module to obtain a channel attention feature map includes:
[0030] Performing global average pooling on the image containing multi-scale feature information to obtain a first vector;
[0031] Performing global max pooling on the image containing multi-scale feature information to obtain a second vector;
[0032] Processing the first vector and the second vector using a fully connected layer to obtain a first weight and a second weight;
[0033] Adding and activating the first weight and the second weight to obtain a target channel attention weight;
[0034] Multiplying the target channel attention weight by the image containing multi-scale feature information to obtain the channel attention feature map.
[0035] Optionally, inputting the channel attention feature map into the spatial attention sub-module to obtain the image with recalibrated channel and spatial dimensions includes:
[0036] Performing average pooling and max pooling on the channel attention feature map for each channel to obtain a pooling result;
[0037] Concatenating the pooling results along the channel dimension to obtain a concatenated feature map;
[0038] Performing a convolution operation on the concatenated feature map to obtain an initial spatial attention weight;
[0039] Activating the initial spatial attention weight to obtain a target spatial attention weight;
[0040] Multiply the target spatial attention weights with the channel attention feature map to obtain the image with re-calibrated channel and spatial dimensions.
[0041] Optionally, before obtaining the remote sensing image and the trained target generative adversarial network, the method further includes:
[0042] Construct the target generative adversarial network; the target generative adversarial network includes an improved generator and the improved discriminator; the improved generator at least includes the multi-scale attention module and the convolutional block attention module;
[0043] Obtain the RICE1 thin cloud dataset; the RICE1 thin cloud dataset includes a training set and a test set;
[0044] Use the training set to perform multiple rounds of training on the target generative adversarial network in combination with a multi-dimensional loss function; the multi-dimensional loss function includes: L1 loss function, mean squared error loss function, and Softplus loss function;
[0045] After each round of training, use the test set to evaluate the target generative adversarial network to obtain an evaluation result;
[0046] When the evaluation result reaches a preset threshold, the training is completed.
[0047] Optionally, both the training set and the test set include multiple image pairs, and each image pair includes a real cloud-free image and a cloud-containing image;
[0048] The step of using the training set to perform multiple rounds of training on the target generative adversarial network in combination with a multi-dimensional loss function includes
[0049] In each round of training, input the cloud-containing image into the improved generator to obtain a cloud-free image to be evaluated;
[0050] Input the cloud-free image to be evaluated and the real cloud-free image into the improved discriminator, and output a discrimination feature and a authenticity probability judgment result;
[0051] Based on the discrimination feature, use the L1 loss function to calculate the L1 loss value between the cloud-free image to be evaluated and the real cloud-free image;
[0052] Based on the discrimination feature, use the mean squared error loss function to calculate the mean squared error loss value between the cloud-free image to be evaluated and the real cloud-free image;
[0053] Based on the authenticity probability judgment result, use the Softplus loss function to calculate the Softplus loss value;
[0054] Determine the total loss value based on the L1 loss value, the mean squared error loss value, and the Softplus loss value;
[0055] Update the network parameters of the improved generator and the network parameters of the improved discriminator based on the total loss value.
[0056] Optionally, inputting the to-be-evaluated cloudless image and the real cloudless image into the improved discriminator, and outputting discriminant features and a authenticity probability judgment result, including:
[0057] Extract features from the to-be-evaluated cloudless image or the real cloudless image to obtain a feature map;
[0058] Perform normalization processing on the feature map to obtain a normalization result;
[0059] Activate the normalization result to obtain the discriminant features;
[0060] Perform a convolution operation on the discriminant features to obtain the authenticity probability judgment result.
[0061] Analysis of the beneficial effects of the present invention: Compared with the prior art, a method for removing thin clouds from remote sensing images provided by the present invention improves the problems existing in the existing model, such as insufficient multi-scale feature processing ability, single feature extraction, high computational complexity, and weak generalization ability: integrating a multi-scale attention mechanism to aggregate multi-scale information to enhance the adaptability of the model to different cloud thicknesses and spatial structures, and improving the processing ability of various cloud features; introducing a convolutional block attention module to adaptively enhance key features through concatenating channel and spatial attention, enriching the feature extraction method, while suppressing irrelevant background noise, improving the cloud removal effect and image quality; based on the RICE dataset, adopting a multi-dimensional loss function (such as the L1 loss function, the mean squared error loss function, and the Softplus loss function) to balance image authenticity and detail restoration, optimizing the training process, effectively improving the model performance and generalization ability, and reducing the adverse effects caused by computational complexity and insufficient generalization.
[0062] In a second aspect, the present invention further provides a device, including:
[0063] An acquisition module, configured to acquire a remote sensing image and a trained target generative adversarial network; the remote sensing image includes a cloud-containing image to be processed; the target generative adversarial network includes an improved generator and an improved discriminator; the improved generator at least includes a multi-scale attention module and a convolutional block attention module;
[0064] A generation module, configured to perform image generation processing on the cloud-containing image to be processed by using the trained target generative adversarial network to obtain a final cloudless image. Brief Description of the Drawings
[0065] The drawings described herein are provided to further understand the present invention and form a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0066] Figure 1 One of the flow diagrams of the method for removing thin clouds from remote sensing images provided for an embodiment of the present invention:
[0067] Figure 2 The structural schematic diagram of the target generative adversarial network provided for an embodiment of the present invention;
[0068] Figure 3 Another flow diagram of the method for removing thin clouds from remote sensing images provided for an embodiment of the present invention;
[0069] Figure 4 The structural schematic diagram of the improved generator provided for an embodiment of the present invention;
[0070] Figure 5 Another flow diagram of the method for removing thin clouds from remote sensing images provided for an embodiment of the present invention;
[0071] Figure 6 The structural schematic diagram of the multi-scale attention module provided for an embodiment of the present invention;
[0072] Figure 7 The structural schematic diagram of the convolutional block attention module provided for an embodiment of the present invention;
[0073] Figure 8 The structural schematic diagram of the device for removing thin clouds from remote sensing images provided for an embodiment of the present invention. Detailed Embodiments
[0074] For the convenience of clearly describing the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. For example, the first threshold and the second threshold are only used to distinguish different thresholds and do not limit their sequence. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. do not necessarily limit differences.
[0075] It should be noted that in the present invention, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0076] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist.
[0077] In recent years, deep learning methods, especially convolutional neural networks (CNNs), generative adversarial networks (GANs), and diffusion models, have made significant progress in the task of thin cloud removal. Compared with traditional methods, deep learning methods exhibit stronger automatic learning and generalization capabilities and can extract more complex features from large-scale data. CNNs can automatically learn the features of ground objects under cloud cover, and a variety of cloud removal algorithms have been derived, such as deep residual symmetric connection networks and Slope-Net, etc.
[0078] Due to its ability to generate high-quality images in unsupervised learning, GAN has become one of the important methods for thin cloud removal. A generative adversarial network is a deep learning model that uses adversarial training to complete data distribution modeling and sample generation through the game between a generator and a discriminator. The task of the generator is to generate as realistic samples as possible from the latent space (usually random noise), while the discriminator evaluates the differences between the generated samples and the real samples. During the training process, the generator minimizes the discriminator's ability to recognize the generated samples. Conversely, the discriminator tries to improve its accuracy in judging real and generated samples, forming a min-max optimization problem.
[0079] During the training process of GAN, the generator and the discriminator update their parameters through continuous games. The generator gradually learns how to generate more realistic samples, while the discriminator becomes increasingly proficient at distinguishing real and generated samples. This adversarial training mechanism prompts the generator and the discriminator to promote each other during the training process, and ultimately enables the generator to generate high-quality samples.
[0080] However, traditional GANs still face challenges in the cloud removal task, such as the generated results being vulnerable to artifacts and having limited ability to restore details. To improve the thin cloud removal effect, researchers introduced an attention mechanism to form the SpAGAN model, enhancing the network's ability to capture key information. For example, the prior-based dense attention network proposed by traditional techniques uses the attention mechanism to extract key information, significantly improving the cloud removal effect. However, many of the above methods and models still face problems such as high computational complexity, single feature extraction, lack of multi-scale feature processing ability, and insufficient model generalization ability, resulting in the need for further improvement in the quality of the finally generated cloud-free images.
[0081] As Figure 1 shown, an embodiment of the present invention provides a method for removing thin clouds from remote sensing images, which may include:
[0082] Step 100: Obtain a remote sensing image and a trained target generative adversarial network; the remote sensing image includes a cloud-containing image to be processed; the target generative adversarial network includes an improved generator and an improved discriminator; the improved generator at least includes a multi-scale attention module and a convolutional block attention module;
[0083] Before step 100, the method for removing thin clouds from remote sensing images further includes:
[0084] (1) Construct a target generative adversarial network; refer to Figure 2 , the target generative adversarial network includes the above-mentioned improved generator (abbreviated as the generator in all figures) and the above-mentioned improved discriminator (abbreviated as the discriminator in all figures); the improved generator at least includes the above-mentioned multi-scale attention module and the above-mentioned convolutional block attention module;
[0085] (2) Obtain the RICE1 thin cloud dataset; the RICE1 thin cloud dataset includes a training set and a test set; wherein, the ratio of the training set to the test set can be 8:2; both the training set and the test set include multiple image pairs, and each image pair includes a real cloud-free image and a cloud-containing image;
[0086] (3) Use the training set and combine a multi-dimensional loss function (refer to Figure 2 for Losses) to perform multiple rounds of training on the target generative adversarial network; the multi-dimensional loss function includes: L1 loss function, mean square error loss function, and Softplus loss function;
[0087] Step (3) specifically includes:
[0088] The first step: In each round of training, input the cloud-containing image (refer to Figure 2 ) into the improved generator to obtain an unclouded image to be evaluated;
[0089] Step 2: Input the cloudless image to be evaluated and the true cloudless image into the improved discriminator, and output discriminative features and the authenticity probability judgment result; among them, the improved discriminator includes a CBR module.
[0090] Specifically, Step 2 includes: extracting features from the cloudless image to be evaluated or the true cloudless image to obtain a feature map; performing normalization processing on the feature map to obtain a normalization result; activating the normalization result to obtain discriminative features; performing a convolution operation on the discriminative features to obtain the authenticity probability judgment result.
[0091] Specifically, the cloudless image to be evaluated and the true cloudless image respectively pass through two processing paths. First, the input image is subjected to feature extraction through a CBR (Convolution-BatchNorm-ReLU Module) module, with the initial output channel number being 32, and the Leaky ReLU activation function is used to enhance the non-linear expression ability of the model; secondly, the feature maps of the two branches are fused through a concatenation operation to form 64 feature channels; then, the feature map is further processed through multiple deep convolutional layers, gradually increasing the number of feature channels until reaching 512 channels. Batch normalization is applied after each convolutional layer to improve the stability of training and accelerate the convergence speed; finally, the output layer uses a 3×3 convolution (3×3 CONV) to map the feature map into a single-channel output, representing the authenticity judgment of the input image by the improved discriminator.
[0092] The CBR module integrates convolution, batch normalization, activation functions, and an optional dropout layer. The constructor of this module accepts parameters such as the number of input channels ch0, the number of output channels ch1, whether to use batch normalization bn, the sampling method sample, the activation function activation, and whether to apply dropout. Inside the module, the choice of convolutional layer depends on the sampling method: if it is "down", a standard convolution (Conv2d) is used for downsampling; if it is "up", a transposed convolution (ConvTranspose2d) is used for upsampling. If batch normalization is enabled, the module normalizes the convolutional output, thereby improving the stability and convergence speed of the model. During the forward propagation process, first, the convolution operation is performed, then batch normalization and dropout are applied according to the configuration, and finally, the feature map is output through the activation function. The CBR module is designed to be flexible and efficient and is widely used in deep learning models to enhance feature extraction and representation capabilities. In addition, the improved discriminator supports multi-GPU parallel computing, which improves the processing speed and efficiency. This design significantly improves the model's ability to distinguish between cloud and non-cloud images, thereby providing more reliable discriminative feedback for the GAN.
[0093] Step 3: Based on the discriminative features, use the L1 loss function to calculate the L1 loss value between the cloudless image to be evaluated and the true cloudless image;
[0094] Among them, the L1 loss function, also known as the Mean Absolute Error (MAE), is used to measure the difference between the predicted value and the true value. Its mathematical expression is formula (1):
[0095]
[0096] where Loss L1 is the L1 loss value, N is the number of samples, y i is the true value, is the predicted value.
[0097] In the training of the target GAN, the main role of the L1 loss is to ensure that the image output by the generator is as close as possible to the real image at the pixel level, thereby improving the quality and visual consistency of the generated image. By minimizing the L1 loss, the generator can gradually learn to reconstruct the target image, making its visual effect more similar to the true cloudless image. In this embodiment, by calculating the L1 distance between the generated image (i.e., the cloudless image to be evaluated) and the true cloudless image, we minimize the absolute error to make the generated image close to the real image at the pixel level, ensuring that the generated image is visually consistent with the target image.
[0098] Step 4: Based on the discriminative features, use the mean squared error loss function to calculate the mean squared error loss value between the cloudless image to be evaluated and the true cloudless image;
[0099] Among them, the mean squared error loss function (Mean Squared Error, MSE), also known as the L2 loss, measures the difference between the predicted value and the true value. Its mathematical expression is formula (2):
[0100]
[0101] where Loss MSE is the mean squared error loss value, N is the number of samples, y i is the true value, is the predicted value.
[0102] In the training of the target GAN, the mean squared error loss is commonly used to evaluate the gap between the generated image and the real image. By minimizing the MSE, the improved generator can gradually improve its output, ensuring that the generated image (i.e., the cloudless image to be evaluated) is as close as possible to the target image visually and in terms of content. In this embodiment, the mean squared error loss is mainly used to calculate the difference between the attention map and the real mask, so as to ensure that the improved generator can accurately capture the attention features required in the cloud removal process and enhance the sensitivity to the input features.
[0103] Step 5: Based on the authenticity probability judgment result, use the Softplus loss function to calculate the Softplus loss value;
[0104] Among them, the Softplus loss function is used for the training of the improved discriminator to calculate the probability between the generated fake image (pred_fake) and the real image (pred_real). This loss function is processed through the Softplus activation function, aiming to improve the judgment ability of the improved discriminator for real images and suppress the false scores of the generated images, thereby enhancing the discrimination ability of the improved discriminator. Softplus is a smooth activation function, and its mathematical expression is formula (3):
[0105] Loss sof = ln(1 + e^x) (3)
[0106] Among them, x is the input value (authenticity probability judgment result).
[0107] In the neural network, the Softplus function is commonly used in the hidden layer, especially when it is necessary to avoid the "dead neuron" problem in the ReLU activation function. Due to its smooth characteristics, Softplus is also commonly used in optimization algorithms to improve the convergence speed and model stability.
[0108] Step 6: Based on the L1 loss value, mean squared error loss value, and Softplus loss value, determine the total loss value;
[0109] Specifically, substitute the L1 loss value, mean squared error loss value, and Softplus loss value into formula (4):
[0110] Loss 总 = aLoss L1 + b × Loss MSE + c × Loss sof (4)
[0111] Calculate the total loss value; among them, Loss 总 is the total loss value; a, b, and c are the corresponding weight coefficients.
[0112] Step 7: Based on the total loss value, update the network parameters of the improved generator and the network parameters of the improved discriminator.
[0113] The specific update strategy can refer to related technologies and will not be elaborated here.
[0114] (4) After each round of training, use the test set to evaluate the target generative adversarial network to obtain the evaluation result;
[0115] For example, the evaluation result can be image quality metrics (PSNR, SSIM): Ensure that the image details and spectral information after thin cloud removal are retained.
[0116] For example, the evaluation result can be the visual evaluation of the images generated by the test set: Ensure that the generated cloud-free images are highly consistent with the real cloud-free images in visual perception.
[0117] These above evaluation metrics work together to help determine whether the generator has practical application value, especially in scenarios where remote sensing data has extremely high requirements for spectral and structural accuracy.
[0118] (5) When the evaluation result reaches the preset threshold, complete the training and separate the trained improved generator from the target generative adversarial network that has completed training.
[0119] The specific preset threshold corresponding to each evaluation result can be set in advance according to the actual situation.
[0120] For example, when the evaluation result is image quality metrics (PSNR, SSIM), the preset threshold corresponding to PSNR can be 30, and the preset threshold corresponding to SSIM can be 0.9.
[0121] The preset thresholds corresponding to other evaluation results can refer to related technologies and will not be elaborated one by one in this embodiment.
[0122] Step 200: Use the target generative adversarial network that has completed training to perform image generation processing on the cloud-containing image to be processed, and obtain the final cloud-free image.
[0123] Analysis of the beneficial effects of this embodiment:
[0124] 1) Existing models lack the ability to process multi-scale features. In this embodiment, a multi-scale attention mechanism is integrated. Through the aggregation of multi-scale information, it helps to enhance the adaptability of the model to different cloud thicknesses and spatial structures, and improve its processing ability for various cloud features, which to a certain extent improves the deficiency of traditional models in multi-scale feature processing.
[0125] 2) The existing model has the problem of single feature extraction. The convolutional block attention module introduced in this embodiment realizes the adaptive enhancement of key features by concatenating channel attention and spatial attention, effectively improving the expression ability of the model. This can enrich the way of feature extraction and has obvious improvement compared with the single feature extraction of the diffusion model. At the same time, the convolutional block attention module can also suppress irrelevant background noise, thereby improving the de-clouding effect, enhancing the image quality, and having a positive effect on the situation where the processing result quality of the diffusion model may be low.
[0126] 3) The existing model has problems such as high computational complexity and insufficient model generalization ability. Based on the RICE dataset, this embodiment uses a multi-dimensional loss function (such as L1 loss function, mean square error loss function, and Softplus loss function) to balance the authenticity and detail recovery of the image. In this way, to a certain extent, the training process of the model is optimized, which helps to improve the performance and generalization ability of the model and reduce the adverse effects caused by high computational complexity and insufficient generalization ability.
[0127] Next, refer to Figure 3 , and elaborate on step 200 in detail. The improved generator includes, in addition to the multi-scale attention module and the convolutional block attention module, an input convolutional layer, a residual network, and an output convolutional layer. Step 200 specifically includes:
[0128] Step 210: Input the cloud-containing image to be processed into the input convolutional layer (in the improved generator) to obtain the image after feature extraction;
[0129] For example, refer to Figure 4 , the input convolutional layer includes a 3×3 convolutional layer (3×3CONV in the figure), the activation function ReLU, and multiple bottleneck layers. The generator first performs preliminary feature extraction on the input image through a 3×3 convolutional layer, and then processes the image after preliminary feature extraction through 3 bottleneck layers before entering the multi-scale attention module to obtain the image after feature extraction;
[0130] Step 220: Input the image after feature extraction into the multi-scale attention module to obtain an image containing multi-scale feature information;
[0131] For example, refer to Figure 4 , 3 bottleneck layers (all abbreviated as bottleneck in the attached figure) and the multi-scale attention module cooperate to perform 4 cycles. After the cycle ends, it enters a bottleneck network composed of 5 bottleneck layers for further processing (such as conventional dimension increase and dimension reduction processing), and finally outputs an image containing multi-scale feature information.
[0132] Step 230: input the image containing multiple scale feature information into the convolution block attention module to obtain the image after the channel and spatial dimension are recalibrated;
[0133] Step 240: input the image after the channel and spatial dimension are recalibrated into the residual network to obtain an optimized image;
[0134] Step 250: Input the optimized image to the output convolution layer for processing to obtain the final cloud-free image.
[0135] From the above, we can see that by inputting the cloud-containing image to be processed into the input convolution layer, preliminary feature extraction can be performed on the cloud-containing image. The convolution layer can automatically learn the local features in the image, extract the basic feature information related to thin clouds and ground objects, provide valuable feature representation for subsequent processing, reduce the complexity of image data, reduce the amount of calculation, and retain important image information for further processing by subsequent modules.
[0136] By inputting the feature-extracted image into the multi-scale attention module, an image containing feature information at multiple scales can be obtained. In the thin cloud removal of remote sensing images, thin clouds and objects of different thicknesses and spatial structures have different characteristics at different scales. The multi-scale attention module can enable the model to pay attention to the characteristics of thin clouds and objects at different scales, capture the detailed information of thin clouds and their relationship with surrounding objects, and enhance the model's adaptability to different cloud thicknesses and spatial structures, thereby more comprehensively processing various thin cloud features and improving the effect of thin cloud removal.
[0137] After the image containing multi-scale feature information passes through the convolutional block attention module, the channel and spatial dimension recalibrated image is obtained. This module can realize adaptive enhancement of key features by connecting channel attention and spatial attention in series. In remote sensing images, key features related to thin cloud removal can be highlighted and irrelevant background noise can be suppressed. For example, the features of thin cloud areas can be enhanced and irrelevant noise features in background objects can be suppressed, which effectively improves the expression ability of the model and makes the model more focused on the thin cloud removal task, thereby improving the declouding effect and the quality of cloud-free images.
[0138] The image with recalibrated channels and spatial dimensions is input into the residual network to obtain the optimized image. The residual network can solve the problem of gradient vanishing or gradient exploding that may occur when the network depth increases in deep learning, allowing the network to learn features at a deeper level. In the thin cloud removal task, the residual network can further learn and optimize the features of the image, dig out more complex relationships between thin clouds and ground objects, repair and optimize the details of the image after thin cloud removal, improve the quality and authenticity of the image, and make the final cloud-free image closer to the actual situation.
[0139] The optimized image is input into the output convolutional layer for processing to obtain the final cloud-free image. The output convolutional layer can integrate and transform the features processed through multiple previous steps, map the features back to the image space, and generate the final cloud-free image. It can fine-tune the details of the image, adjust the resolution, number of channels, etc. of the image, so that the generated cloud-free image meets the expected output format and quality requirements, and completes the conversion process from feature representation to the final cloud-free image.
[0140] See Figure 5 , step 220: Input the image after feature extraction into the multi-scale attention module to obtain an image containing multi-scale feature information, including:
[0141] Step 221: Extract the context information of specific levels from the image after feature extraction and output multiple feature vectors;
[0142] In remote sensing images, the distribution and morphology of thin clouds are complex and diverse. The context information at different levels contains the relationships between thin clouds and surrounding ground objects at different scales. Extracting the context information of specific levels can capture the features of thin clouds at different spatial scales, such as the edge details of thin clouds and the relative positions with large-area ground objects. Outputting multiple feature vectors can represent this context information from multiple perspectives, providing a rich data basis for subsequent comprehensive analysis, helping the model to more comprehensively understand the image content, and thus more accurately remove thin clouds.
[0143] The context information at different levels can also reflect features such as the thickness change of thin clouds, which is crucial for accurately identifying and removing thin clouds of different thicknesses, and improves the model's processing ability for complex thin cloud situations.
[0144] Step 222: Perform element-wise addition on the multiple feature vectors to obtain a comprehensive attention weight map;
[0145] The element-wise addition operation fuses the information of multiple feature vectors, highlighting the regions of common concern in different feature vectors while suppressing unimportant regions. In the removal of thin clouds in remote sensing images, this helps the model determine the importance of thin cloud regions and ground object regions in the image. The generated comprehensive attention weight map can more accurately reflect the position and range of thin clouds in the image, enabling the model to more specifically focus on the thin cloud regions in subsequent processing and improving the effect of thin cloud removal.
[0146] The comprehensive attention weight map obtained in this way can integrate information at different scales, comprehensively consider the overall features of thin clouds, and avoid the situation of incomplete removal of thin clouds or misjudging ground objects as thin clouds caused by only focusing on local features.
[0147] Step 223: Perform cross-channel information integration on the comprehensive attention weight map to form an attention mask;
[0148] Remote sensing images usually contain information in multiple channels (such as multi - spectral channels like RGB), and each channel provides different feature information. Cross - channel information integration can make full use of the correlation between these multi - channel information, further optimizing the attention weight map. The formed attention mask can more accurately locate thin cloud regions because it comprehensively considers the features of thin clouds in different channels, enhancing the recognition ability of thin clouds.
[0149] The attention mask can uniformly focus on and process thin cloud regions on different channels, ensuring that the information of each channel can be reasonably utilized and adjusted when removing thin clouds, avoiding image quality problems caused by inconsistent information between channels, and improving the quality and consistency of the finally generated cloud - free image.
[0150] Step 224: Perform per - pixel weighting on the attention mask and the image after feature extraction to obtain an image containing multi - scale feature information.
[0151] The per - pixel weighting operation integrates the information of the attention mask into the image after feature extraction, enabling the model to process each pixel in the image according to the indication of the attention mask. In thin cloud removal, for pixels in thin cloud regions, the model will give more attention and adjustment according to the attention mask, thus more effectively removing thin clouds; for pixels in ground object regions, their original features are maintained as much as possible to avoid over - processing.
[0152] The image obtained in this way, which contains multi - scale feature information, not only includes thin cloud features at different scales but also can perform reasonable weighting on the image according to the attention mask, enabling the model to better utilize these multi - scale feature information in subsequent processing, further improving the accuracy of thin cloud removal and the quality of the image, and generating a more natural and realistic cloud - free remote sensing image.
[0153] In an optional embodiment, referring to Figure 6 the multi - scale attention module in, the dimension of the input feature map is (B, C, H, W), where B represents the batch size, C is the number of channels, and H and W are the height and width of the feature map respectively. Three two - dimensional convolutional layers with different sizes are used to process the input feature map, and the number of output channels after the convolution operation is the same as the number of input channels. These 3 convolutional layers can effectively extract features from different scales and fuse them in subsequent processing. The output of each convolutional layer is used to generate Figure 6The three attention tensors of different scales in [the object] are fused through a summation operation to form a comprehensive attention tensor (i.e., the comprehensive attention weight map). Then, the fused tensor is further processed through a series of 1×1 convolutional layers, aiming to refine and integrate the fused attention information to make it more suitable for the final attention mechanism calculation. Finally, the processed attention tensor is normalized using the Sigmoid activation function, restricting its value within the range of [0,1], thereby generating the final attention map (i.e., the attention mask). The generated final attention map is then multiplied element-wise with the input feature map to achieve weighted modulation of the feature map. This modulation process enables the model to highlight important features and suppress irrelevant features, thereby improving the model's expressiveness and accuracy in processing complex data.
[0154] For example, the convolutional block attention module includes at least a channel attention sub-module and a spatial attention sub-module. Step 230: Input an image containing multi-scale feature information into the convolutional block attention module to obtain an image with re-calibrated channel and spatial dimensions, specifically including:
[0155] Input the image output by the multi-scale attention module into the channel attention sub-module to obtain a channel attention feature map;
[0156] Input the channel attention feature map into the spatial attention sub-module to obtain an image with re-calibrated channel and spatial dimensions.
[0157] Among them, in combination with Figure 7 For understanding, inputting the image output by the multi-scale attention module into the channel attention sub-module to obtain a channel attention feature map includes:
[0158] Step 1: Perform global average pooling on the image containing multi-scale feature information to obtain a first vector;
[0159] This step has the following beneficial effects: Performing global average pooling on the image containing multi-scale feature information can capture the overall spectral intensity features of the channels.
[0160] In remote sensing images, the reflectance of thin clouds in multi-spectral channels (such as red, green, blue, near-infrared) has regularity (such as overall high brightness, abnormal reflectance in the near-infrared channel). Global average pooling can capture the overall intensity features of thin clouds in each channel by calculating the global average value of each channel (for example: the average value of the blue channel in the thin cloud-covered area is usually higher than that in the cloud-free area), providing basic data for channel importance assessment.
[0161] Remote sensing images are vulnerable to local interferences such as sensor noise and ground object textures. Average pooling weakens the impact of local noise on channel features by aggregating spatial information, enabling the model to focus more on the global spectral trends of thin clouds (such as the overall brightness shift of large-scale thin clouds) rather than fragmented noise.
[0162] Step 2: Perform global max pooling on the image containing multi-scale feature information to obtain a second vector;
[0163] This step has the following beneficial effects: capturing extreme spectral features of channels.
[0164] Extremely high reflectance values may appear at the edges of thin clouds or in locally thicker regions in specific channels (such as the red channel). Global max pooling can extract the peak features of each channel, helping the model identify the boundaries of thin clouds or locally dense regions (for example, the transition zone between thin clouds and clear sky shows significant differences after max pooling).
[0165] Average pooling focuses on the overall trend and may overlook local key features (such as small-scale thick cloud patches in thin clouds). Max pooling complements it, ensuring that channel attention includes both global statistical information and retains local extreme spectral signals, avoiding missed judgments of thin cloud details (such as broken cloud blocks).
[0166] Step 3: Use a fully connected layer ( Figure 7 the shared fully connected layer therein) to process the first vector and the second vector to obtain a first weight and a second weight;
[0167] This step has the following beneficial effects: learning non-linear dependencies between channels.
[0168] The removal of thin clouds from remote sensing images relies on the collaboration of multi-spectral information (such as using the reflectance difference between the near-infrared and visible light channels to distinguish thin clouds from ground objects). Through non-linear transformation (such as ReLU activation), the fully connected layer can learn the complex dependencies between different channels (for example, the combination of high values in the near-infrared channel and high values in the blue channel is more likely to correspond to thin cloud regions), generating channel-specific weights.
[0169] Regarding the spectral variations of thin clouds in different scenarios (such as the infrared characteristics of thin clouds in high-latitude regions being different from those in low-latitudes), the fully connected layer can adaptively adjust the weights, enhancing the channels related to thin clouds in the current scenario (such as enhancing the channels significantly affected by thin clouds), suppressing irrelevant channels (such as the thermal infrared channel with more noise), and improving the model's adaptability to complex spectral conditions.
[0170] Step 4: Add and activate the first weight and the second weight to obtain the target channel attention weight (which is also the Figure 7 channel attention mask in
[0171] This step has the following beneficial effects: adding the global trend (first weight) of average pooling and the local extreme value (second weight) of max pooling enables the channel attention weight to contain both the "overall distribution of thin clouds" and the "local significant features", avoiding information loss caused by single pooling. For example, it can identify both the overall increase in brightness of large-scale thin clouds and capture the high reflection peaks of local thick clouds.
[0172] Compress the weight to [0, 1] through activation functions such as Sigmoid to ensure the interpretability of the weight (the closer the weight is to 1, the higher the channel importance). In thin cloud removal, this operation can accurately suppress noise channels (such as short-wave infrared channels with low signal-to-noise ratio), while enhancing thin cloud sensitive channels (such as blue light and near-infrared channels), avoiding spectral distortion (such as abnormal ground object colors after cloud removal).
[0173] Step 5: Multiply the target channel attention weight by the image containing multi-scale feature information to obtain the channel attention feature map.
[0174] This step has the following beneficial effects: achieving adaptive enhancement and suppression of spectral features.
[0175] Assign high weights to thin cloud sensitive channels (such as near-infrared channels) to enhance their feature expressions and help the model more accurately detect the spectral anomalies of thin clouds (such as increased near-infrared reflectance); maintain reasonable weights for channels rich in surface features (such as green light channels) to avoid over-weakening the spectral information of ground objects during cloud removal, such as the chlorophyll absorption characteristics of vegetation.
[0176] Assign low weights to channels severely affected by atmospheric scattering or irrelevant to thin clouds (such as certain low-frequency noise channels) to reduce their interference with the model's decision-making and lower the misjudgment rate of thin clouds (such as avoiding misidentifying high-reflectance ground objects as thin clouds).
[0177] Generally speaking, each step of the channel attention sub-module is designed according to the multi-spectral characteristics of remote sensing images and the spectral distribution law of thin clouds. Through "global-local" pooling fusion, channel dependence learning, weight normalization, and feature weighting, it realizes the enhancement of thin cloud sensitive channels and the protection of ground object channels, provides purer and more discriminative spectral features for the subsequent spatial attention sub-module, and ultimately improves the accuracy and spectral authenticity of thin cloud removal.
[0178] Combined with Figure 7 For understanding, input the channel attention feature map into the spatial attention sub-module to obtain the image with re-calibrated channel and spatial dimensions, including:
[0179] Step 1: Perform average pooling and max pooling on the channel attention feature map for each channel to obtain the pooling results;
[0180] This step has the following beneficial effects: The thin clouds in the remote sensing image are characterized by uneven spatial distribution (such as thin cloud coverage in local areas), and the differences in ground object details (such as texture and edges) in the spatial dimension are the key to distinguishing thin clouds from real ground objects. Average pooling captures the global average information of the pixels within the channel, reflecting the overall brightness / radiance mean of the thin clouds in space (thin cloud areas are generally brighter as a whole). Max pooling retains the extreme value information of the pixels within the channel, highlighting high-frequency features such as the edges and textures of ground objects (such as building outlines and vegetation boundaries under thin cloud coverage).
[0181] Step 2: Concatenate the pooling results along the channel dimension to obtain the concatenated feature map;
[0182] This step has the following beneficial effects: Concatenating the results of average pooling and max pooling forms cross-channel spatial feature fusion, retaining the comprehensive responses of each spatial position under different pooling methods, providing richer input features for subsequent convolutional operations, enabling the model to perceive the dual differences in "overall brightness" and "local details" at the same spatial position, and better distinguishing thin cloud-covered areas from real ground objects.
[0183] Step 3: Perform a convolutional operation on the concatenated feature map to obtain the initial spatial attention weights;
[0184] This step has the following beneficial effects: The convolutional operation captures the dependency relationships of spatially adjacent pixels through local receptive fields, generating the attention weights for each spatial position. For the thin cloud removal task, a weight pattern of "suppressing thin cloud areas with continuous high brightness and blurred details and enhancing ground object areas with clear textures / edges" can be learned to achieve spatial localization and suppression of thin cloud-covered areas.
[0185] Step 4: Activate the initial spatial attention weights to obtain the target spatial attention weights (i.e., the spatial attention mask in Figure 7 , which can also be called the spatial attention mask);
[0186] This step has the following beneficial effects: The activation function normalizes the initial weights into the [0,1] interval, generating an adaptive spatial attention mask. For remote sensing images, the pixel values in thin cloud areas usually lie between those of real ground objects and thick clouds, with blurred boundaries. The activated weights can finely adjust the importance of each pixel: assigning low weights (suppression) to low-texture areas covered by thin clouds and high weights (enhancement) to areas with rich ground object details; avoiding the problem of gradient disappearance and enabling the model to learn the distribution law of thin clouds in the spatial dimension end-to-end.
[0187] Step 5: Multiply the target spatial attention weights by the channel attention feature map to obtain the image with re-calibrated channel and spatial dimensions.
[0188] This step has the following beneficial effects: By combining channel attention and spatial attention, dual feature calibration is achieved. Channel attention has weighted the importance of each spectral channel (e.g., enhancing the infrared channel less affected by thin clouds and suppressing the visible light channel more affected); spatial attention further weights the pixel positions within each channel (e.g., suppressing the highlighted areas covered by thin clouds in the visible light channel and retaining the true spectral signals of ground objects). The final result can more accurately remove the "spatial non-uniform noise" of thin clouds while retaining the spectral features of ground objects (such as the red edge effect of vegetation) and spatial structures (such as the geometric shape of buildings), avoiding common problems such as "over-smoothing" or "detail loss" in traditional methods.
[0189] Generally speaking, the spatial attention sub-module solves the core challenges of "spatial non-uniformity" and "ground object detail retention" in thin cloud removal from remote sensing images through multi-pooling fusion, local feature modeling, and adaptive weight generation, enabling the model to accurately distinguish thin clouds from real ground objects at the pixel level and ultimately improving the spectral fidelity and spatial resolution of the final cloud-free image.
[0190] Finally, conducting research on cloud removal methods for remote sensing images is crucial for improving data utilization, enhancing image quality, and promoting time series analysis and dynamic monitoring. This study deeply explores the effectiveness of a generative adversarial network (GAN) that combines convolutional block attention mechanism and multi-scale attention mechanism in thin cloud removal and compares it with HazeRemoval, FFA-Net, CGAN, and SpAGAN.
[0191] Table 1: Comparison of quantitative results of different models on the RICE1 dataset
[0192]
[0193]
[0194] The experimental results show that combining the convolutional block attention mechanism and the multi-scale attention mechanism significantly improves the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) of the image, reaching 31.321 and 0.894 respectively, which is better than other methods. This improvement reflects the advantages of the model in retaining detail and structural information and shows its adaptability and robustness in complex image scenes. Ablation experiments further prove the unique contributions of each attention mechanism and their synergistic effects, providing theoretical support for future attention mechanism-based image processing technologies. Therefore, the present invention provides a new idea for the development of thin cloud removal technology and is expected to be widely applied in the field of remote sensing image processing.
[0195] The device provided by the present invention will be described below, and the device described below can be correspondingly referred to the method described above.
[0196] As Figure 8 shown, an embodiment of the present invention further provides a thin cloud removal device for remote sensing images, which is used to implement the method in any of the above embodiments. The thin cloud removal device for remote sensing images may include:
[0197] An acquisition module 810, configured to acquire a remote sensing image and a trained target generative adversarial network; the remote sensing image includes a cloud-containing image to be processed; the improved generator is optimized based on a pre-constructed target generative adversarial network and in combination with a multi-dimensional loss function during training, and the target generative adversarial network includes an improved generator and an improved discriminator; the target generative adversarial network includes an improved generator and an improved discriminator; the improved generator at least includes a multi-scale attention module and a convolutional block attention module;
[0198] A generation module 820, configured to perform image generation processing on the cloud-containing image to be processed by using the trained target generative adversarial network to obtain a final cloud-free image.
[0199] Optionally, the improved generator further includes an input convolutional layer, a residual network, and an output convolutional layer. The generation module 820 is specifically configured to:
[0200] Input the cloud-containing image to be processed into the input convolutional layer to obtain an image after feature extraction;
[0201] Input the image after feature extraction into the multi-scale attention module to obtain an image containing multi-scale feature information;
[0202] Input the image containing multi-scale feature information into the convolutional block attention module to obtain an image with recalibrated channel and spatial dimensions;
[0203] Input the image with recalibrated channel and spatial dimensions into the residual network to obtain an optimized image;
[0204] Input the optimized image into the output convolutional layer for processing to obtain a final cloud-free image.
[0205] The generation module 820 is specifically configured to:
[0206] Extract specific-level context information from the image after feature extraction and output multiple feature vectors;
[0207] Perform element-wise addition on the multiple feature vectors to obtain a comprehensive attention weight map;
[0208] Perform cross-channel information integration on the comprehensive attention weight map to form an attention mask;
[0209] Perform pixel-wise weighting of the attention mask and the image after feature extraction to obtain an image containing multi-scale feature information.
[0210] The generation module 820 is specifically configured to:
[0211] Input the image output by the multi-scale attention module into the channel attention sub-module to obtain a channel attention feature map;
[0212] Input the channel attention feature map into the spatial attention sub-module to obtain an image with recalibrated channel and spatial dimensions.
[0213] The generation module 820 is specifically configured to:
[0214] Perform global average pooling on the image containing multi-scale feature information to obtain a first vector;
[0215] Perform global max pooling on the image containing multi-scale feature information to obtain a second vector;
[0216] Use a fully connected layer to process the first vector and the second vector to obtain a first weight and a second weight;
[0217] Add and activate the first weight and the second weight to obtain a target channel attention weight;
[0218] Multiply the target channel attention weight by the image containing multi-scale feature information to obtain a channel attention feature map.
[0219] The generation module 820 is specifically configured to:
[0220] Perform average pooling and max pooling on each channel of the channel attention feature map to obtain a pooling result;
[0221] Concatenate the pooling results along the channel dimension to obtain a concatenated feature map;
[0222] Perform a convolution operation on the concatenated feature map to obtain an initial spatial attention weight;
[0223] Activate the initial spatial attention weight to obtain a target spatial attention weight;
[0224] Multiply the target spatial attention weight by the channel attention feature map to obtain an image with recalibrated channel and spatial dimensions.
[0225] The remote sensing image thin cloud removal device further includes a construction training module for:
[0226] Obtain the RICE1 thin cloud dataset; the RICE1 thin cloud dataset includes a training set and a test set; both the training set and the test set include multiple image pairs, and each image pair includes a real cloud-free image and a cloud-containing image;
[0227] Use the training set to perform multiple rounds of training on the target generative adversarial network in combination with a multi-dimensional loss function; the multi-dimensional loss function includes: L1 loss function, mean squared error loss function, and Softplus loss function;
[0228] After each round of training, use the test set to evaluate the target generative adversarial network to obtain an evaluation result;
[0229] When the evaluation result reaches the preset threshold, the training is completed, and the trained improved generator is separated from the target generative adversarial network that has completed training.
[0230] Construct a training module, specifically used for:
[0231] In each round of training, input the cloud-containing image into the improved generator to obtain the cloud-free image to be evaluated;
[0232] Input the cloud-free image to be evaluated and the real cloud-free image into the improved discriminator, and output the discriminant feature and the authenticity probability judgment result;
[0233] Based on the discriminant feature, use the L1 loss function to calculate the L1 loss value between the cloud-free image to be evaluated and the real cloud-free image;
[0234] Based on the discriminant feature, use the mean squared error loss function to calculate the mean squared error loss value between the cloud-free image to be evaluated and the real cloud-free image;
[0235] Based on the authenticity probability judgment result, use the Softplus loss function to calculate the Softplus loss value;
[0236] Based on the L1 loss value, mean squared error loss value, and Softplus loss value, determine the total loss value;
[0237] Based on the total loss value, update the network parameters of the improved generator and the network parameters of the improved discriminator.
[0238] Construct a training module, specifically used for:
[0239] Extract features from the cloud-free image to be evaluated or the real cloud-free image to obtain a feature map;
[0240] Perform normalization processing on the feature map to obtain a normalization result;
[0241] Activate the normalization result to obtain the discriminant feature;
[0242] Perform a convolution operation on the discriminant feature to obtain the authenticity probability judgment result.
[0243] An embodiment of the present invention further provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus. A computer program that can be run by the processor is stored on the memory; when the processor runs the computer program, it can execute the thin cloud removal method for remote sensing images in any of the above embodiments.
[0244] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs that can store program codes.
[0245] On the other hand, the present invention also provides a non-transitory computer-readable storage medium. Instructions are stored in the computer storage medium, and when the instructions are run, the thin cloud removal method for remote sensing images in any of the above embodiments is implemented.
[0246] Although the present invention is described in combination with various embodiments herein, however, in the process of implementing the claimed present invention, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the drawings, the disclosure content, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0247] Although the present invention is described in combination with specific features and their embodiments, obviously, various modifications and combinations can be made without departing from the spirit and scope of the present invention. Accordingly, this specification and the drawings are only exemplary descriptions of the present invention defined by the appended claims, and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present invention. Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A method for removing thin clouds from remote sensing images, characterized in that, Including: Obtaining remote sensing images and a trained target generative adversarial network; The remote sensing images include cloud-containing images to be processed; The target generative adversarial network includes an improved generator and an improved discriminator; the improved generator at least includes a multi-scale attention module and a convolutional block attention module; Using the trained target generative adversarial network to perform image generation processing on the cloud-containing image to be processed to obtain a final cloud-free image.
2. The method for removing thin clouds from remote sensing images according to claim 1, wherein, The improved generator further includes an input convolutional layer, a residual network, and an output convolutional layer; The using the trained target generative adversarial network to perform image generation processing on the cloud-containing image to be processed to obtain a final cloud-free image at least includes: Inputting the cloud-containing image to be processed into the input convolutional layer to obtain an image after feature extraction; Inputting the image after feature extraction into the multi-scale attention module to obtain an image containing multi-scale feature information; Inputting the image containing multi-scale feature information into the convolutional block attention module to obtain an image with re-calibrated channel and spatial dimensions; Inputting the image with re-calibrated channel and spatial dimensions into the residual network to obtain an optimized image; Inputting the optimized image into the output convolutional layer for processing to obtain the final cloud-free image.
3. The method for removing thin clouds from remote sensing images according to claim 2, characterized in that The inputting the image after feature extraction into the multi-scale attention module to obtain an image containing multi-scale feature information includes: Extracting context information at a specific level from the image after feature extraction and outputting multiple feature vectors; Performing element-wise addition on the multiple feature vectors to obtain a comprehensive attention weight map; Performing cross-channel information integration on the comprehensive attention weight map to form an attention mask; Performing pixel-wise weighting of the attention mask and the image after feature extraction to obtain the image containing multi-scale feature information.
4. The method for removing thin clouds from remote sensing images according to claim 2, wherein, The convolutional block attention module at least includes a channel attention sub-module and a spatial attention sub-module; The inputting the image containing multi-scale feature information into the convolutional block attention module to obtain an image with re-calibrated channel and spatial dimensions includes: Inputting the image output by the multi-scale attention module into the channel attention sub-module to obtain a channel attention feature map; Inputting the channel attention feature map into the spatial attention sub-module to obtain the image with re-calibrated channel and spatial dimensions.
5. The method for removing thin clouds from remote sensing images according to claim 4, wherein The inputting the image output by the multi-scale attention module into the channel attention sub-module to obtain a channel attention feature map includes: Performing global average pooling on the image containing multi-scale feature information to obtain a first vector; Performing global max pooling on the image containing multi-scale feature information to obtain a second vector; Using a fully connected layer to process the first vector and the second vector to obtain a first weight and a second weight; Adding and activating the first weight and the second weight to obtain a target channel attention weight; Multiplying the target channel attention weight by the image containing multi-scale feature information to obtain the channel attention feature map.
6. The method for removing thin clouds from remote sensing images according to claim 4, wherein Inputting the channel attention feature map into the spatial attention sub-module to obtain the image with re-calibrated channel and spatial dimensions includes: Performing average pooling and max pooling on the channel attention feature map for each channel to obtain pooling results; Concatenating the pooling results along the channel dimension to obtain a concatenated feature map; Performing a convolution operation on the concatenated feature map to obtain an initial spatial attention weight; Activating the initial spatial attention weight to obtain a target spatial attention weight; Multiplying the target spatial attention weight by the channel attention feature map to obtain the image with re-calibrated channel and spatial dimensions.
7. The method for removing thin clouds from remote sensing images according to claim 1, wherein Before obtaining the remote sensing image and the trained target generative adversarial network, the method further includes: Constructing the target generative adversarial network; the target generative adversarial network includes an improved generator and the improved discriminator; the improved generator at least includes the multi-scale attention module and the convolutional block attention module; Obtaining the RICE1 thin cloud dataset; the RICE1 thin cloud dataset includes a training set and a test set; Using the training set, and combining with a multi-dimensional loss function to perform multiple rounds of training on the target generative adversarial network; the multi-dimensional loss function includes: L1 loss function, mean square error loss function, and Softplus loss function; After each round of training, using the test set to evaluate the target generative adversarial network to obtain an evaluation result; When the evaluation result reaches a preset threshold, the training is completed.
8. The method for removing thin clouds from remote sensing images according to claim 7, characterized in that, Both the training set and the test set include multiple image pairs, and each image pair includes a real cloud-free image and a cloud-containing image; The using the training set, and combining with a multi-dimensional loss function to perform multiple rounds of training on the target generative adversarial network includes: In each round of training, inputting the cloud-containing image into the improved generator to obtain a cloud-free image to be evaluated; Inputting the cloud-free image to be evaluated and the real cloud-free image into the improved discriminator, and outputting discriminative features and a authenticity probability judgment result; Based on the discriminative features, using the L1 loss function to calculate the L1 loss value between the cloud-free image to be evaluated and the real cloud-free image; Based on the discriminative features, using the mean square error loss function to calculate the mean square error loss value between the cloud-free image to be evaluated and the real cloud-free image; Based on the authenticity probability judgment result, using the Softplus loss function to calculate the Softplus loss value; Based on the L1 loss value, the mean square error loss value, and the Softplus loss value, determining the total loss value; Based on the total loss value, updating the network parameters of the improved generator and the network parameters of the improved discriminator.
9. The method for removing thin clouds from remote sensing images according to claim 8, characterized in that, The inputting the cloud-free image to be evaluated and the real cloud-free image into the improved discriminator, and outputting discriminative features and a authenticity probability judgment result includes: Performing feature extraction on the cloud-free image to be evaluated or the real cloud-free image to obtain a feature map; Performing normalization processing on the feature map to obtain a normalization result; Activate the normalization result to obtain the discriminant feature; Perform a convolution operation on the discriminant feature to obtain the authenticity probability judgment result.
10. A thin cloud removal device for remote sensing images, characterized in that, It includes: An acquisition module for acquiring a remote sensing image and a trained target generative adversarial network; The remote sensing image includes a cloud-containing image to be processed; The target generative adversarial network includes an improved generator and an improved discriminator; the improved generator at least includes a multi-scale attention module and a convolutional block attention module; A generation module for performing image generation processing on the cloud-containing image to be processed by using the trained target generative adversarial network to obtain a final cloud-free image.
Citation Information
Patent Citations
Remote sensing image cloud removal method and system
CN116823664A
Remote sensing image space-time fusion method based on generative adversarial network
CN119314007A
Cited By
Satellite landslide data generation method and device based on multi-scale adversarial neural network, equipment and medium
CN121837441A
Satellite landslide data generation method, device, equipment, and medium based on multi-scale adversarial neural network.
CN121837441B