Degradation-aware and semantic-guided generative thin cloud remote sensing image de-clouding method
A generative thin cloud remote sensing image declouding method guided by multi-scale degradation spatial feature extraction and cloud pollution semantic features solves the problems of detail loss and texture blurring in thin cloud image reconstruction, and realizes the generation of high-fidelity, semantically accurate cloudless remote sensing images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2026-03-18
- Publication Date
- 2026-06-02
AI Technical Summary
Existing methods for removing clouds from thin cloud remote sensing images are ineffective at capturing the complex degradation patterns of thin cloud images at different spatial scales, resulting in loss of detail, blurred texture, and semantic shift in the reconstructed images.
A generative thin-cloud remote sensing image declouding method based on degradation perception and semantic guidance is adopted. The generative thin-cloud remote sensing image declouding model is constructed through a multi-scale degradation spatial feature extraction module and a cloud pollution semantic feature guidance module. The thin-cloud and cloudless images in the training set are used to train the data, and the U-Net network is used for prediction and post-processing to generate high-fidelity cloudless remote sensing images.
It effectively captures degradation information at different cloud scales, enhances the model's ability to preserve the structure and details of ground features under clouds, ensures the accuracy of generated images in terms of semantic consistency and ground feature classification, and improves radiometric fidelity.
Smart Images

Figure CN122134577A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud removal technology for thin cloud remote sensing images, and in particular to a generative method for cloud removal of thin cloud remote sensing images based on degradation perception and semantic guidance. Background Technology
[0002] With the improvement of Earth observation capabilities, remote sensing imagery plays an important role in large-scale, high-frequency natural resource monitoring. However, clouds alter the spectral characteristics of surface reflectance through atmospheric scattering and cause spatial discontinuities in images due to their anisotropic texture distribution, severely affecting the quality of remote sensing imaging. Cloud morphology is generally classified into thin clouds and thick clouds based on different optical thicknesses; thick clouds almost completely block surface radiation signals, rendering the image unusable; thin clouds, on the other hand, have some transmissivity and can partially recover surface information. Traditional cloud removal methods rely on prior assumptions or specific sensor parameters, making it difficult to cope with the complex and varied cloud contamination scenarios in high-resolution imagery, and easily causing distortion of ground features or spectral aberration.
[0003] With the development of deep learning, data-driven declouding methods have demonstrated powerful feature learning capabilities and have become the mainstream approach for declouding. However, the complexity of cloud morphology and sub-cloud scenes in high-resolution imagery still poses a challenge to high-quality declouding reconstruction. Different deep learning models need to be designed to address the different characteristics of thin and thick clouds. The non-uniform spatial distribution of thin clouds significantly increases the complexity of cloud pollution degradation modeling. Specifically, the non-uniform spatial distribution of thin clouds makes it difficult for models to effectively capture and represent the complex degradation patterns of thin cloud images at different spatial scales, thus hindering the accurate extraction of multi-scale ground feature structures and leading to fidelity issues such as local detail loss and texture blurring in the reconstructed image. Simultaneously, the occlusion effect of thin clouds causes the loss of important scene radiation information, which further shifts the semantic features of ground features in the reconstructed image, affecting the accuracy of subsequent downstream tasks.
[0004] Therefore, it is necessary to provide a generative thin cloud remote sensing image declouding method based on degradation perception and semantic guidance to solve the problems of loss of detail, texture blurring and semantic shift in reconstructed images caused by non-uniform thin cloud coverage. Summary of the Invention
[0005] The purpose of this invention is to provide a generative thin cloud remote sensing image declouding method based on degradation perception and semantic guidance. The specific technical solution is as follows: De-clouding methods for generative thin-cloud remote sensing images based on degradation perception and semantic guidance include: Step S1: Obtain thin cloud-cloud-free image pairs of the target area and divide them into training set and test set according to the ratio; Step S2: Use the thin cloud-cloud-free images in the training set to train the data to obtain a generative thin cloud remote sensing image cloud removal model based on degradation perception and semantic guidance. Step S3: Input the thin cloud-cloud-free image pairs from the test set into the trained cloud removal model for prediction to obtain the latent space features of the test set. Then, post-process the latent space features of the test set to obtain the remote sensing image after removing the thin clouds.
[0006] Optionally, in step S2, the step of training the cloud removal model includes: Step S2.1: Select the cloudless image data from the thin cloud-cloudless image pair data in the training set. y After being processed by the VAE encoder in the multi-scale degraded spatial feature extraction module, cloudless latent spatial features are obtained. Z y ; will the Z y After being processed in the U-Net network, each time step is obtained. t Conditional latent space features of noisy images and by each Determine the intermediate features of each layer of the corresponding U-Net network ;in, c Indicates cloudless image conditions; This represents the number of layers in the U-Net network at a certain spatial scale. This indicates the generation process of the U-Net network during diffusion denoising; The thin cloud image data in the thin cloud-clear image pair data of the training set. x After being processed by the VAE encoder in the multi-scale degradation spatial feature extraction module, the thin cloud hidden spatial features are obtained. Z x ; will the Z x With time step t After being processed by the multi-scale conditional feature encoder of the multi-scale degradation space feature extraction module, the multi-scale conditional features are obtained. ;in, N This represents the total number of spatial scale levels in the U-Net network; cond Indicates thin cloud image conditions; Represents the number of levels at a certain spatial scale in the U-Net network. i Corresponding scale-conditional features; Step S2.2: Extract the scale-conditional features from each of the aforementioned features using convolutional layers. Predict a set of modulation parameter scaling factors With bias term , will the The above With the After being input into the U-Net network and modulated, modulation features are obtained. ; Step S2.3: Input the thin cloud image data from the thin cloud-cloud-free image pairs in the training set into the RAM encoder of the cloud pollution semantic feature guidance module for processing to obtain cross-modal semantic features. C R ; will the C R With noise characteristics Z Cross-attention is obtained through combined computation Z R ; wherein, the noise characteristics Z By the modulation features Transformed to obtain; the said Z R The transformation yields updated modulation features. ; Step S2.4, the... The above Z R and stated Substitute into the loss function L G In the process, the loss value is calculated. If the loss value is greater than the target threshold, the loss value is backpropagated to update the parameters of the entire U-Net network model. If the loss value is less than or equal to the target threshold, the trained cloud removal model is obtained.
[0007] Optionally, in step S2.4, the backpropagation method includes: The chain rule is used to calculate the gradient of the loss function with respect to the parameters of each layer of the U-Net network. The weights of the U-Net network are then updated using these gradients. Subsequently, steps S2.1 to S2.3 are re-run using the updated U-Net network to obtain the updated... , Z R and Substitute into the loss function L G In the middle, the loss value is calculated.
[0008] Optionally, in step S2.2, the modulation process is obtained by combining Equation 1) and Equation 2); Equations 1) and 2) are as follows: Formula 1); Equation 2); In equation 1), This represents a convolutional neural network with learnable parameters; In equation 2), and Each of the above The mean and standard deviation; The multiplication symbol is used to represent multiplication.
[0009] Optionally, in step S2.3, cross-attention is obtained using equation 3). Z R ; where Equation 3) is as follows: Equation 3); In equation 3), ; Indicates the query conversion parameters; ; Indicates index conversion parameters; ; Indicates content conversion parameters; T Represents a matrix Perform a transpose operation; d This represents the scaling factor.
[0010] Optionally, in step S2.3, the noise characteristics Z The modulation features are ordered in two dimensions. The modulation features are obtained by arranging them into a one-dimensional sequence along the column direction; the updated modulation features are obtained by transformation. It is described by a one-dimensional sequence. Z R It is obtained by transforming it back into a two-dimensional sequence.
[0011] Optionally, in step S2.4, the loss function L G Equation 4) is used to represent: Equation 4); In equation 4), This indicates a VAE encoder with frozen parameters; This represents a multi-scale conditional feature encoder; This represents the training parameters in a multi-scale conditional feature encoder. This represents a cloud-aware cross-modal semantic extractor; R This represents the training parameters in the cloud-aware cross-modal semantic extractor; This represents a parameterized denoising U-Net; This represents noise that follows a standard normal distribution and is randomly sampled. express Extracted modulation features ; express Extracted cross attention ZR ; This represents the average value of the prediction error of the cloud removal model; This represents a standard normal distribution.
[0012] Optionally, in step S1, the method for acquiring the thin cloud-cloud-free image pair data includes: Step S1.1: Perform median filtering on the raw optical images of the target area to obtain preprocessed data; Step S1.2: Perform registration processing on the preprocessed data to obtain thin cloud-cloud-free image pairs.
[0013] Optionally, in step S1.2, the registration process includes: First, local SIFT feature points are extracted in the non-cloud-covered areas of the image pairs to be registered in the preprocessed data for registration between thin cloud and cloudless images. By calculating the Euclidean distance between the SIFT feature points of the cloudless image and the thin cloud image, the matching point with the smallest distance to the corresponding position in the thin cloud image is found for each feature point in the cloudless image, thus achieving the matching of the same feature points; wherein, the total number of extracted local SIFT feature points is not less than 50. Secondly, by using quadratic polynomial correction, the pixel coordinates of SIFT feature points on the thin cloud image are mapped to the coordinates of the corresponding SIFT feature points on the cloudless image, thus achieving registration between the thin cloud and cloudless images. The quadratic polynomial is represented by equations 5) and 6): Equation 5); Formula 6); In equation 5), Represents thin cloud image data; and All are transformation parameters of Equation 5); Represents the pixel coordinates of SIFT feature points on thin cloud images ; In equation 6), y This indicates cloudless imagery data; and All are transformation parameters of Equation 6); Finally, the least squares method is used to solve for the transformation parameters. and The root mean square error (RMSE) is used to verify whether the registration accuracy between thin cloud and cloudless images is controlled within 0.5 pixels. If the RMSE is greater than 0.5 pixels, the residuals of each SIFT feature point are calculated, the point with the largest residual is removed, and then the least squares method is used again to solve for the transformation parameters. and This continues until the root mean square error is within 0.5 pixels.
[0014] Optionally, in step S1, the thin cloud-cloudless image pair data is divided into a training set and a test set in a ratio of 6:4 to 9:1. In step S2.4, the target threshold value is 2; In step S3, the post-processing involves using a VAE decoder to process the latent space features in the test set.
[0015] The application of the technical solution of the present invention has at least the following beneficial effects: (1) The present invention provides a generative thin cloud remote sensing image declouding method based on degradation perception and semantic guidance, which can solve the problems of loss of detail, blurred texture and semantic shift in reconstructed images caused by non-uniform coverage of thin clouds. Specifically, the present invention uses thin cloud-free images in the training set to train a generative thin cloud remote sensing image declouding model based on degradation perception and semantic guidance. By constructing a multi-scale degradation spatial feature extraction mechanism, different scale spatial degradation features of thin cloud images are extracted and injected during the diffusion generation process, effectively capturing degradation information at different scales of the cloud layer, enhancing the model's ability to preserve the structure and details of ground objects under the cloud, and ensuring clear texture. In addition, the model introduces a cloud pollution semantic feature guidance mechanism to achieve precise guidance at the semantic level, thereby ensuring the accuracy of the generated image in terms of semantic consistency and ground object category identification. Furthermore, the present invention inputs the thin cloud-free images in the test set into the trained declouding model for prediction, obtains the latent spatial features of the test set, and then processes the latent spatial features of the test set to obtain a high-fidelity thin cloud removal and detail reconstruction remote sensing image.
[0016] (2) The present invention introduces a multi-scale degradation spatial feature extraction module and a cloud pollution semantic feature guidance module to coordinate and optimize the training strategy, and obtains a well-trained cloud removal model, which can improve the radiation fidelity while ensuring structural consistency. Finally, a VAE decoder is used to generate cloudless remote sensing images with clear details, accurate semantics, and consistent radiation.
[0017] In addition to the objectives, features, and advantages described above, the present invention has other objectives, features, and advantages. The invention will now be described in further detail with reference to the figures. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the generative thin cloud remote sensing image declouding method based on degradation perception and semantic guidance in an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0021] Example: See Figure 1 A generative thin-cloud remote sensing image declouding method based on degradation perception and semantic guidance includes: Step S1: Obtain thin cloud-cloud-free image pairs of the target area (specifically, a selected target city) and divide them into training and test sets at a ratio of 8:2. Step S2: Use the thin cloud-cloud-free images in the training set to train the data to obtain a generative thin cloud remote sensing image cloud removal model based on degradation perception and semantic guidance. Step S3: Input the thin cloud-free image pairs from the test set into the trained cloud removal model for prediction to obtain the latent space features of the test set. Then, process the latent space features of the test set (i.e., process them with a VAE decoder) to obtain the remote sensing image after removing the thin clouds.
[0022] In step S2, the steps for training the cloud removal model include: Step S2.1: Select the cloudless image data from the thin cloud-cloudless image pair data in the training set. y After being processed by the VAE encoder in the multi-scale degraded spatial feature extraction module, cloudless latent spatial features are obtained. Z y ; will the Z y After being processed in the U-Net network, each time step is obtained. t Conditional latent space features of noisy images and by each Determine the intermediate features of each layer of the corresponding U-Net network ;in, c Indicates cloudless image conditions; This represents the number of layers in the U-Net network at a certain spatial scale. This indicates the generation process of the U-Net network during diffusion denoising; The thin cloud image data in the thin cloud-clear image pair data of the training set. x After being processed by the VAE encoder in the multi-scale degradation spatial feature extraction module, the thin cloud hidden spatial features are obtained. Z x ; will the Z x With time step t After being processed by the multi-scale conditional feature encoder of the multi-scale degradation space feature extraction module, the multi-scale conditional features are obtained. ;in, N This represents the total number of spatial scale levels in the U-Net network; cond Indicates thin cloud image conditions; Represents the number of levels at a certain spatial scale in the U-Net network. i Corresponding scale-conditional features; Step S2.2: Extract the scale-conditional features from each of the aforementioned features using convolutional layers. Predict a set of modulation parameter scaling factors With bias term , will the The above With the After being input into the U-Net network and modulated, modulation features are obtained. ; Step S2.3: Input the thin cloud image data from the thin cloud-cloud-free image pairs in the training set into the RAM encoder of the cloud pollution semantic feature guidance module for processing to obtain cross-modal semantic features. C R ; will the C R With noise characteristics Z Cross-attention is obtained through combined computation Z R ; wherein, the noise characteristics Z By the modulation features Transformed to obtain; the said Z R The transformation yields updated modulation features. ; Step S2.4, the... The above Z R and stated Substitute into the loss function L G In the process, the loss value is calculated. If the loss value is greater than the target threshold, the loss value is backpropagated to update the parameters of the entire U-Net network model. If the loss value is less than or equal to the target threshold, the trained cloud removal model is obtained. The target threshold is 2.
[0023] In step S2.4, the backpropagation method includes: The chain rule is used to calculate the gradient of the loss function with respect to the parameters of each layer of the U-Net network. The weights of the U-Net network are then updated using these gradients. Subsequently, steps S2.1 to S2.3 are re-run using the updated U-Net network to obtain the updated... , Z R and Substitute into the loss function L G In the middle, the loss value is calculated.
[0024] In step S2.2, the modulation process is obtained by combining Equation 1) and Equation 2); Equations 1) and 2) are as follows: Formula 1); Equation 2); In equation 1), This represents a convolutional neural network with learnable parameters; In equation 2), and Each of the above The mean and standard deviation; The multiplication symbol is used to represent multiplication.
[0025] In step S2.3, cross-attention is obtained using equation 3). Z R ; where Equation 3) is as follows: Equation 3); In equation 3), ; Indicates the query conversion parameters; ; Indicates index conversion parameters; ; Indicates content conversion parameters; T Represents a matrix Perform a transpose operation; d This represents the scaling factor, specifically 512.
[0026] In step S2.3, the noise characteristics Z The modulation features are ordered in two dimensions. The modulation features are obtained by arranging them into a one-dimensional sequence along the column direction; the updated modulation features are obtained by transformation. It is described by a one-dimensional sequence. Z R It is obtained by transforming it back into a two-dimensional sequence.
[0027] In step S2.4, the loss function L G Equation 4) is used to represent: Equation 4); In equation 4), This indicates a VAE encoder with frozen parameters; This represents a multi-scale conditional feature encoder; This represents the training parameters in a multi-scale conditional feature encoder. This represents a cloud-aware cross-modal semantic extractor; R This represents the training parameters in the cloud-aware cross-modal semantic extractor; This represents a parameterized denoising U-Net; This represents noise that follows a standard normal distribution and is randomly sampled. express Extracted modulation features ; express Extracted cross attention Z R ; This represents the average value of the prediction error of the cloud removal model; This represents a standard normal distribution.
[0028] In step S1, the method for acquiring the thin cloud-cloud-free image pair data includes: Step S1.1: Perform median filtering on the raw optical images of the target area (specifically, by purchasing and downloading the raw optical images of the target area from PlanetScope satellite imagery products) to obtain preprocessed data; Step S1.2: Perform registration processing on the preprocessed data to obtain thin cloud-cloud-free image pairs.
[0029] In step S1.2, the registration process includes: First, local SIFT feature points are extracted in the non-cloud-covered areas of the image pairs to be registered in the preprocessed data for registration between thin cloud and cloudless images. By calculating the Euclidean distance between the SIFT feature points of the cloudless image and the thin cloud image, the matching point with the smallest distance to the corresponding position in the thin cloud image is found for each feature point in the cloudless image, thus achieving the matching of the same feature points; wherein, the total number of extracted local SIFT feature points is not less than 50. Secondly, by using quadratic polynomial correction, the pixel coordinates of SIFT feature points on the thin cloud image are mapped to the coordinates of the corresponding SIFT feature points on the cloudless image, thus achieving registration between the thin cloud and cloudless images. The quadratic polynomial is represented by equations 5) and 6): Equation 5); Formula 6); In equation 5), Represents thin cloud image data; and All are transformation parameters of Equation 5); Represents the pixel coordinates of SIFT feature points on thin cloud images ; In equation 6), y This indicates cloudless imagery data; and All are transformation parameters of Equation 6); Finally, the least squares method is used to solve for the transformation parameters. and The root mean square error (RMSE) is used to verify whether the registration accuracy between thin cloud and cloudless images is controlled within 0.5 pixels. If the RMSE is greater than 0.5 pixels, the residuals of each SIFT feature point are calculated, the point with the largest residual is removed, and then the least squares method is used again to solve for the transformation parameters. and This continues until the root mean square error is within 0.5 pixels.
[0030] Solving for transformation parameters using the least squares method a 0~ a 5 and b 0~ b The process of 5 includes: ① n (In this embodiment) n Six sets of corresponding pixel coordinate equations can be selected and converted into matrix form for thin cloud image data. coordinate transformation The equation is reconstructed as follows: X=D x A ; in,X Represents the observation vector of thin cloud image; D x Represents the thin cloud image coefficient matrix; A Represents the vector of thin cloud image coefficients; Thin cloud image observation vector X For one n A vector of row 1 and column 1, containing the coordinates of thin cloud image data for SIFT feature points with the same name for cloudless images. : ; Thin cloud image coefficient matrix D x For one n A 6-row matrix containing the pixel coordinates of SIFT feature points on thin cloud images. u , v ); ; Vector of thin cloud image coefficients to be determined It is a 6x1 vector containing the parameters of the quadratic polynomial transformation to be determined: ; Similarly, for cloudless image data The transformation equation for the coordinate system is: Y=D y B ; in, Y Represents the observation vector of cloudless imagery; D y This represents the coefficient matrix of cloudless images; B Represents the vector of cloudless image coefficients; Cloudless image observation vector X For one n A vector with row 1 and column 1, containing the coordinates of cloudless image data for SIFT feature points of the same name. : ; Cloudless Image Coefficient Matrix D y For one n A matrix of 6 rows and 6 columns containing the pixel coordinates of SIFT feature points on a cloudless image. u , v ); ; Vector coefficients of cloudless image to be determined B It is a 6x1 vector containing the parameters of the quadratic polynomial transformation to be determined: ; ② Solve for the parameters of the quadratic polynomial transformation; Solve the equation as follows: ; ; ③Accuracy evaluation and feedback: After obtaining the parameters, the root mean square error of the coordinates for a single pair of images needs to be obtained: ; and After solving for the parameters of the quadratic polynomial transformation, the coordinates of the corresponding points in the cloudless image are obtained by transforming the coordinates of the corresponding points in the thin cloud image through a quadratic mathematical model. n Indicates the number of groups of pixels with the same name; j Take 1, 2, ... n Any value in; x j The x-coordinate of the corresponding point in the cloudless image for the feature point in the thin cloud image; y j This represents the ordinate of the corresponding point in the cloudless image.
[0031] like If the value is greater than 0.5 pixels, then calculate the residual of each SIFT feature point. The points with the largest residuals are removed, and then the least squares method is used again to solve for the transformation parameters. a 0~ a 5 and b 0~ b 5. Continue until the root mean square error is within 0.5 pixels.
[0032] residual The specific calculation process is as follows: ; ; .
[0033] Thin cloud removal tests were conducted on the target area using the Pix2Pix method, SpA GAN method, UnCRtainTS method, SwinIR method, DehazeFormer method, DiffCR method, and the cloud removal method from the example (i.e., StableCR in Table 1). The test results are shown in Table 1. The test steps for obtaining the test results in Table 1 are as follows: H1) Download the publicly available code corresponding to the methods described in the above documents to the same training computer used in this invention, and configure the same experimental conditions. H2) The above method is trained using the thin cloud-cloudless image pair data in the training set described in step S1. The paired thin cloud images are used as the input conditions of the method, and the corresponding cloudless images are used as the training labels of the above method. The mapping ability from input conditions to training labels is trained. The training process is completed on the same NVIDIA Tesla A100 (80G) GPU graphics card. H3) After training the above methods for 20,000 iterations, stop training. Use the thin cloud-cloudless images in the test set described in step S1 to test each method. Set the input condition to the paired thin cloud images and require the output to be the corresponding cloudless images. H4) The cloudless images obtained by the above method are compared and evaluated with the cloudless images in the test set. The evaluation criteria use the following parameters: PSNR measures the overall reconstruction quality using the signal-to-noise ratio; the higher the PSNR, the higher the quality of the reconstructed image. MAE reflects the average pixel deviation, while RMSE highlights significant errors through a squared penalty mechanism; the smaller the MAE and RMSE, the higher the quality of the reconstructed image. SSIM is used to evaluate the overall retention of brightness, contrast and structure. The SSIM value ranges from [-1, 1]. When two images are identical, the SSIM value is 1. SAM is used to analyze the fidelity of high-dimensional spectral features. The lower the SAM, the higher the quality of the reconstructed image. LPIPS captures visual similarity to the human eye based on deep network features. The lower the LPIPS, the higher the similarity between the reconstructed image and the reference image.
[0034] Table 1. Test results of thin cloud removal in the target area using different methods.
[0035] The Pix2Pix method, SpA GAN method, UnCRtainTS method, SwinIR method, DehazeFormer method, DiffCR method, and the cloud removal method of this invention (i.e., StableCR in Table 2) were used to test thin cloud removal in a cross-regional area (specifically, a selected target city within the cross-regional area). The test results are shown in Table 2. The only difference between the test steps for obtaining the test results in Table 2 and those in Table 1 is that in step H3), the thin cloud-to-cloudless imagery from the cross-regional area was used to test each method.
[0036] Table 2. Test results of different methods for thin cloud removal across regions.
[0037] The methods in Tables 1 and 2 are cited from the following references: The Pix2Pix method is cited from the following literature: Isola P, Zhu JY, Zhou T, et al. Image-to-Image Translation with Conditional Adversarial Networks[C] / / IEEE Conference on Computer Vision and Pattern Recognition. 2017: 5967-5976; The SpA GAN method is cited from the following literature: Pan H. Cloud Removal for Remote Sensing Imagery via Spatial Attention Generative Adversarial Network[J]. arXiv preprintarXiv:2009.13015, 2020; The UnCRtainTS method is quoted from the literature: Ebel P, Fare Garnot VS, Schmitt M, et al.UnCRtainTS: Uncertainty Quantification for Cloud Removal in Optical SatelliteTime Series[C] / / IEEE Conference on Computer Vision and Pattern RecognitionWorkshops. 2023: 2086-2096; The SwinIR method is cited in the following literature: Liang J, Cao J, Sun G, et al. SwinIR: Image Restoration Using Swin Transformer[C] / / IEEE International Conference on Computer Vision Workshops. 2021: 1833-1844; The DehazeFormer method is cited in the following literature: Song Y, He Z, Qian H, et al. VisionTransformers for Single Image Dehazing[J]. IEEE Transactions on ImageProcessing, 2023, 32: 1927-1941; The DiffCR method is quoted from the literature: Zou X, Li K, Xing J, et al. DiffCR: A FastConditional Diffusion Framework for Cloud Removal From Optical SatelliteImages[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 1-14.
[0038] As shown in Table 1, compared with the methods mentioned in existing literature, StableCR has achieved state-of-the-art performance on all visual evaluation metrics. It not only significantly reduces pixel-level errors in reconstructed images, but also generates spatial structures and spectral textures that conform to visual cognition.
[0039] As shown in Table 2, compared with the methods mentioned in existing literature, StableCR has achieved state-of-the-art performance on all visual evaluation metrics and can maintain stable performance even when there are differences in data distribution. This fully demonstrates that the cloud removal method proposed in the embodiment has excellent generalization ability.
[0040] The study investigated the thin cloud removal tests on the target area using the Concat comparison method, without RAM, without MS, and the cloud removal method of the example (i.e., StableCR in Table 3). The test results are shown in Table 3. The test steps for obtaining the test results in Table 3 are as follows: Concat comparison method: This method omits the multi-scale degradation spatial feature extraction module and the cloud pollution semantic feature guidance module from the model in this embodiment of the invention. It uses the multi-scale conditional features extracted by the multi-scale degradation spatial feature extraction module. Cross-modal semantic features extracted by the cloud pollution semantic feature guidance module C R The thin cloud imagery is merged with the cloudless imagery in the channel dimension and used as input to the model, without modulating the two features into the U-Net. The cloud removal method is then trained using the thin cloud-free imagery pair concat comparison data from the training set described in step S1. Finally, the trained model is tested on the thin cloud-free imagery pair data from the test set described in step S1.
[0041] w / o RAM: The cloud pollution semantic feature guidance module in the model is omitted. The thin-cloud-free imagery from the training set described in step S1 is used to train the w / o RAM to remove clouds. Finally, the trained model is tested on the thin-cloud-free imagery from the test set described in step S1.
[0042] w / o MS: The multi-scale degradation spatial feature extraction module in the model is omitted. The w / o MS model is trained using the thin-cloud-free imagery from the training set described in step S1. Finally, the trained model is tested on the thin-cloud-free imagery from the test set described in step S1.
[0043] Table 3. Test results of thin cloud removal in the target area using the Concat comparison method, without RAM, without MS, and the cloud removal method in the examples.
[0044] Table 3 shows that Concat fails to fully utilize the multi-scale information of thin cloud imagery, exhibiting significant limitations. Compared to Concat, both w / o RAM and w / o MS show gains across all visual evaluation metrics, indicating the positive impact of the multi-scale degradation spatial feature extraction module and the cloud pollution semantic feature guidance module on the thin cloud removal process. StableCR achieves the best performance across all visual evaluation metrics, demonstrating the superiority of coupling multi-scale visual features with cross-modal semantic features.
[0045] The above description is only a preferred embodiment of the present invention and does not limit the scope of the present invention. All equivalent structural transformations made under the inventive concept of the present invention using the contents of the present invention specification and drawings, or direct / indirect applications in other related technical fields, are included within the protection scope of the present invention.
Claims
1. A generative thin cloud remote sensing image declouding method based on degradation perception and semantic guidance, characterized in that, include: Step S1: Obtain thin cloud-cloud-free image pairs of the target area and divide them into training set and test set according to the ratio; Step S2: Use the thin cloud-cloud-free images in the training set to train the data to obtain a generative thin cloud remote sensing image cloud removal model based on degradation perception and semantic guidance. Step S3: Input the thin cloud-cloud-free image pairs from the test set into the trained cloud removal model for prediction to obtain the latent space features of the test set. Then, post-process the latent space features of the test set to obtain the remote sensing image after removing the thin clouds.
2. The cloud removal method for generative thin cloud remote sensing images based on degradation perception and semantic guidance as described in claim 1, characterized in that, In step S2, the steps for training the cloud removal model include: Step S2.1: Select the cloudless image data from the thin cloud-cloudless image pair data in the training set. y After being processed by the VAE encoder in the multi-scale degraded spatial feature extraction module, cloudless latent spatial features are obtained. Z y ; will the Z y After being processed in the U-Net network, each time step is obtained. t Conditional latent space features of noisy images and by each Determine the intermediate features of each layer of the corresponding U-Net network ;in, c Indicates cloudless image conditions; This represents the number of layers in the U-Net network at a certain spatial scale. This indicates the generation process of the U-Net network during diffusion denoising; The thin cloud image data in the thin cloud-clear image pair data of the training set. x After being processed by the VAE encoder in the multi-scale degradation spatial feature extraction module, the thin cloud hidden spatial features are obtained. Z x ; will the Z x With time step t After being processed by the multi-scale conditional feature encoder of the multi-scale degradation space feature extraction module, the multi-scale conditional features are obtained. ;in, N This represents the total number of spatial scale levels in the U-Net network; cond Indicates thin cloud image conditions; Represents the number of levels at a certain spatial scale in the U-Net network. i Corresponding scale-conditional features; Step S2.2: Extract the scale-conditional features from each of the aforementioned features using convolutional layers. Predict a set of modulation parameter scaling factors With bias term , will the The above With the After being input into the U-Net network and modulated, modulation features are obtained. ; Step S2.3: Input the thin cloud image data from the thin cloud-cloud-free image pairs in the training set into the RAM encoder of the cloud pollution semantic feature guidance module for processing to obtain cross-modal semantic features. C R ; will the C R With noise characteristics Z Cross-attention is obtained through combined computation Z R ; wherein, the noise characteristics Z By the modulation features Transformed to obtain; the said Z R The transformation yields updated modulation features. ; Step S2.4, the... The above Z R and stated Substitute into the loss function L G In the process, the loss value is calculated. If the loss value is greater than the target threshold, the loss value is backpropagated to update the parameters of the entire U-Net network model. If the loss value is less than or equal to the target threshold, the trained cloud removal model is obtained.
3. The cloud removal method for generative thin cloud remote sensing images based on degradation perception and semantic guidance as described in claim 2, characterized in that, In step S2.4, the backpropagation method includes: The chain rule is used to calculate the gradient of the loss function with respect to the parameters of each layer of the U-Net network. The weights of the U-Net network are then updated using these gradients. Subsequently, steps S2.1 to S2.3 are re-run using the updated U-Net network to obtain the updated... , Z R and Substitute into the loss function L G In the middle, the loss value is calculated.
4. The cloud removal method for generative thin cloud remote sensing images based on degradation perception and semantic guidance as described in claim 2, characterized in that, In step S2.2, the modulation process is obtained by combining Equation 1) and Equation 2); Equations 1) and 2) are as follows: Formula 1); Equation 2); In equation 1), This represents a convolutional neural network with learnable parameters; In equation 2), and Each of the above The mean and standard deviation; The multiplication symbol is used to represent multiplication.
5. The cloud removal method for generative thin cloud remote sensing images based on degradation perception and semantic guidance as described in claim 2, characterized in that, In step S2.3, cross-attention is obtained using equation 3). Z R ; where Equation 3) is as follows: Equation 3); In equation 3), ; Indicates the query conversion parameters; ; Indicates index conversion parameters; ; Indicates content conversion parameters; T Represents a matrix Perform a transpose operation; d This represents the scaling factor.
6. The cloud removal method for generative thin cloud remote sensing images based on degradation perception and semantic guidance as described in claim 2, characterized in that, In step S2.3, the noise characteristics Z The modulation features are ordered in two dimensions. The modulation features are obtained by arranging them into a one-dimensional sequence along the column direction; the updated modulation features are obtained by transformation. It is described by a one-dimensional sequence. Z R It is obtained by transforming it back into a two-dimensional sequence.
7. The cloud removal method for generative thin cloud remote sensing images based on degradation perception and semantic guidance as described in claim 2, characterized in that, In step S2.4, the loss function L G Equation 4) is used to represent: Equation 4); In equation 4), This indicates a VAE encoder with frozen parameters; This represents a multi-scale conditional feature encoder; This represents the training parameters in a multi-scale conditional feature encoder. This represents a cloud-aware cross-modal semantic extractor; R This represents the training parameters in the cloud-aware cross-modal semantic extractor; This represents a parameterized denoising U-Net; This represents noise that follows a standard normal distribution and is randomly sampled. express Extracted modulation features ; express Extracted cross attention Z R ; This represents the average value of the prediction error of the cloud removal model; This represents a standard normal distribution.
8. The cloud removal method for generative thin cloud remote sensing images based on degradation perception and semantic guidance as described in any one of claims 2 to 7, characterized in that, In step S1, the method for acquiring the thin cloud-cloud-free image pair data includes: Step S1.1: Perform median filtering on the raw optical images of the target area to obtain preprocessed data; Step S1.2: Perform registration processing on the preprocessed data to obtain thin cloud-cloud-free image pairs.
9. The generative thin cloud remote sensing image declouding method based on degradation perception and semantic guidance as described in claim 8, characterized in that, In step S1.2, the registration process includes: First, local SIFT feature points are extracted in the non-cloud-covered areas of the image pairs to be registered in the preprocessed data for registration between thin cloud and cloudless images. By calculating the Euclidean distance between the SIFT feature points of the cloudless image and the thin cloud image, the matching point with the smallest distance to the corresponding position in the thin cloud image is found for each feature point in the cloudless image, thus achieving the matching of the same feature points; wherein, the total number of extracted local SIFT feature points is not less than 50. Secondly, by using quadratic polynomial correction, the pixel coordinates of SIFT feature points on the thin cloud image are mapped to the coordinates of the corresponding SIFT feature points on the cloudless image, thus achieving registration between the thin cloud and cloudless images. The quadratic polynomial is represented by equations 5) and 6): Equation 5); Formula 6); In equation 5), Represents thin cloud image data; and All are transformation parameters of Equation 5); Represents the pixel coordinates of SIFT feature points on thin cloud images ; In equation 6), y This indicates cloudless imagery data; and All are transformation parameters of Equation 6); Finally, the least squares method is used to solve for the transformation parameters. and The root mean square error (RMSE) is used to verify whether the registration accuracy between thin cloud and cloudless images is controlled within 0.5 pixels. If the RMSE is greater than 0.5 pixels, the residuals of each SIFT feature point are calculated, the point with the largest residual is removed, and then the least squares method is used again to solve for the transformation parameters. and This continues until the root mean square error is within 0.5 pixels.
10. The generative thin cloud remote sensing image declouding method based on degradation perception and semantic guidance as described in claim 7, characterized in that, In step S1, the thin cloud-cloudless image data is divided into a training set and a test set in a ratio of 6:4 to 9:
1. In step S2.4, the target threshold value is 2; In step S3, the post-processing involves using a VAE decoder to process the latent space features in the test set.