Thin cloud removal method fusing physical prior and feedback enhancement
By constructing a thin cloud removal network that integrates physical priors and feedback enhancement, and combining it with an atmospheric scattering model and a generative adversarial network, the problem of thin cloud removal in high-concentration clouds and complex scenes is solved, achieving efficient detail preservation and improved generalization capabilities.
Patent Information
- Application Number
- CN202510576953.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-05
AI Technical Summary
Existing thin cloud removal methods in high-density cloud layers or complex scenes suffer from detail loss or artifacts, insufficient generalization capabilities, and high computational costs for iterative training, which affects the declouding effect and application efficiency.
A thin cloud removal network that integrates physical priors and feedback enhancement is constructed. It combines a feature fusion residual declouding generator, an atmospheric scattering restoration module, and a spatial detail enhancement discriminator. Physical prior knowledge is provided through the atmospheric scattering model, and adversarial loss and dual reconstruction loss mechanisms are introduced to optimize the declouding effect.
It improves the declouding effect in high-density clouds and complex scenes, enhances the ability to capture ground texture and details, reduces the complexity of declouding, and improves the robustness and generalization ability of the model.
Smart Images

Figure CN120598804A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing and relates to a thin cloud removal method combining an atmospheric scattering model with a generative adversarial network. Background Art
[0002] Most current cloud removal methods for remote sensing images based on atmospheric scattering models rely on specific assumptions, lack in-depth research on the coupling relationship between cloud and ground object information, and have insufficient generalization capabilities, making them difficult to adapt to large-scale datasets. While generative adversarial networks improve cloud removal, they fail to fully consider the dynamic changes in cloud density and distribution, resulting in suboptimal cloud removal performance. High cloud density in remote sensing images obscures ground object information, making it prone to loss of detail or artifacts during removal. In complex scenes, thin clouds often resemble high-brightness features, making it difficult for the network to fully preserve ground objects and edge textures during cloud removal, impacting downstream tasks.
[0003] The paper "Xu Z, Wu K, Wang W, et al. Semi-supervised thin cloud removal with mutually beneficial guides[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2022, 192:327-343" proposes a semi-supervised thin cloud removal framework based on mutually beneficial guides. By collaboratively training a model mutual donor network and a target mutual donor network, this approach improves the utilization of cloudy images and the reliability of declouding results. This approach introduces beneficial pattern guidance and target guidance to expand training data and optimize declouding results, respectively, and employs a cross-scale content fusion structure to generate high-resolution images. However, this approach suffers from the following challenges: First, in high-density cloud layers or complex scenes, insufficient cloudy image quality can lead to loss of detail or artifacts, compromising declouding effectiveness. Second, iterative training of multispectral images is computationally expensive, and the collaborative optimization of the network requires fine-tuning of parameters, increasing implementation complexity. This limitation limits the approach's application efficiency in real-world scenarios and its scalability to large-scale datasets. Summary of the Invention
[0004] In order to overcome the shortcomings of the existing technology, improve the cloud removal effect in high-concentration clouds or complex scenes, and reduce the complexity of cloud removal, the present invention provides a thin cloud removal method that integrates physical priors and feedback enhancement.
[0005] A thin cloud removal method integrating physical prior and feedback enhancement includes the following steps:
[0006] Step 1: Construct a thin cloud removal network that integrates physical priors and feedback enhancement;
[0007] The thin cloud removal network that integrates physical prior and feedback enhancement is a conditional adversarial generation network cGAN;
[0008] The thin cloud removal network that integrates physical priors and feedback enhancement includes a feature fusion residual declouding generator, an atmospheric scattering restoration module, and a spatial detail enhancement discriminator. The feature fusion residual declouding generator outputs a declouded image. The atmospheric scattering restoration module extracts atmospheric light values and transmittance maps to provide physical prior knowledge. The extracted atmospheric light values, transmittance maps, and declouded images are reconstructed to obtain a pseudo thin cloud image. The true cloud-free image, declouded image, thin cloud image, and pseudo thin cloud image are then input into the spatial detail enhancement discriminator for discrimination. The spatial detail enhancement discriminator outputs a true or false discrimination result. The discrimination result updates the training parameters of the feature fusion residual declouding generator through a loss function.
[0009] Step 2: Use the dataset, training set, and validation set to train the thin cloud removal network that integrates physical priors and feedback enhancement;
[0010] Step 3: Use the trained thin cloud removal network that integrates physical priors and feedback enhancement to perform remote sensing thin cloud removal and identification.
[0011] Furthermore, the feature fusion residual declouding generator includes an encoding module, an intermediate block and a decoding module;
[0012] The input of the feature fusion residual declouding generator is a thin cloud image; the output of the feature fusion residual declouding generator is a declouded image; the data flow between the encoding module, the intermediate block and the decoding module adopts a dense connection; the dense connection is a DenseNet-like connection;
[0013] The encoding module includes a first encoding block, a second encoding block, a third encoding block and a fourth encoding block; the input of the encoding module is a thin cloud image; the output of the encoding module is a high-dimensional feature map; the first encoding block includes a first convolutional layer and a first deep network layer; the step size of the first convolutional layer is 1; the first deep network layer includes three identical residual blocks; the number of channels of the three identical residual blocks is 16; the second encoding block includes a second convolutional layer, a first feature fusion residual declouding module and a second deep network layer; the second convolutional layer, the first feature fusion residual declouding module and the second deep network layer are connected in sequence; the second convolutional layer is a 3×3 convolution kernel, and the second convolutional layer is a 3×3 convolution kernel. The stride of the convolution layer is 2; the first feature fusion residual declouding module includes DenseNet and a feature fusion block; the feature fusion block includes a fifth encoding block and a fifth decoding block; the fifth encoding block includes a 3×3 convolution kernel and a PReLU activation function; the fifth decoding block includes a 4×4 convolution kernel and a PReLU activation function; the DenseNet includes four identical network layers; each network layer includes a fourth convolution layer and a ReLU activation function; the fourth convolution layer is a 3×3 convolution kernel; the structure of the third encoding block and the structure of the fourth encoding block are the same as the structure of the second encoding block; the structure of the second deep network layer is the same as the structure of the first deep network layer;
[0014] The intermediate block includes a third convolutional layer, a second feature fusion residual declouding module, and a first declouding module; the third convolutional layer, the second feature fusion residual declouding module, and the first declouding module are connected in sequence; the first declouding module includes eighteen identical residual blocks; the structure of the third convolutional layer is the same as that of the second convolutional layer; the input of the intermediate block is a high-dimensional feature map; the output of the intermediate block is a reconstructed high-dimensional feature map;
[0015] The structure of the second feature fusion residual declouding module is the same as that of the first feature fusion residual declouding module;
[0016] The decoding module includes a first decoding block, a second decoding block, a third decoding block and a fourth decoding block; the input of the decoding module is a reconstructed high-dimensional feature map; the output of the decoding module is a declouded image; the first decoding block includes a first deconvolution layer, a second declouding module and a third feature fusion residual declouding module; the first deconvolution layer, the second declouding module and the third feature fusion residual declouding module are connected in sequence; the kernel size of the first deconvolution layer is 3×3, and the stride of the first deconvolution layer is 2; the structure of the second decoding block and the third decoding block are the same as that of the first decoding block; the structure of the third feature fusion residual declouding module is the same as that of the first feature fusion residual declouding module;
[0017] The second declouding module includes upsampling, three identical residual blocks and an activation function layer; the structures of the second decoding block and the third decoding block are the same as those of the first decoding block; the input of the second declouding module is the feature map output by the first deconvolution layer; the output of the second declouding module is the high-dimensional feature map after declouding; the fourth decoding block is based on the first decoding block and adds a fifth convolution layer after the third feature fusion residual declouding module; the structure of the fifth convolution layer is the same as that of the first convolution layer; the output of the fourth decoding block is the transmission feature and the declouded image;
[0018] The atmospheric scattering restoration module includes an upsampling-convolution module and a U-Net network; the input of the atmospheric scattering restoration module is a thin cloud image and a transmission feature; the output of the atmospheric scattering restoration module is an atmospheric light value and a transmittance map;
[0019] The upsampling-convolution module includes bilinear interpolation, a sixth convolution layer, and a seventh convolution layer; the bilinear interpolation, the sixth convolution layer, and the seventh convolution layer are connected in sequence; the input of the upsampling-convolution module is the transmittance feature; the output of the upsampling-convolution module is a transmittance map; the transmittance feature is amplified by bilinear interpolation and then input into the sixth convolution layer and the seventh convolution layer to obtain the transmittance map; the U-Net network includes downsampling and upsampling; the input of the U-Net network is a thin cloud image; the output of the U-Net is the atmospheric light value;
[0020] The spatial detail enhancement discriminator is a Transformer discriminator; the input of the spatial detail enhancement discriminator is a thin cloud image, a cloud removal image, a reconstructed thin cloud image and a real cloud-free image;
[0021] The reconstructed thin cloud image I_fake is:
[0022] I_fake=T_tran×J_fake+A_atm(1-T_tran)
[0023] T_tran is the transmittance map, J_fake is the cloud removal image, and A_atm is the atmospheric light value;
[0024] The output of the spatial detail enhancement discriminator is the true and false discrimination results D(J), D(I), D(G(I)) and D(G(J)); among them, D(J) is the prediction of the spatial detail enhancement discriminator for the real cloud-free image J, that is, the probability that the spatial detail enhancement discriminator believes that J is a real image; D(I) is the prediction of the spatial detail enhancement discriminator for the thin cloud image I, that is, the probability that the spatial detail enhancement discriminator believes that I is a real image; D(G(I)) is the prediction of the spatial detail enhancement discriminator for the generated de-clouding image, that is, the probability that the spatial detail enhancement discriminator believes that the de-clouding image is a real image; D(G(J)) is the prediction of the discriminator for the reconstructed pseudo thin cloud image; G(I) is the de-clouding image generated by the feature fusion residual de-clouding generator from the thin cloud image I.
[0025] Furthermore, the total loss function L of the thin cloud removal network that integrates physical prior and feedback enhancement is total for:
[0026] L total =λ1L adv +λ2L L1 +λ3L mse
[0027] Among them L adv is the adversarial loss, λ1 is the weight parameter of the adversarial loss; L L1 is the double reconstruction loss, λ2 is the weight parameter of the double reconstruction loss; L mse is the detail enhancement loss, λ3 is the weight parameter of the detail enhancement loss;
[0028] Adversarial loss L adv for:
[0029] L adv =L I2J (G,D,I,J)+L J2I (G,D,I,J)
[0030] Among them, L I2J (G, D, I, J) is the adversarial loss between the de-clouded image and the true cloud-free image, L J2I (G, D, I, J) is the adversarial loss between the thin cloud image and the reconstructed thin cloud image;
[0031] The adversarial loss L between the de-clouded image and the true cloud-free image J2I (G,D,I,J) is:
[0032]
[0033] is the true cloud-free image discrimination expectation; It is to de-cloud the image to deceive expectations;
[0034] The true cloud-free image discrimination expectation is calculated on the true cloud-free image J, p data (J) is the data distribution of the real cloud-free image; when D(J) is close to 0, it means the probability that J is a real image is lower; when D(J) is close to 1, it means the probability that J is a real image is higher; where p data (I) is the data distribution of thin cloud images;
[0035] The adversarial loss L between the thin cloud image and the reconstructed thin cloud image J2I (G,D,I,J) is:
[0036]
[0037] is the discriminant expectation of thin cloud image; The deception expectation of the reconstructed thin cloud image; the discriminant expectation of the thin cloud image is calculated on the thin cloud image I; when D(I) is close to 0, the probability that I is a real image is lower; when D(I) is close to 1, the probability that I is a real image is higher; the deception expectation of the reconstructed thin cloud image is calculated on the real cloud-free image J; G(J) is the reconstructed thin cloud image generated by the feature fusion residual declouding generator from the real cloud-free image J; D(G(J)) is the prediction of the spatial detail enhancement discriminator on the reconstructed image, that is, the probability that the spatial detail enhancement discriminator believes that the reconstructed thin cloud image is a real image;
[0038] Dual reconstruction loss L L1 for:
[0039]
[0040] is the forward reconstruction loss; is the reverse reconstruction loss;
[0041] Forward reconstruction loss for:
[0042]
[0043] λ4 is the weight parameter of the forward reconstruction loss, which is used to balance the contribution of the forward reconstruction loss in the reconstruction loss; ‖·‖1 represents the L1 norm; the L1 norm calculates the sum of the absolute values of the vector elements; ||G(I)-J||1 measures the difference between the declouded image and the true cloud-free image;
[0044] Reverse reconstruction loss for:
[0045]
[0046] λ5 is the weight parameter of the reverse reconstruction loss; A(G(I)) is the reconstructed thin cloud image; ||A(G(I))-I||1 measures the difference between the reconstructed thin cloud image and the thin cloud image;
[0047] Detail enhancement loss L mse :
[0048]
[0049] Where N represents the total number of training set samples, J is the real cloud-free image, and G(I) is the cloud-free image;
[0050] When the total loss function reaches convergence, that is, the training error fluctuation range is less than the preset threshold, and the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) of the declouded image both reach the preset performance targets, the training of the thin cloud removal network that integrates physical priors and feedback enhancement is considered completed, and the thin cloud removal network that integrates physical priors and feedback enhancement at this time is saved as the optimal model.
[0051] Furthermore, the preset threshold is 3%.
[0052] Furthermore, the thin cloud images and the real cloud-free images are selected from remote sensing cloud removal datasets; the remote sensing cloud removal datasets include the RICE1 dataset, the WHUS2-CR dataset and the Landsat8 dataset.
[0053] Furthermore, the values of the weight parameter λ1 of the adversarial loss, the weight parameter λ2 of the dual reconstruction loss, and the weight parameter λ3 of the detail enhancement loss are:
[0054] λ1=λ2=1;λ3=20。
[0055] The beneficial effects of the present invention are:
[0056] The present invention proposes a method for removing thin clouds from remote sensing images that integrates physical priors and feedback enhancement. The method of the present invention addresses the problems of severe detail loss and insufficient generalization ability of existing methods in high-density clouds and complex scenes. By combining the physical prior knowledge of the atmospheric scattering model with the feedback enhancement mechanism of the generative adversarial network, the model performance is effectively improved. In addition, a spatial detail enhancement discriminator and a cross-scale feature fusion module are introduced to enable the network to adjust the dynamic weights of high-density thin clouds, enhance the ability to capture the texture and details of ground objects, and make it more robust in complex environments. The cloud removal experiments of the present invention on the public datasets RICE1, WHUS2-CR and Landsat8 show that the four image quality evaluation indicators PSNR, SSIM, MSE and NIQE are improved by 4.85, 1.22, 16.69 and 28.98 percentage points respectively compared with the literature methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 A framework for thin cloud removal networks that fuses physical priors with feedback enhancement;
[0058] Figure 2 This is the structural diagram of the feature fusion residual declouding module;
[0059] Figure 3 This is the structural diagram of the second cloud removal module;
[0060] Figure 4 This is the structure diagram of the spatial enhancement detail discriminator; DETAILED DESCRIPTION
[0061] A thin cloud removal method integrating physical prior and feedback enhancement includes the following steps:
[0062] Step 1: Construct a thin cloud removal network that integrates physical priors and feedback enhancement;
[0063] The thin cloud removal network that integrates physical prior and feedback enhancement is a conditional adversarial generation network cGAN;
[0064] The thin cloud removal network that integrates physical priors and feedback enhancement includes a feature fusion residual declouding generator, an atmospheric scattering restoration module, and a spatial detail enhancement discriminator. The atmospheric scattering restoration module extracts atmospheric light values and transmittance maps to provide physical prior knowledge. The extracted atmospheric light values and transmittance maps are reconstructed with the declouded image output by the feature fusion residual declouding generator to obtain a pseudo thin cloud image. The real cloud-free image, the declouded image, the thin cloud image, and the pseudo thin cloud image are then input into the spatial detail enhancement discriminator for discrimination. The spatial detail enhancement discriminator outputs a true or false discrimination result. The discrimination result guides the feature fusion residual declouding generator to update training parameters through a loss function.
[0065] The feature fusion residual de-clouding generator includes an encoding module, an intermediate block and a decoding module;
[0066] The input of the feature fusion residual declouding generator is a thin cloud image; the output of the feature fusion residual declouding generator is a declouded image;
[0067] The encoding module includes a first encoding block, a second encoding block, a third encoding block and a fourth encoding block; the input of the encoding module is a thin cloud image; the output of the encoding module is a high-dimensional feature map;
[0068] The first coding block includes a first convolutional layer and a first deep network layer; the step size of the first convolutional layer is 1; the deep network layer includes three identical residual blocks; the number of channels of the three identical residual blocks is 16;
[0069] The second encoding block includes a second convolutional layer, a first feature fusion residual declouding module, and a second deep network layer; the second convolutional layer, the first feature fusion residual declouding module, and the second deep network layer are connected in sequence; the second convolutional layer is a 3×3 convolution kernel, and the stride of the second convolutional layer is 2; the first feature fusion residual declouding module includes a DenseNet and a feature fusion block;
[0070] The structure of the third coding block and the structure of the fourth coding block are the same as the structure of the second coding block;
[0071] The intermediate block includes a third convolutional layer, a second feature fusion residual declouding module, and a first declouding module; the third convolutional layer, the second feature fusion residual declouding module, and the first declouding module are connected in sequence; the first declouding module includes eighteen identical residual blocks; the structure of the third convolutional layer is the same as that of the second convolutional layer; the input of the intermediate block is a high-dimensional feature map; the output of the intermediate block is a reconstructed high-dimensional feature map;
[0072] The structure of the second feature fusion residual declouding module is the same as that of the first feature fusion residual declouding module;
[0073] The structure of the third feature fusion residual declouding module is the same as that of the first feature fusion residual declouding module;
[0074] The feature fusion block includes a fifth encoding block and a fifth decoding block; the fifth encoding block includes a 3×3 convolution kernel and a PReLU activation function; the fifth decoding block includes a 4×4 convolution kernel and a PReLU activation function;
[0075] DenseNet consists of four identical network layers; each network layer includes a fourth convolutional layer and a ReLU activation function; the fourth convolutional layer is a 3×3 convolution kernel;
[0076] The decoding module includes a first decoding block, a second decoding block, a third decoding block and a fourth decoding block; the input of the decoding module is a reconstructed high-dimensional feature map; the output of the decoding module is a de-clouded image;
[0077] The first decoding block includes a first deconvolution layer, a second declouding module, and a third feature fusion residual declouding module; the first deconvolution layer, the second declouding module, and the third feature fusion residual declouding module are connected in sequence; the kernel size of the first deconvolution layer is 3×3, and the stride of the deconvolution layer is 2;
[0078] The structure of the second decoding block and the third decoding block are the same as the structure of the first decoding block;
[0079] The second declouding module includes upsampling, three identical residual blocks, and an activation function layer; the structures of the second decoding block and the third decoding block are the same as those of the first decoding block; the input of the second declouding module is the feature map output by the deconvolution layer; the output of the second declouding module is the high-dimensional feature map after declouding;
[0080] The fourth decoding block is based on the first decoding block, and adds a fifth convolutional layer after the third feature fusion residual declouding module; the structure of the fifth convolutional layer is the same as that of the first convolutional layer; the output of the fourth decoding block is the transmission feature and the declouded image;
[0081] The data flow between the encoding module, the intermediate block and the decoding module adopts a dense connection; the dense connection is a DenseNet-like connection;
[0082] The atmospheric scattering restoration module includes an upsampling-convolution module and a U-Net network;
[0083] The input of the atmospheric scattering restoration module is thin cloud images and transmission features;
[0084] The output of the atmospheric scattering restoration module is the atmospheric light value and transmittance map;
[0085] The upsampling-convolution module includes bilinear interpolation, a sixth convolution layer, and a seventh convolution layer; the bilinear interpolation, the sixth convolution layer, and the seventh convolution layer are connected in sequence; the input of the upsampling-convolution module is the transmittance feature; the output of the upsampling-convolution module is the transmittance map;
[0086] The transmittance feature is amplified by bilinear interpolation and then input into the sixth and seventh convolutional layers to obtain the transmittance map;
[0087] The U-Net network includes downsampling and upsampling; the input of the U-Net network is a thin cloud image; the output of the U-Net is the atmospheric light value;
[0088] The spatial detail enhancement discriminator is a Transformer discriminator;
[0089] The input of the spatial detail enhancement discriminator is a thin cloud image, a de-clouded image, a reconstructed thin cloud image and a real cloud-free image;
[0090] The reconstructed thin cloud image is calculated using the atmospheric scattering formula as I_fake:
[0091] I_fake=T_tran×J_fake+A_atm(1-T_tran)
[0092] I_fake is the reconstructed thin cloud image, T_tran is the transmittance map, J_fake is the cloud-removed image, and A_atm is the atmospheric light value. The reconstructed thin cloud image is calculated using the atmospheric light value, transmittance map, and cloud-removed image using the atmospheric scattering formula.
[0093] The thin cloud images and real cloud-free images are selected from the remote sensing cloud removal dataset; the remote sensing cloud removal dataset includes the RICE1 dataset, the WHUS2-CR dataset, and the Landsat8 dataset;
[0094] The output of the spatial detail enhancement discriminator is the true and false discrimination results D(J), D(I), D(G(I)) and D(G(J));
[0095] Where D(J) is the prediction of the spatial detail enhancement discriminator on the real cloud-free image J, that is, the probability that the spatial detail enhancement discriminator believes that J is a real image; D(I) is the prediction of the spatial detail enhancement discriminator on the thin cloud image I, that is, the probability that the spatial detail enhancement discriminator believes that I is a real image; D(G(I)) is the prediction of the spatial detail enhancement discriminator on the generated de-clouded image, that is, the probability that the spatial detail enhancement discriminator believes that the de-clouded image is a real image; D(G(J)) is the prediction of the discriminator on the reconstructed pseudo thin cloud image; G(I) is the de-clouded image generated by the feature fusion residual de-clouding generator from the thin cloud image I;
[0096] The total loss function L of the thin cloud removal network that integrates physical prior and feedback enhancement total for:
[0097] L total =λ1L adv +λ2L L1 +λ3L mse
[0098] Among them L adv is the adversarial loss, λ1 is the weight parameter of the adversarial loss; L L1 is the dual reconstruction loss, λ2 is the weight parameter of the reconstruction loss; L mse is the detail enhancement loss, λ3 is the weight parameter of the detail enhancement loss; λ3, λ2 and λ1 are used to balance the contributions of different loss terms; where λ1 = λ2 = 1; λ3 = 20;
[0099] Adversarial loss L adv for:
[0100] L adv =L I2J (G,D,I,J)+L J2I (G,D,I,J)
[0101] Among them, L I2J(G, D, I, J) is the adversarial loss between the de-clouded image and the true cloud-free image, L J2I (G, D, I, J) is the adversarial loss between the thin cloud image and the reconstructed thin cloud image;
[0102] The adversarial loss L between the de-clouded image and the true cloud-free image J2I (G,D,I,J) is:
[0103]
[0104] is the true cloud-free image discrimination expectation; It is to de-cloud the image to deceive expectations;
[0105] The true cloud-free image discrimination expectation is calculated on the true cloud-free image J, p data (J) is the data distribution of the real cloud-free image; D(J) is the prediction of the spatial detail enhancement discriminator for the real cloud-free image J, that is, the probability that the spatial detail enhancement discriminator believes that J is a real image; when D(J) is close to 0, it means that the probability that J is a real image is lower; when D(J) is close to 1, it means that the probability that J is a real image is higher; the deception expectation of the declouded image is calculated on the thin cloud image I, where p data (I) is the data distribution of the thin cloud image; G(I) is the declouded image generated by the feature fusion residual declouding generator from the thin cloud image I. D(G(I)) is the prediction of the spatial detail enhancement discriminator on the generated declouded image, that is, the probability that the spatial detail enhancement discriminator believes that the declouded image is a real image;
[0106] The adversarial loss L between the thin cloud image and the reconstructed thin cloud image J2I (G,D,I,J) is:
[0107]
[0108] is the discriminant expectation of thin cloud image; The deception expectation of the reconstructed thin cloud image; the thin cloud image discrimination expectation is calculated on the thin cloud image I; D(I) is the prediction of the spatial detail enhancement discriminator on the thin cloud image I, that is, the probability that the spatial detail enhancement discriminator believes that I is a real image; when D(I) is close to 0, it means that the probability that I is a real image is lower; when D(I) is close to 1, it means that the probability that I is a real image is higher; the deception expectation of the reconstructed thin cloud image is calculated on the real cloud-free image J; G(J) is the reconstructed thin cloud image generated by the feature fusion residual declouding generator from the real cloud-free image J. D(G(J)) is the prediction of the spatial detail enhancement discriminator on the reconstructed image, that is, the probability that the spatial detail enhancement discriminator believes that the reconstructed thin cloud image is a real image;
[0109] Dual reconstruction loss L L1 for:
[0110]
[0111] is the forward reconstruction loss, is the reverse reconstruction loss;
[0112] Forward reconstruction loss for:
[0113]
[0114] λ4 is the weight parameter of the forward reconstruction loss, which is used to balance the contribution of the forward reconstruction loss in the reconstruction loss; ‖·‖1 represents the L1 norm; the L1 norm calculates the sum of the absolute values of the vector elements; ||G(I)-J||1 measures the difference between the declouded image and the true cloud-free image;
[0115] Reverse reconstruction loss for:
[0116]
[0117] λ5 is the weight parameter of the reverse reconstruction loss, which is used to balance the contribution of the reverse reconstruction loss in the reconstruction loss; A(G(I)) is the reconstructed thin cloud image; ||A(G(I))-I||1 measures the difference between the reconstructed thin cloud image and the thin cloud image;
[0118] Detail enhancement loss L mse :
[0119]
[0120] Where N represents the total number of training set samples, J is the real cloud-free image, and G(I) is the cloud-free image;
[0121] The total loss function guides the network parameter updates of the feature fusion residual declouding generator and the spatial detail enhancement discriminator in each round of training to improve training results. With the increase in training rounds, the quality of the declouded images generated by the feature fusion residual declouding generator gradually improves, and the discriminative ability of the spatial detail enhancement discriminator also continues to increase. This adversarial training mechanism forces the generator and discriminator to compete with each other and improve together, ultimately achieving the following results:
[0122] The feature-fused residual declouding generator is able to produce high-quality declouded images that are visually indistinguishable from real cloud-free images. The spatial detail enhancement discriminator can accurately distinguish between real and generated images, thus effectively guiding the training of the generator.
[0123] When the total loss function reaches convergence, that is, the training error fluctuation range is less than the preset threshold, and the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) of the declouded image both reach the preset performance targets, the training of the thin cloud removal network that integrates physical priors and feedback enhancement is considered complete, and the thin cloud removal network that integrates physical priors and feedback enhancement is saved as the optimal model.
[0124] The preset threshold is 3%;
[0125] Step 2: Use the dataset, training set, and validation set to train the thin cloud removal network that integrates physical priors and feedback enhancement;
[0126] Step 3: Use the trained thin cloud removal network that integrates physical priors and feedback enhancement to perform remote sensing thin cloud removal and identification.
[0127] The present invention will be further described below with reference to the accompanying drawings and examples.
[0128] The present invention features the following steps: constructing a feature fusion residual declouding generator, integrating a feature fusion residual module into the encoder to perform feature accumulation and fusion. The output high-scale feature map undergoes preliminary declouding processing in the first declouding module, enhancing the model's adaptability to the terrain environment and its ability to restore detail. The decoder uses a multi-layer feature decoding module to decode and fuse features layer by layer, and combines this with the second declouding module to refine the declouding effect, further improving image resolution and detail preservation. An atmospheric scattering restoration module is designed, combining a U-Net network and an upsampling-convolution module to accurately estimate atmospheric light values and transmittance maps, providing physical prior knowledge for the declouding process. The parameters of the transmittance map are dynamically optimized through a feedback mechanism between the generator and the discriminator, allowing the network to focus more on cloud distribution and terrain details, improving declouding effectiveness and generalization. The spatial detail enhancement discriminator incorporates a dual reconstruction loss mechanism and a grid attention mechanism to block the input image and generate position-encoded feature maps. Furthermore, Transformer blocks are used to extract features and capture long-range dependencies, enhancing the ability to capture image details and driving the generator to produce higher-quality declouded images. Because the present invention integrates physical prior knowledge with feedback enhancement mechanism, combines multi-scale feature fusion and detail enhancement discriminator, effectively extracts global and local features of the image, and accelerates network convergence by dynamically optimizing the thin cloud area weights and transmittance map, thus achieving high-quality thin cloud removal in complex scenes.
[0129] The steps of the technical solution adopted by the present invention are as follows:
[0130] Step 1: Build data sets, training sets, and validation sets;
[0131] The datasets include the Rice 1 (RICE1), the WHUS2-CR (WHUS2-CR), and the Landsat 8 (Landsat 8). Atmospheric correction and radiometric calibration were performed on the WHUS2-CR and Landsat 8 datasets, and the images were cropped to 256×256 pixels. The training and test sets were divided into two groups with a ratio of 8:2.
[0132] Step 2: Build a thin cloud removal generative adversarial network;
[0133] The thin cloud removal network is based on the conditional generative adversarial network (cGAN) framework, which includes a feature fusion residual declouding generator, an atmospheric scattering restoration module, and a spatial detail enhancement discriminator. Figure 1 As shown in the figure, the thin cloud images in the training set are input into the generator and restoration module respectively. The generator produces a preliminary de-clouded image and transmittance features, while the restoration module uses U-Net to calculate the atmospheric light value. The transmittance features are then input into the upsampling-convolution module of the restoration module to generate a transmittance map. Using the atmospheric scattering model formula, the transmittance map, atmospheric light value, and de-clouded image are combined to reconstruct a pseudo thin cloud image. This image, along with the true cloud-free image, de-clouded image, and thin cloud image, is input into the discriminator for authenticity judgment. The output of the discriminator guides the update of the generator training parameters through the loss function, and the next round of training is repeated.
[0134] Step 3: Initialize the parameters of the thin cloud removal generative adversarial network model;
[0135] The training parameters were initialized as follows: optimizer momentum Adam was 0.9, weight decay coefficient was 0.0001, bias decay coefficient was 0, number of training rounds was 200, batch_size was set to 1, number of data loading threads was 8, and the cloud image size was cropped to 256×256. The initial learning rate was 0.0002, decaying to 0.4 times the original value every 50 epochs. The total loss function weights were [1, 1, 20].
[0136] Step 4: Train the thin cloud removal generative adversarial network model;
[0137] The RICE1, WHUS2-CR, and Landsat8 datasets were respectively input into the thin cloud removal generative adversarial network model. Each dataset was trained for 200 iterations. The trained models were saved, and the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) of each model were calculated. The model with the largest PSNR and structural similarity was the best model.
[0138] Feature fusion residual declouding generator construction:
[0139] The feature fusion residual de-clouding generator is a U-Net architecture, such as Figure 2As shown in Figure 2, the feature fusion residual declouding generator consists of four encoding blocks, one intermediate block, and four decoding blocks. The thin cloud image is first input to the first encoding block, then processed by all modules in sequence, and output from the last decoding block to obtain the declouded result.
[0140] In the first encoding block, the thin cloud image is passed to a convolutional layer with a stride of 1 to extract shallow information. The output is stored in a feature matrix and then passed to a deep network with 16 output channels. The deep network contains three identical residual blocks, and the input of the deep network is summed with its output. In the second encoding block, the feature map out_conv1 output of the first encoding block is passed to a convolutional layer with a stride of 2 and an output channel of 32. The output and the feature matrix are then passed to the feature fusion residual declouding module and the output is stored in the feature matrix. The output is then passed to the deep network and the output is summed with its input. In the third encoding block, the feature map out_conv2 output of the second encoding block is passed to a convolutional layer with a stride of 2 and an output channel of 64. The output and the feature matrix are then passed to the feature fusion residual declouding module and the output is stored in the feature matrix. The output is then passed to the deep network and the output is summed with its input. In the fourth encoding block, the feature map out_conv3 output by the third encoding block is passed to a convolutional layer with a stride of 2 and an output channel of 128. The output and the feature matrix are then passed to the feature fusion residual declouding module, and the output is saved in the feature matrix. The output is then passed to the deep network, and the output is added to its input.
[0141] The middle block consists of a convolutional layer with a stride of 2, a feature fusion residual declouding module, and a first declouding module. The output of the fourth encoding block, out_conv4, is first input to the convolutional layer, with an output channel of 256. This output, along with the feature matrix, is then passed to the feature fusion residual declouding module, where the output is stored in the feature matrix. The output is then passed to the first declouding module, where the feature map is processed by 18 residual blocks, resulting in the output of the high-dimensional restored image, res.
[0142] The first decoding block accepts res as input and first passes it through a deconvolution layer with a stride of 2 and an output channel of 128. The output is then passed to the second declouding module. The output and the feature matrix are then passed to the feature fusion residual declouding module for feature fusion, and the output is saved in the feature matrix. In the second decoding block, the feature map out_deonv1 output by the first decoding block is passed to a deconvolution layer with a stride of 2 and an output channel of 64. It is then passed to the second declouding module. The output of the second declouding module and the feature matrix are passed to the feature fusion residual declouding module, and its output is saved in the feature matrix. In the third decoding block, the feature map out_deonv2 output by the second decoding block is passed to a deconvolution layer with a stride of 2 and an output channel of 32. It is then passed to the second declouding module. The output of the second declouding module and the feature matrix are passed to the feature fusion residual declouding module, and its output is saved in the feature matrix. In the fourth decoding block, the feature map out_deonv3 output by the third decoding block is passed to a deconvolution layer with a stride of 2 and an output channel of 16. It is then passed to the second declouding module. The output of the second declouding module and the feature matrix are passed to the feature fusion residual declouding module, and its output is saved in the feature matrix. Finally, it is passed to a deconvolution layer with a stride of 1, and the output is a declouded image with a channel number of 3.
[0143] The specific structure of the feature fusion residual declouding module is as follows Figure 3 As shown in the figure, the network consists of two branches: a feature fusion block and a DenseNet. This module splits the input feature map into two parts. Both feature maps retain their original height and width, but only have half the number of channels. The feature map with the first half of the channels is fed into the feature fusion block, which includes an encoding block and a decoding block, respectively, consisting of a convolutional layer with a 3×3 convolution kernel and a convolutional layer with a 4×4 convolution kernel, followed by a Pre-ReLU activation function. The decoded feature map is fused with the feature matrix in an intermediate layer and passed to the encoder. The output is then fed into the next feature fusion block for repeated processing. The feature map with the second half of the channels is fed into the DenseNet. It passes through four convolutional layers with 3×3 convolution kernels and ReLU activation functions. Each module receives the input of all previous modules, adds the output to the module, and then feeds the next module. The output is then fed into a convolutional layer with a 1×1 convolution kernel, adds the output to the feature map with the second half of the channels, and passes through a LeakyReLU activation function. Finally, the two processed feature maps are concatenated by channel, and the output is the fused feature.
[0144] The specific structure of the second cloud removal module is as follows: Figure 4As shown in the figure, it includes upsampling, three residual blocks, and a Relu activation function. The feature map output by the deconvolution layer is fed into the upsampling layer for amplification and then added to the output features of the encoding block. Specifically, in the first, second, third, and fourth decoding blocks, the output features of the encoding block added by the second declouding module are out_conv4, out_conv3, out_conv2, and out_conv1, respectively. The output is fed into a residual block, which includes a convolutional layer, a Relu activation function, and a convolutional layer. The output is added to the input and then fed into a LeakyReLU activation function. The output is then passed to two residual blocks with the same structure for further processing. The outputs of these two residual blocks are then element-wise added to the output of the upsampling layer and the output features of the encoding block to fuse multi-scale feature information. The fused features are then fed into a Relu activation function to introduce nonlinearity, thereby enhancing the model's expressiveness. Ultimately, the feature map processed by the Relu activation function is the desired declouding feature map.
[0145] This paper uses dual reconstruction loss and MSE loss to optimize the global consistency and pixel-level accuracy of the image. By calculating the L1 loss and MSE loss functions to guide the generator training, the network declouding result is closer to the real cloud-free image. For the cloud image I_real and the reconstructed cloud image I_fake, as well as the cloud-free image J_real and the reconstructed cloud-free image J_fake, the dual reconstruction loss function is expressed as:
[0146]
[0147] in, and They represent the forward reconstruction loss and the reverse reconstruction loss respectively. λ1 and λ2 are two weight parameters used to adjust the impact of L1 loss in the total loss. The MSE loss function is used to measure the sum of squares of the differences between the model prediction value and the actual observation value. Its calculation formula is defined as:
[0148]
[0149] Among them, N represents the number of samples, J is the true value of the sample, is the value predicted by the model. By minimizing the MSE loss, the model is able to learn a more accurate cloud removal map.
[0150] Atmospheric scattering restoration module construction:
[0151] The atmospheric scattering restoration module comprises a U-Net network and an upsampling-convolution module. This module is based on atmospheric scattering theory and provides physical priors for the network. The upsampling-convolution module consists of one upsampling layer and two convolutional layers. The input transmittance features are amplified using bilinear interpolation and then fed into the convolutional layers to produce a transmittance map. The U-Net network incorporates both downsampling and upsampling to extract global and local information from the cloud image. The downsampling layer comprises seven convolutional layers, each consisting of a 3×3 convolutional kernel, a normalization layer, and an activation layer. The first through seventh convolutional layers are connected sequentially, with each convolutional layer taking a feature map as input and the next layer taking the output of the previous layer as input. The upsampling layer comprises seven neural network layers, each including a deconvolution layer, a normalization layer, and an activation layer. The results of the seventh downsampling layer are gradually amplified using deconvolution and then concatenated channel-by-channel with the results of the first through sixth downsampling layers and the feature map obtained through upsampling to produce the final atmospheric light value. Subsequently, the transmittance map, the cloud removal image, and the atmospheric light value are used to generate a realistic fake cloud image I_fake. The generation of I_fake follows the following formula:
[0152] I_fake=T_tran×J_fake+A_atm(1-T_tran)
[0153] Where T_tran is the transmittance map, J_fake is the declouded image, and A_atm is the atmospheric light value. By generating fake cloud images and calculating the loss with real cloud images, the atmospheric scattering restoration module effectively trains the generator, enabling it to learn more accurate transmittance maps and atmospheric light values, thereby improving declouding performance and generalization capabilities.
[0154] 4. Construction of spatial detail enhancement discriminator:
[0155] This paper proposes a spatial enhancement detail discriminator, which includes two neural network layers and two pooling layers. The neural network layer consists of a pooling layer and a Transformer layer, which work together to extract and process image features. Figure 4The network structure and processing flow of the discriminator are shown. The discriminator inputs include pairs of cloud- and cloud-free images. These images are segmented into 64-pixel patches in an 8×8 format and fed into a Transformer for feature encoding. The feature maps processed by the Transformer are then downsampled through a pooling layer to reduce the spatial dimension of the feature maps while retaining important feature information. These pooled feature maps are then channel-wise concatenated with the 4×4 image patches to form fused feature maps, which are then fed into the next Transformer block. This step enables the network to capture image details at different scales. After processing through the second Transformer block and the corresponding pooling layer, the output feature maps are channel-wise concatenated with the 2×2 image patches. Finally, these feature maps, which have undergone multiple feature extractions and fusions, are fed into two pooling layers to further compress the spatial dimension of the feature maps, facilitating discrimination by the final classifier.
[0156] In this way, the discriminator can gradually reduce the resolution of the feature map and expand the receptive field, effectively fusing local and global information, thereby improving the accuracy of the discrimination. The final output is a binary label indicating the authenticity of the input image pair, accurately distinguishing between real and generated images. Under this architecture, the discriminator can effectively handle images in complex scenes, with higher robustness and accuracy.
[0157] This paper uses adversarial loss to distinguish the authenticity of images. By calculating the adversarial loss function to guide network training, the cloud-free images generated by the generator are more realistic and rich in details. The adversarial loss function is expressed as:
[0158] L adv =L I2J (G,D,I,J)+L J2I (G,D,I)
[0159]
[0160] Where I is the cloud image input to the generator, J is the real cloud-free image, G(I) is the image produced by the generator, D(J) is the discriminator's prediction of the real image, and D(G(I))) is the prediction of the generated image. The adversarial loss of the cloud image reconstruction process can be expressed as:
[0161]
[0162] Among them, A is the atmospheric scattering restoration module.
[0163] The final combined loss function is expressed as:
[0164] L total =Ladv +L L1 +αL mse
[0165] Here, α represents the variance balancing factor. In the experiments in this paper, α is set to 20, and the network is trained by optimizing the combined loss function.
Claims
1. A thin cloud removal method integrating physical prior and feedback enhancement, comprising the following steps; Step 1: Construct a thin cloud removal network that integrates physical priors and feedback enhancement; The thin cloud removal network that integrates physical prior and feedback enhancement is a conditional adversarial generation network cGAN; The thin cloud removal network that integrates physical prior and feedback enhancement includes a feature fusion residual declouding generator, an atmospheric scattering restoration module and a spatial detail enhancement discriminator; The feature fusion residual declouding generator outputs a declouded image; The atmospheric scattering restoration module extracts atmospheric light values and transmittance maps, providing physical prior knowledge. The extracted atmospheric light values, transmittance maps, and cloud-free images are reconstructed to obtain a pseudo-thin cloud image. The true cloud-free image, cloud-free image, thin cloud image, and pseudo-thin cloud image are then input into a spatial detail enhancement discriminator for discrimination. The spatial detail enhancement discriminator outputs a true or false judgment result, which is then used to update the training parameters of the feature fusion residual cloud removal generator through a loss function. Step 2: Use the dataset, training set, and validation set to train the thin cloud removal network that integrates physical priors and feedback enhancement; Step 3: Use the trained thin cloud removal network that integrates physical priors and feedback enhancement to perform remote sensing thin cloud removal and identification.
2. The thin cloud removal method integrating physical prior and feedback enhancement according to claim 1 is characterized by: The feature fusion residual de-clouding generator includes an encoding module, an intermediate block and a decoding module; The input of the feature fusion residual declouding generator is a thin cloud image; the output of the feature fusion residual declouding generator is a declouded image; the data flow between the encoding module, the intermediate block and the decoding module adopts a dense connection; the dense connection is a DenseNet-like connection; The encoding module includes a first encoding block, a second encoding block, a third encoding block and a fourth encoding block; the input of the encoding module is a thin cloud image; the output of the encoding module is a high-dimensional feature map; the first encoding block includes a first convolutional layer and a first deep network layer; the step size of the first convolutional layer is 1; the first deep network layer includes three identical residual blocks; the number of channels of the three identical residual blocks is 16; the second encoding block includes a second convolutional layer, a first feature fusion residual declouding module and a second deep network layer; the second convolutional layer, the first feature fusion residual declouding module and the second deep network layer are connected in sequence; the second convolutional layer is a 3×3 convolution kernel, and the step size of the second convolutional layer is 2; the first feature fusion residual declouding module includes DenseNet and feature fusion blocks; The feature fusion block includes a fifth encoding block and a fifth decoding block; the fifth encoding block includes a 3×3 convolution kernel and a PReLU activation function; the fifth decoding block includes a 4×4 convolution kernel and a PReLU activation function; the DenseNet includes four identical network layers; each network layer includes a fourth convolution layer and a ReLU activation function; the fourth convolution layer is a 3×3 convolution kernel; the structures of the third encoding block and the fourth encoding block are the same as those of the second encoding block; the structure of the second deep network layer is the same as that of the first deep network layer; The intermediate block includes a third convolutional layer, a second feature fusion residual declouding module, and a first declouding module; the third convolutional layer, the second feature fusion residual declouding module, and the first declouding module are connected in sequence; the first declouding module includes eighteen identical residual blocks; the structure of the third convolutional layer is the same as that of the second convolutional layer; the input of the intermediate block is a high-dimensional feature map; the output of the intermediate block is a reconstructed high-dimensional feature map; The structure of the second feature fusion residual declouding module is the same as that of the first feature fusion residual declouding module; The decoding module includes a first decoding block, a second decoding block, a third decoding block and a fourth decoding block; the input of the decoding module is a reconstructed high-dimensional feature map; the output of the decoding module is a declouded image; the first decoding block includes a first deconvolution layer, a second declouding module and a third feature fusion residual declouding module; the first deconvolution layer, the second declouding module and the third feature fusion residual declouding module are connected in sequence; the kernel size of the first deconvolution layer is 3×3, and the stride of the first deconvolution layer is 2; the structure of the second decoding block and the third decoding block are the same as that of the first decoding block; the structure of the third feature fusion residual declouding module is the same as that of the first feature fusion residual declouding module; The second declouding module includes upsampling, three identical residual blocks and an activation function layer; the structure of the second decoding block and the structure of the third decoding block are the same as the structure of the first decoding block; The input of the second declouding module is the feature map output by the first deconvolution layer; The output of the second declouding module is a high-dimensional feature map after declouding. The fourth decoding block is based on the first decoding block and adds a fifth convolutional layer after the third feature fusion residual declouding module. The structure of the fifth convolutional layer is the same as that of the first convolutional layer. The output of the fourth decoding block is the transmission feature and the declouded image. The atmospheric scattering restoration module includes an upsampling-convolution module and a U-Net network; the input of the atmospheric scattering restoration module is a thin cloud image and a transmission feature; the output of the atmospheric scattering restoration module is an atmospheric light value and a transmittance map; The upsampling-convolution module includes bilinear interpolation, a sixth convolution layer, and a seventh convolution layer; the bilinear interpolation, the sixth convolution layer, and the seventh convolution layer are connected in sequence; the input of the upsampling-convolution module is the transmittance feature; the output of the upsampling-convolution module is a transmittance map; the transmittance feature is amplified by bilinear interpolation and then input into the sixth convolution layer and the seventh convolution layer to obtain the transmittance map; the U-Net network includes downsampling and upsampling; the input of the U-Net network is a thin cloud image; the output of the U-Net is the atmospheric light value; The spatial detail enhancement discriminator is a Transformer discriminator; the input of the spatial detail enhancement discriminator is a thin cloud image, a cloud removal image, a reconstructed thin cloud image and a real cloud-free image; The reconstructed thin cloud image I_fake is: I_fake=T_tran×J_fake+A_atm(1-T_tran) T_tran is the transmittance map, J_fake is the cloud removal image, and A_atm is the atmospheric light value; The output of the spatial detail enhancement discriminator is the true and false discrimination results D(J), D(I), D(G(I)) and D(G(J)); among them, D(J) is the prediction of the spatial detail enhancement discriminator for the real cloud-free image J, that is, the probability that the spatial detail enhancement discriminator believes that J is a real image; D(I) is the prediction of the spatial detail enhancement discriminator for the thin cloud image I, that is, the probability that the spatial detail enhancement discriminator believes that I is a real image; D(G(I)) is the prediction of the spatial detail enhancement discriminator for the generated de-clouding image, that is, the probability that the spatial detail enhancement discriminator believes that the de-clouding image is a real image; D(G(J)) is the prediction of the discriminator for the reconstructed pseudo thin cloud image; G(I) is the de-clouding image generated by the feature fusion residual de-clouding generator from the thin cloud image I.
3. The thin cloud removal method integrating physical prior and feedback enhancement according to claim 1 is characterized in that , the total loss function L of the thin cloud removal network that integrates physical prior and feedback enhancement total for: L total =λ1L adv +λ2L L1 +λ3L mse Among them L adv is the adversarial loss, λ1 is the weight parameter of the adversarial loss; L L1 is the double reconstruction loss, λ2 is the weight parameter of the double reconstruction loss; L mse is the detail enhancement loss, λ3 is the weight parameter of the detail enhancement loss; Adversarial loss L adv for: L adv =L I2J (G,D,I,J)+L J2I (G,D,I,J) Among them, L I2J (G, D, I, J) is the adversarial loss between the de-clouded image and the true cloud-free image, L J2I (G, D, I, J) is the adversarial loss between the thin cloud image and the reconstructed thin cloud image; The adversarial loss L between the de-clouded image and the true cloud-free image J2I (G,D,I,J) is: is the true cloud-free image discrimination expectation; It is to de-cloud the image to deceive expectations; The true cloud-free image discrimination expectation is calculated on the true cloud-free image J, p data (J) is the data distribution of the real cloud-free image; when D(J) is close to 0, it means the probability that J is the real image is lower; when D(J) is close to 1, it means the probability that J is the real image is higher; where p data (I) is the data distribution of thin cloud images; The adversarial loss L between the thin cloud image and the reconstructed thin cloud image J2I (G,D,I,J) is: is the discriminant expectation of thin cloud image; The deception expectation of the reconstructed thin cloud image; the discriminant expectation of the thin cloud image is calculated on the thin cloud image I; when D(I) is close to 0, the probability that I is a real image is lower; when D(I) is close to 1, the probability that I is a real image is higher; the deception expectation of the reconstructed thin cloud image is calculated on the real cloud-free image J; G(J) is the reconstructed thin cloud image generated by the feature fusion residual declouding generator from the real cloud-free image J; D(G(J)) is the prediction of the spatial detail enhancement discriminator on the reconstructed image, that is, the probability that the spatial detail enhancement discriminator believes that the reconstructed thin cloud image is a real image; Dual reconstruction loss L L1 for: is the forward reconstruction loss; is the reverse reconstruction loss; Forward reconstruction loss for: λ4 is the weight parameter of the forward reconstruction loss, which is used to balance the contribution of the forward reconstruction loss in the reconstruction loss; ‖·‖1 represents the L1 norm; the L1 norm calculates the sum of the absolute values of the vector elements; ||G(I)-J||1 measures the difference between the declouded image and the true cloud-free image; Reverse reconstruction loss for: λ5 is the weight parameter of the reverse reconstruction loss; A(G(I)) is the reconstructed thin cloud image; ||A(G(I))-I||1 measures the difference between the reconstructed thin cloud image and the thin cloud image; Detail enhancement loss L mse : Where N represents the total number of training set samples, J is the real cloud-free image, and G(I) is the cloud-free image; When the total loss function reaches convergence, that is, the training error fluctuation range is less than the preset threshold, and the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) of the declouded image both reach the preset performance targets, the training of the thin cloud removal network that integrates physical priors and feedback enhancement is considered completed, and the thin cloud removal network that integrates physical priors and feedback enhancement at this time is saved as the optimal model.
4. The thin cloud removal method integrating physical prior and feedback enhancement according to claim 3 is characterized in that ,Furthermore, the preset threshold is 3%.
5. The thin cloud removal method integrating physical prior and feedback enhancement according to claim 3 is characterized in that ,The thin cloud images and real cloud-free images are selected from remote sensing cloud removal datasets; ,remote sensing cloud removal datasets include RICE1 dataset, WHUS2-CR dataset and Landsat8 dataset.
6. The thin cloud removal method integrating physical prior and feedback enhancement according to claim 3 is characterized in that ,λ1=λ2=1;λ3=20; Where λ1 is the weight parameter of the adversarial loss, λ2 is the weight parameter of the dual reconstruction loss, and λ3 is the weight parameter of the detail enhancement loss.
7. A computer-readable storage medium storing a computer program; characterized in that: When the computer program is executed by a processor, a thin cloud removal method integrating physical prior and feedback enhancement according to any one of claims 1 to 6 is implemented.