A Remote Sensing Image Cloud Removal Method Based on a Microwave Guided Diffusion Model
Through the microwave-guided diffusion model combined with ControlNet, the problem of insufficient feature extraction and difficulty in cross-modal fusion of microwave-assisted decloud methods in the existing technology is solved, and the generation of high-quality remote sensing images and accurate recovery of surface information is achieved, and the interpretability and data fusion capabilities of the model are improved.
Patent Information
- Application Number
- CN202510248048.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The existing microwave-assisted decloud-based method based on deep learning has problems such as insufficient SAR feature extraction, difficulty in fusion of cross-modal features, poor model interpretability and limited multi-source data fusion capabilities, which are difficult to meet the needs of actual deployment and large-scale applications.
The microwave-guided diffusion model is used to extract multi-band microwave data features through principal component analysis, and a conditional branch network is introduced into the diffusion model in combination with ControlNet to achieve accurate control of the generated results. The microwave data is used to penetrate the cloud layer, and high-quality images are generated in combination with the diffusion model, retaining the texture and detailed information of the ground object.
It significantly improves the generation quality of remote sensing image cloud-default tasks, realizes the generation of high-quality images and the accurate recovery of surface information, and enhances the interpretability of the model and the ability of multi-source data fusion.
Smart Images

Figure CN119784621B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a method for removing clouds from remote sensing images based on a microwave-guided diffusion model. Background Art
[0002] Optical remote sensing images, with their advantages of strong timeliness, wide coverage, and obvious visual features, are widely used in fields such as land planning, natural resource monitoring, and disaster response. However, cloud occlusion is the main challenge in the application of optical remote sensing images, leading to difficulties in extracting ground information and significantly reducing the utilization value of the images. Synthetic Aperture Radar (SAR), as a microwave active earth observation system, can obtain high-resolution radar images under low visibility weather conditions, being almost unaffected by clouds and meteorological conditions, enabling all-weather observation, and making up for the data missing problem of optical images under complex meteorological conditions.
[0003] Although existing microwave-assisted cloud removal methods based on deep learning have made some progress, there are still many deficiencies. For example, SAR feature extraction is insufficient, cross-modal feature fusion is difficult, the model interpretability is poor, and the multi-source data fusion ability is limited. The time series or multi-angle observation data is not fully utilized, making it difficult to meet the requirements of actual deployment and large-scale applications. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for removing clouds from remote sensing images based on a microwave-guided diffusion model. By performing principal component analysis on multi-band microwave data to extract significant features and inputting them into the diffusion model, while introducing ControlNet, a conditional branch network is added to the diffusion model using ControlNet, and additional conditional information (such as cloud masks, multi-band microwave principal component features, or edge features, etc.) is embedded into specific stages of the diffusion process to achieve precise control of the generated results, thereby further improving the performance of the diffusion model in the task of removing clouds from remote sensing images, realizing high-quality image generation and precise restoration of surface information, and better retaining the texture and detail information of ground objects.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] A method for removing clouds from remote sensing images based on a microwave-guided diffusion model, comprising the following steps:
[0007] S1. Establish a cloud mask: Use the maximum likelihood classification method to establish classification training samples of clouds and shadows, and obtain the mask map of the clouds.
[0008] S2. Analyze and extract multi-band microwave features:
[0009] S21. Expand the pixel values of two microwave images taken at the same location into a matrix form, and perform standardization processing on each column of data so that the mean is 0 and the variance is 1;
[0010] S22. Calculate the covariance matrix of the standardized matrix, and perform eigenvalue decomposition on it to obtain two corresponding eigenvectors;
[0011] S23. Select the eigenvector corresponding to the largest eigenvalue, project the original data onto the direction of the first principal component PC1, and reshape the projection result into an image form to replace the two microwave images in step S21; where the eigenvectors are in turn starting from the original data: , , ,... ;
[0012] S3. Use the microwave-guided diffusion model to train and generate visible light images:
[0013] S31. Forward diffusion: Starting from the original data , the process of transforming the complex real data distribution into a standard Gaussian distribution by gradually adding noise to the data; specifically, forward diffusion starts from the original data (such as an image), and a small amount of random noise is introduced in each step, and the amplitude of this noise increases with the increase of the time step, gradually destroying the structure and content of the original data;
[0014] Forward diffusion starts from the original data , and noise is introduced at each time step t:
[0015] ,
[0016] where, is the original data, is the result after forward diffusion to time t, is the noise attenuation coefficient, which determines the ratio of data and noise in each step, is the standard Gaussian distribution noise;
[0017] In this process, each moment t is only related to the moment t - 1, and is represented by a Markov process:
[0018] ,
[0019] As t increases, gets closer and closer to pure noise; when , is completely Gaussian noise;
[0020] S32. Reverse diffusion: Starting from the random noise Start by following the diffusion steps in sequence to gradually remove noise; at each time step , the diffusion model will predict the noise component based on the current data , and update the data with the estimated result of this noise , thereby approaching the target data distribution; among them, the diffusion model is constructed by a U-Net neural network, and the network architecture of U-Net includes a ControlNet network and a Stable Diffusion network;
[0021] From to , each step will be modeled by a U-Net neural network. This network learns the noise distribution in the forward diffusion process and masters how to recover results closer to the original data from the noise data at any time point. The main calculation formula is:
[0022] ,
[0023] where is the predicted mean, is the predicted variance, is the noise attenuation coefficient in the forward diffusion process, is the cumulative noise attenuation coefficient, is the noise predicted by the neural network;
[0024] The learning method of the U-Net network includes training the ControlNet network and training the Stable Diffusion network; the ControlNet network copies the encoder of the U-Net network into its own network and trains together with the U-Net network through zero convolution without generating interference during the process; the specific steps are:
[0025] S321. Additional condition input: Input the microwave image as an additional condition into the ControlNet network, initialize the weights to zero through zero convolution to ensure that the initial influence of the condition input is controllable and does not interfere with the basic network;
[0026] S322. Conditional Guidance Encoding: The ControlNet network combines and passes the microwave image, noise image, and time step encoding to the encoder module. This part consists of multiple Encoder Blocks, and each block includes the following operations: Convolution operation: Extract input features and compress the spatial resolution; Normalization: Use normalization to standardize the feature distribution; Activation function: Use the GELU function to increase non-linearity; Downsampling: Reduce the size of the feature map through convolution, changing the input image from 64×64 to 32×32, then 16×16, and finally 8×8 in sequence. During the training of ControlNet, the weights of the U-Net network in Stable Diffusion are frozen to maintain the generation ability of the original model.
[0027] S323. Intermediate Feature Processing: At each scale, the output of the ControlNet network is fused into the encoder features of the main U-Net through the SkipConnection module, and then a multi-scale feature map that incorporates conditional input information is output.
[0028] S324. Conditional Decoding: The decoder upsamples the low-resolution features through transposed convolution. Each decoding block increases the feature resolution from 8×8 to 16×16, then to 32×32, and finally restores it to the original resolution of 64×64. At the same time, the Skip Connection module is used to retain the local detailed information at high resolution.
[0029] After these four steps of operations, the prediction from to is completed, and the original image is gradually restored.
[0030] S325. Loss Function: At a given time step t, the text prompt and the microwave image , through the learning network , are used to predict the noise added to the noise image . The loss function is expressed as follows:
[0031] ,
[0032] is the overall learning objective of the entire diffusion model, and this objective can be directly used to fine-tune the diffusion model.
[0033] S33. Style Transfer: Image style transfer based on optimal transport transfers the color style of the target domain visible light image space to the source domain, where the image space is composed of the numerical information of the image plane and color channels. The specific steps are as follows:
[0034] Assume that the color distributions of the source domain and target domain image spaces are and , where represents a Gaussian distribution; optimal transport determines and the closed - form mapping between ,
[0035] where is the transport matrix, and represent source - domain and target - domain samples respectively; the optimal mapping is determined by optimal transport to minimize the distance between the source - domain and target - domain distributions, that is:
[0036] ,
[0037] where represents the cost function; the optimal transport matrix is:
[0038] .
[0039] S4. Replace the cloud image with the microwave - generated image: Use the ControlNet network to generate a visible - light image from the microwave band, and use the mask image established in step S1 to replace the cloud image with the generated part, thus completing the cloud - removing work.
[0040] Compared with the prior art, the present invention has the following advantages or beneficial effects:
[0041] 1. The present invention first proposes a method combining microwave and diffusion model for cloud - removing processing. By utilizing the characteristic that microwaves can penetrate clouds, the interference of clouds on surface observation is eliminated, and at the same time, combined with the generation ability of the diffusion model, the microwave characteristics of the surface can be restored almost realistically.
[0042] 2. To further extract the microwave characteristics of the surface, principal component analysis (PCA) is used to perform spatial dimensionality reduction on multi - band data, effectively improving the generation quality and efficiency of the diffusion model. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as a limitation of the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.
[0044] Figure 1 FIG. is a schematic flow chart of a remote - sensing image cloud - removing method based on a microwave - guided diffusion model;
[0045] Figure 2Schematic diagram of the training process of the diffusion model;
[0046] Figure 3 Schematic diagram of the training of the ControlNet network;
[0047] Figure 4 Diagram showing the generation results of C, Ka, and PC1;
[0048] Figure 5 Diagram showing the comparison of the results with or without microwave guidance. Detailed implementation method
[0049] Next, the technical solutions in the embodiments of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0050] As Figure 1 shown, the present invention provides a method for removing clouds from remote sensing images based on a microwave-guided diffusion model. When there are clouds in the remote sensing satellite images, the corresponding microwave band is used as a conditional input into the latent space, and the diffusion model is used to learn the denoising method from each time point to the next time point during the reverse diffusion process, so as to realize using microwave data to guide the model to obtain a remote sensing image close to the real from a noise map and replace the part with clouds in the original image to complete the cloud removal work. In addition, in order to better extract the data features of microwaves, the present invention uses principal component analysis to gather the features of the microwave C band and Ka band into the first principal component, and uses the first principal component to guide the removal of clouds in the visible light band. The specific steps are as follows:
[0051] S1. Establish a cloud mask:
[0052] Adopt the method of maximum likelihood classification to establish the classification training samples of clouds and shadows, and obtain the mask map of clouds.
[0053] S2. Analyze and extract multi-band microwave features:
[0054] S21. Expand the pixel values of two microwave images taken at the same location into a matrix form, and perform standardization processing on each column of data so that the mean is 0 and the variance is 1.
[0055] S22. Calculate the covariance matrix of the standardized matrix, and perform eigenvalue decomposition on it to obtain the corresponding two eigenvectors.
[0056] S23. Select the eigenvector corresponding to the largest eigenvalue, project the original data onto the direction of the first principal component PC1, and reshape the projection result into an image form to replace the two microwave images in step S21; among them, the eigenvectors start from the original data in turn as: 、 、 、... 。
[0057] S3. Use the microwave-guided diffusion model to train (as Figure 2 shown) to generate visible light images:
[0058] S31. Forward diffusion: Starting from the original data and gradually adding noise to the data to transform the complex real data distribution into a standard Gaussian distribution; specifically, forward diffusion starts from the original data (such as an image), and in each step, a small amount of random noise is introduced, and the amplitude of this noise increases with the increase of the time step, gradually destroying the structure and content of the original data.
[0059] Forward diffusion starts from the original data and introduces noise at each time step t:
[0060] ,
[0061] where is the original data, is the result after forward diffusion to time t, is the noise attenuation coefficient, which determines the ratio of data and noise in each step, is the standard Gaussian distribution noise.
[0062] In this process, each moment t is only related to the moment t - 1, which is represented by a Markov process:
[0063] ,
[0064] As t increases, gets closer and closer to pure noise; when , is completely Gaussian noise.
[0065] S32. Reverse diffusion: Starting from the random noise and in the order of the diffusion steps , gradually remove the noise; at each time step , the diffusion model will predict the noise component based on the current data and update the data with the estimated result of this noise, so as to approximate the target data distribution; among them, the diffusion model is constructed by a U-Net neural network, and the network architecture of U-Net includes a ControlNet network and a Stable Diffusion network.
[0066] From to In this process, each step is modeled by the U-Net neural network. By learning the noise distribution in the forward diffusion process, this network masters how to recover results closer to the original data from the noise data at any time point. The main calculation formula is as follows:
[0067] ,
[0068] Among them, is the predicted mean, is the predicted variance, is the noise attenuation coefficient in the forward diffusion process, is the cumulative noise attenuation coefficient, is the noise predicted by the neural network.
[0069] The learning method of the U-Net network includes training the ControlNet network and training the Stable Diffusion network; as Figure 3 shown, the ControlNet network copies the encoder of the U-Net network into its own network and trains together with the U-Net network through zero convolution without generating interference during the process; the specific steps are as follows:
[0070] S321. Additional condition input:
[0071] Input the microwave image as an additional condition into the ControlNet network, and initialize the weights to zero through zero convolution to ensure that the initial influence of the condition input is controllable and does not interfere with the basic network.
[0072] S322. Conditional guided encoding:
[0073] The ControlNet network combines and transmits the microwave image, noise image, and time step encoding to the encoder module; this part consists of multiple Encoder Blocks, and each block includes the following operations: Convolution operation: Extract input features and compress the spatial resolution; Normalization: Use normalization to standardize the feature distribution; Activation function: Use the GELU function to increase non-linearity; Downsampling: Reduce the size of the feature map through convolution, and change the input image from 64×64 to 32×32, 16×16, 8×8 in sequence; When training ControlNet, the weights of the U-Net network of Stable Diffusion are frozen to maintain the generation ability of the original model.
[0074] S323. Intermediate feature processing:
[0075] At each scale, the output of the ControlNet network is fused into the encoder features of the main U-Net through the Skip Connection module, and then a multi-scale feature map fused with conditional input information is output.
[0076] S324. Conditional Decoding:
[0077] The decoder upsamples the low-resolution features through deconvolution; each decoding block increases the feature resolution from 8×8 to 16×16, then to 32×32, and finally restores it to the original resolution of 64×64. At the same time, the SkipConnection module is used to retain the local detailed information of the high resolution.
[0078] After these four operations, the prediction from to is completed, and the original image is gradually restored.
[0079] S325. Loss Function:
[0080] At a given time step t, the text prompt and the microwave image , through the learning network , are used to predict the noise added to the noisy image . The loss function is expressed as follows:
[0081] ,
[0082] is the overall learning objective of the entire diffusion model, and this objective can be directly used to fine-tune the diffusion model.
[0083] S33. Style Transfer: Image style transfer based on optimal transport transfers the color style in the visible light image space of the target domain to the source domain, where the image space consists of the numerical information of the image plane and color channels. The specific steps are as follows:
[0084] Assume that the color distributions in the image spaces of the source domain and the target domain are and , respectively, where represents the Gaussian distribution; optimal transport determines the closed-form mapping between and , which satisfies , that is:
[0085] ,
[0086] where is the transport matrix, and represent the source domain and target domain samples respectively; the optimal mapping is determined through optimal transport to minimize the distance between the source domain and target domain distributions, that is:
[0087] ,
[0088] Among them, represents the cost function; the optimal transport matrix is:
[0089] ,
[0090] The specific steps of optimal transport in the image space are as follows:
[0091] (1) Statistically analyze the color space distribution histograms of the source domain and target domain images to obtain the color distribution parameters of the source domain and target domain , , , ;
[0092] (2) Calculate the optimal transport matrix ;
[0093] (3) Transform the source domain image to obtain the source domain image with the color style of the target domain.
[0094] S4. Replace the cloud image with the microwave generated image:
[0095] Utilize the ControlNet network to generate visible light images in the microwave band, and use the mask image established in step S1 to replace the cloud image with the generated part, thereby completing the cloud removal work.
[0096] Conduct the following experiments using the remote sensing image cloud removal method provided in the above embodiment:
[0097] 1. Effectiveness of extraction by principal component analysis:
[0098] Table 1 Comparison of C, Ka, and PC1 band indicators
[0099] Index C Ka PCA (C / Ka) SSIM 0.515494667 0.506486111 0.557158667 PSNR 17.56272222 16.46173333 18.38193333
[0100] The above three groups of comparative experiments are all the results obtained after running with a training step of 10,000 steps, ensuring the comparability of the index results.
[0101] From the data in the table, it can be seen that the PCA (C / Ka) method performs better in terms of indicators than the data fusion using only the C-band or Ka-band data. Specifically, the SSIM value of the PCA method is 0.5572, significantly higher than 0.5155 of the C-band and 0.5065 of the Ka-band, indicating that the structure similarity of the image generated after PCA fusion is higher and can more accurately reflect the structural characteristics of the original image. In addition, the PSNR value of PCA is 18.3819, also higher than 17.5627 of the C-band and 16.4617 of the Ka-band, indicating that the image after PCA method fusion has a higher signal-to-noise ratio, lower noise level, and better image quality. These results show that the PCA fusion method significantly improves the clarity and detail fidelity of the generated image by effectively integrating the feature information of the C-band and Ka-band, and it is a more superior data fusion method.
[0102] Combined with Figure 4 it can be known that the advantages of principal component generation in denoising and detail enhancement are particularly significant: the building outlines are clear, the vegetation area textures are rich, and the boundaries of ground objects are accurate, while the C-band and Ka-band are weak in detail retention, with obvious noise or blurred areas. This effect shows that the principal component generation method not only improves the image quality by integrating the information of the C-band and Ka-band, but also retains richer spatial and spectral features, and its comprehensive performance is significantly better than the generation results of a single band.
[0103] 2. Comparison of the generation effects with or without microwave diffusion:
[0104] Table 2 Images generated with or without microwave guidance
[0105] Index No Yes SSIM 0.34224225 0.6345115 PSNR 15.511325 20.730675
[0106] The above two groups of experiments are all the results obtained after running with a training step of 40,000 steps to ensure the comparability of the index results.
[0107] From the table, it can be seen that after using microwave guidance, the SSIM increases from 0.3422 to 0.6345, and the PSNR increases from 15.5113 to 20.7307, indicating that microwave guidance significantly improves the quality of the generation results. The increase in SSIM reflects the enhancement of the structural similarity of the generated image, indicating that microwave guidance can better capture the detail information and texture features of the image; while the increase in PSNR indicates that the noise level of the generated image decreases and the signal quality improves significantly. This shows that introducing microwave guidance can not only strengthen the restoration ability of the image content, but also effectively improve the clarity and authenticity of the generation results, further verifying the important role of microwave guidance in multi-source data fusion.
[0108] From Figure 5As can be seen, in the case of cloud cover, there are significant differences between the non-microwave guidance and microwave guidance methods in eliminating cloud interference. Although the images generated by the non-microwave guidance method can partially restore ground information, obvious afterimages or information loss still exist in the cloud area, especially in complex terrain and building areas, and the detail restoration is not accurate enough. The microwave guidance method, on the other hand, shows better cloud removal ability. The generated images have richer detail information in the cloud area and clearer ground object features. Especially in areas such as forests, buildings, and roads, the restoration effect is significantly better than that of non-microwave guidance. The improvement of this effect indicates that microwave guidance can effectively utilize the ability of microwave data to penetrate clouds to make up for the problem of information loss in optical images under cloud cover, significantly improving the image restoration quality and environmental perception ability. At the same time, after introducing the style transfer method, the visual consistency and detail performance of the generated images are further enhanced. Style transfer can effectively combine the structural features of microwave images and the color information of optical images to achieve more natural texture generation and ground object restoration in the cloud-covered area. Especially in complex scenes (such as dense vegetation, urban building groups, and intersecting roads), the style transfer method can more accurately capture the spatial relationship of ground object features, and the generated images not only have a high visual realism macroscopically but also can retain rich texture features microscopically.
[0109] A method for removing clouds from remote sensing images based on a microwave-guided diffusion model provided in this embodiment includes: obtaining a cloud mask: using the maximum likelihood classification method to classify and identify the cloud regions in the optical remote sensing images, thereby generating a cloud mask. By the spectral characteristics of the optical images, combined with the prior probability distribution and within-class variance of the ground object categories, the cloud regions in the images are accurately segmented. The cloud mask marks the parts of the images covered by clouds, providing an accurate regional scope for subsequent microwave guidance and image replacement; establishing a dataset corresponding to microwave and visible light, and establishing a diffusion model for generating visible light guided by microwave: collecting remote sensing data with both microwave images and optical images, establishing a paired training dataset, and through spatial registration and time synchronization, ensuring the consistency of the microwave images and optical images at the pixel level. Based on this dataset, training a diffusion model with the microwave image as the conditional input, the model can learn the mapping relationship from the microwave image to generate high-quality visible light images. The generated images have both the structural characteristics of the microwave images and the texture features of the optical images. During the training process, a style transfer module is introduced to extract the texture style features of the microwave images through a pre-trained convolutional neural network and combine them with the content features of the optical images to ensure the texture consistency and visual quality of the generated images in the cloud-obscured regions; replacement and restoration of cloud-obscured regions: the generated cloud-free optical images will contain complete ground object information, but only the cloud-obscured regions are the parts that need to be restored. Through the cloud mask generated in the first step, accurately locate the cloud-covered regions in the optical images, and crop out the cloud-free parts at the corresponding positions of the generated images. These cropped parts contain the ground object details in the cloud-covered regions, generated with the structural guidance provided by the microwave images, with high-quality texture details and authenticity. After cropping, replace the corresponding cloud regions in the original optical images.
[0110] Experimental results show that compared with traditional cloud removal methods, the model combining microwave guidance and style transfer shows significant advantages in ground object restoration accuracy, image continuity, and overall style consistency.
[0111] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the part of the module, program segment, or code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based device that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A method for removing clouds from remote sensing images based on a microwave-guided diffusion model, characterized in that, It includes the following steps: S1. Establish a cloud mask: Adopt the maximum likelihood classification method to classify and identify the cloud area in the optical remote sensing image, so as to generate a cloud mask map; S2. Analyze and extract multi-band microwave features: S21. Expand the pixel values of two microwave images taken at the same position into a matrix form, and perform normalization processing on each column of data to make the mean 0 and the variance 1; S22. Calculate the covariance matrix of the normalized matrix, and perform eigenvalue decomposition on it to obtain two corresponding eigenvectors; S23. Select the eigenvector corresponding to the largest eigenvalue, project the original data onto the direction of the first principal component PC1, and reshape the projection result into an image form to replace the two microwave images in step S21. Among them, the eigenvectors starting from the original data are successively: , , ... ; S3. Use the microwave-guided diffusion model to train and generate visible light images: S31. Forward diffusion: Starting from the original data and gradually adding noise to the data, the process of transforming the complex real data distribution into a standard Gaussian distribution; The specific steps of S31 are: Forward diffusion starts from the original data and introduces noise at each time step t : , Among them, is the original data, is the result after forward diffusion to time t, is the noise attenuation coefficient, which determines the ratio of data to noise in each step, is the standard Gaussian distribution noise; In the process of forward diffusion, each moment t is only related to the moment t - 1, which is represented by a Markov process: , As t increases, it gets closer and closer to pure noise; when , it is completely Gaussian noise; S32. Reverse Diffusion: Starting from random noise and following the order of diffusion steps , gradually remove the noise step by step; at each time step , the diffusion model predicts the noise component based on the current data and updates the data with the estimated result of the noise so as to approximate the target data distribution; wherein, the diffusion model is constructed by a U-Net neural network, and the network architecture of U-Net includes a ControlNet network and a Stable Diffusion network; The specific steps of S32 are: From to in each step, it will be modeled by the U-Net neural network. By learning the noise distribution in the forward diffusion process, this network masters how to recover results closer to the original data from the noise data at any time point. The main calculation formula is: Among them, is the predicted mean, is the predicted variance, is the noise attenuation coefficient in the forward diffusion process, is the cumulative noise attenuation coefficient, is the noise predicted by the neural network; The learning method of the U-Net network includes training the ControlNet network and the Stable Diffusion network; the ControlNet network copies the encoder of the U-Net network into its own network and trains it together with the U-Net network through zero convolution, and will not generate interference during the process; the specific steps are: S321. Additional condition input: Input the microwave image as an additional condition into the ControlNet network, and initialize the weight to zero through zero convolution to ensure that the initial influence of the condition input is controllable and does not interfere with the basic network; S322. Conditional guided encoding: The ControlNet network combines and transmits the microwave image, the noise image, and the time step encoding to the encoder module; this part consists of multiple Encoder Blocks, and each block includes the following operations: Convolution operation: Extract the input features and compress the spatial resolution; Normalization: Use normalization to standardize the feature distribution; Activation function: Use the GELU function to increase non-linearity; Downsampling: Reduce the size of the feature map through convolution, and change the input image from 64×64 to 32×32, 16×16, 8×8 in sequence; When the ControlNet is trained, the weights of the U-Net network of Stable Diffusion are frozen to maintain the generation ability of the original model; S323. Intermediate feature processing: At each scale, the output of the ControlNet network is fused into the encoder features of the main U-Net through the Skip Connection module, and then a multi-scale feature map fused with the conditional input information is output; S324. Conditional decoding: The decoder upsamples the low-resolution features through transposed convolution; each decoding block increases the feature resolution from 8×8 to 16×16, then to 32×32, and finally restores it to the original resolution of 64×64; at the same time, use the SkipConnection module to retain the high-resolution local detail information; S325, Loss Function: At a given time step t, the text prompt and the microwave image , through the learning network , are used to predict the noise added to the noisy image . The loss function is expressed as follows: , It is the overall learning objective of the entire diffusion model, and this objective can be directly used to fine-tune the diffusion model; S33. Style transfer: The image style transfer based on optimal transport transfers the color style of the target domain visible light image space to the source domain, where the image space is composed of the numerical information of the image plane and the color channels; The specific steps of S33 are: Suppose the color distributions of the source domain and target domain image spaces are respectively and , where represents a Gaussian distribution; Optimal transport determines a closed-form mapping between and that satisfies , that is: , where is the transfer matrix, and represent source domain and target domain samples respectively; the optimal mapping is determined by optimal transport to minimize the distance between the source domain and target domain distributions, i.e.: , Among them, represents the cost function; the optimal transport matrix is: ; S4. Replace the cloud map with the microwave-generated map Using the ControlNet network, generate visible light images in the microwave band, and use the mask map established in step S1 to replace the cloud map with the generated part, thereby completing the cloud removal work.