Remote Sensing Image Cloud Removal Method Based on the Fusion of Synthetic Aperture Radar and Visible Light

The feature decoupling and self-attention-based approach in multi-modal image fusion enhances cloud removal accuracy and consistency by leveraging optical image content in SAR images, addressing the challenges of modal differences in existing methods.

CN116993598BActive Publication Date: 2025-07-15NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310543571.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-15
Publication Date
2025-07-15
Estimated Expiration
2043-05-15

AI Technical Summary

Technical Problem

When the existing multimodal remote sensing image declouding method uses synthetic aperture radar and optical image information, there are problems such as large image differences, difficulty in accurately obtaining cloud locations, and insufficient effective information fusion, resulting in poor cloud removal effect.

Method used

A multimodal image decloud algorithm based on feature decoupling and a modal difference supplement mechanism of self-attention are adopted. By decoupling the content information of different modes, a modal difference supplement mechanism is designed, and the content information under optical modes is used for pixel-by-pixel correction, which improves image fusion efficiency and accuracy.

Benefits of technology

It significantly reduces the difficulty of multimodal image fusion tasks, improves the accuracy and robustness of cloud removal results, and can be quickly applied under extreme weather conditions and can be put into actual use without fine-tuning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116993598B_ABST
    Figure CN116993598B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for removing clouds from remote sensing images based on the fusion of synthetic aperture radar and visible light, belonging to the fields of computer vision, image restoration and enhancement. This method decouples the same content information and modal features expressed by paired images of different modalities to reduce the differences between multi-modal images; a modal difference compensation mechanism is designed, which can make full use of the content information in the optical modality to correct the content information in the SAR modality pixel by pixel, thereby significantly improving the consistency inside and outside the restored area. Using the method of the present invention, the difficulty of the multi-modal fusion cloud removal task can be greatly reduced, and the accuracy of the image cloud removal result can be improved. Moreover, the model has strong robustness to changes in cloud shape and size, so in extreme meteorological observation scenarios, it can be quickly put into practical application without fine-tuning.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Field of the Invention

[0002] The present invention belongs to the fields of computer vision, image restoration and enhancement, and particularly relates to a method for removing clouds from remote sensing images based on multimodal image restoration and contrast learning. Background Art

[0003] Optical images (RGB) do not contain speckle noise and geometric distortion, and have high readability by the naked eye. Therefore, using an optical camera to image the ground is the most common and intuitive means of remote sensing for ground observation. However, since the optical camera uses a passive shooting mode, that is, objects on the earth's surface reflect sunlight into the lens of the optical camera for imaging, and light from the ultraviolet band to the infrared band cannot penetrate thick clouds, clouds will block ground objects in the acquired optical images, seriously reducing the quality of ground observation by the optical remote sensing camera.

[0004] High-resolution synthetic aperture radar imaging is another means of ground observation. It uses an active imaging mode, actively emits microwave signals to the earth's surface, and the radar sensor receives the microwave signals reflected by the target ground objects to obtain an image. Since microwave signals can penetrate clouds, synthetic aperture radar can perform ground observation day and night under any weather conditions, effectively overcoming the disadvantages of optical cameras. However, synthetic aperture radar images (SAR) also have problems such as speckle noise, geometric distortion, and poor readability by the naked eye.

[0005] Although synthetic aperture radar images and optical images have complementary information characteristics, and information fusion is helpful for ground observation, due to the huge differences in the imaging principles between multi-source remote sensing devices, the captured images will have large differences in style, content, features, etc. At the same time, there are many types of ground objects in the scene, with complex structures and large scale differences, which pose great challenges to the multi-source remote sensing image fusion work. Therefore, how to construct a more effective multi-modal data fusion model, correctly match various ground object features in the image, and ensure the authenticity of the image restoration result and the robustness of the model under different weather conditions has very important value. Currently, according to the different ways of using multi-source image features, related work can be divided into the following two categories:

[0006] The first type is the method based on image translation between multi-source images. Such methods commonly use the conditional generative adversarial network model (cGANs, conditional Generative Adversarial Networks) to perform image translation on the input synthetic aperture radar image to obtain the corresponding optical image, and then replace the pixels in the occluded area of the image to be restored with the corresponding pixels in the model output image, so as to achieve the task of cloud removal from the optical image. Bermudez et al. in the literature "Bermudez J.D, Happ P.N, Oliveira D.A.B, and Feitosa R.Q. SAR to Optical Image Synthesis for Cloud Removal with Generative Adversarial Networks[J]. ISPRS Annals of Photogrammetry, Remote Sensing & Spatial Information Sciences, 4(1): 5-11, 2018." cut the unoccluded area in the optical image and the corresponding area in the SAR image into image patches, and then put them into the cGAN model for supervised training. Finally, all the image patches predicted by the model are stitched together to form the target image. Turnes et al. in the paper "Turnes J, Castro J, Torres D, Vega P, Feitosa R, and Happ P. Atrous cGAN for SAR to Optical Image Translation[J]. IEEE Geoscience and Remote Sensing Letters, 19: 1-5, 2022." used the cGAN model, and in the encoding stage of the generator, a ResNet network with dilated convolutional layers was used for feature extraction, and then an atrous spatial pyramid pooling (ASPP) module and a global average pooling layer were added in parallel to extract multi-scale feature information. Adding multiple ASPP modules in the discriminator can improve the performance of the discriminator.Li et al. "Li Y, Fu R, Meng X, Jin W, and Shao F. A SAR-to-Optical Image Translation Method Based on Conditional Generation Adversarial Network (cGAN) [J]. IEEE Access, 8: 60338-60343, 2020." added a structural similarity constraint to the loss function of cGAN, which can reduce the structural information difference between images. However, in the training process of this type of method, only SAR images are input, so the corresponding relationship between SAR images and RGB images, as well as the valid information not occluded in RGB images, are not fully utilized, resulting in a general restoration effect.

[0007] The second type is the method based on multi-source image fusion and restoration. Such methods usually pair multi-source images and input them into the model together for end-to-end processing, and finally output the RGB image predicted by the model. Grohnfeldt et al. proposed in the paper "Grohnfeldt C, Schmitt M, Zhu X. A Conditional Generative Adversarial Network to Fuse SAR and Multispectral Optical Data for Cloud Removal from Sentinel-2 Images[C]. IEEE International Geoscience and Remote Sensing Symposium, 1726-1729, 2018." to discard the random noise that should be input at the generator end, and input the RGB image occluded by clouds and the SAR image together as conditional information into the generator to predict the RGB image. Meraner et al. used a fully convolutional neural network based on the residual structure in the paper "Meraner A, Ebel P, Zhu X, and Schmitt M. Cloud Removal in Sentinel-2 Imagery Using A Deep Residual Neural Network and SAR-Optical Data Fusion[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 166:333-346, 2020.", and input the paired SAR image and multi-spectral image into the network together to obtain the cloud-free optical image. However, in the process of image fusion, it is difficult to explicitly obtain the location of the clouds, and the effective information in the optical image will be interfered by the SAR image during the fusion process, resulting in inefficient utilization of multi-modal information. Summary of the Invention

[0008] The purpose of the present invention is to improve the accuracy of multimodal image declouding results and utilize the effective information in RGB images in a more efficient and accurate manner. The present invention proposes a multimodal image declouding algorithm based on feature decoupling and a modal difference supplementation mechanism based on self-attention. The main idea of the method is to decouple the same content information expressed by paired images of different modalities from the modal features to reduce the differences between multimodal images; design a modal difference supplementation mechanism that can make full use of the content information under the optical modality and correct the content information under the SAR modality pixel by pixel, thereby significantly improving the consistency inside and outside the repair area. Using the method of the present invention, the difficulty of the multimodal fusion declouding task can be greatly reduced, and the accuracy of the image declouding results can be improved. In addition, the model is very robust to changes in cloud shape and size, so it can be quickly put into practical application in extreme meteorological observation scenarios without fine-tuning.

[0009] The technical solution of the present invention is to propose a remote sensing image cloud removal method based on synthetic aperture radar and visible light fusion, which mainly includes the following steps:

[0010] 1. Draw a binary mask image to simulate cloud occlusion. The pixels in the mask M are set to 0 or 255. 0 means the current pixel is blocked by clouds, and 255 means the current pixel is valid information. Then the mask M is combined with the clean RGB image through alpha channel blending to synthesize the simulated RGB image X blocked by clouds. R At the same time, some real cloud images are selected from the 38-Cloud dataset and the corresponding binary mask labels M are used. r After filtering the ground background information, the real image X blocked by clouds is generated by mixing it with the clean RGB image through the alpha channel method. Rr , the pixel value in this mask is also 0 or 255, with the same meaning as above.

[0011] 2. Design a multimodal image declouding algorithm based on feature decoupling, which can decouple the same content information expressed by the paired SAR images and RGB images from the features of the SAR modality and RGB modality respectively, so as to reduce the difference between images of different modalities and improve the efficiency of multimodal image fusion. The algorithm consists of an encoder and a decoder:

[0012] The encoder is composed of a content feature extractor under SAR mode A content feature extractor in RGB mode and a style feature extractor in RGB mode composition.

[0013] Content Feature Extractor It consists of two dilated convolutions with a stride of 2 and four residual blocks built by dilated convolutions. The former is used for downsampling the feature map and extracting shallow features in the image; the latter is used for extracting deep semantic features. Each layer of convolution in the content extractor is paired with an Instance Normalization layer to remove the style information in the feature map.

[0014] Style Feature Extractor It consists of two dilated convolutions with a stride of 2, a global pooling layer, and a 1×1 convolution. The dilated convolution with a stride of 2 is used for downsampling; the global average pooling layer compresses the feature map into a feature vector that can represent the style of each channel to remove the spatial information in the image; finally, the 1×1 convolution is used to concentrate and purify the above feature vector to obtain the final style feature vector.

[0015] The decoder first uses a Multilayer Perceptron (MLP) to convert the style features of the RGB image into the parameters required for the corresponding AdaIN layer, and then passes the parameters to a group of residual blocks configured with AdaIN normalization layers to perform style rendering on the content features in the RGB modality. Finally, the upsampling module composed of alternating nearest neighbor upsampling layers and ordinary convolutions converts the feature map into a cloud-free RGB image.

[0016] 3. Design a self-attention-based modal difference supplement mechanism that can make full use of the effective content information in the RGB modality to correct the content information in the SAR modality pixel by pixel, improve the consistency of the content information extracted from multi-modal images, and ultimately enhance the consistency inside and outside the repaired area. This step includes the following sub-steps:

[0017] Sub-step 301: Calculate the difference information C between the feature maps in the RGB and SAR modalities DF , that is, C DF = C R - C S ⊙ M C , where C S is the content feature refined from the SAR image, C R is the content feature of the RGB image, M C is a binary mask predicted by the dilated convolution to describe the invalid area in the current feature map, and ⊙ is the Hadamard product of matrices.

[0018] Sub-step 302: Use a gating function composed of a Sigmoid function and a 1×1 convolution Conv 1×1 to adaptively select the information in C DF to obtain the supplementary information C PC , which is expressed by the formula as CPC = C DF ⊙ Sigmoid(Conv 1×1 (C DF ))。

[0019] Sub-step 303: Use the self-attention mechanism to model the relationships between the pixels in C S and, based on these relationships, convert the known supplementary information C in the valid region into the required supplementary information in the region to be repaired, that is PC where where the subscript i represents the position of the currently queried pixel, and the subscript j traverses all N p pixels at all positions in the attention map. W v1 , W v2 and W v3 are all linear transformation matrices, ReLU is the activation function, and LN represents layer normalization (LayerNormalization).

[0020] 4. Embed the self-attention-based modality difference supplementation mechanism into the feature decoupling-based multi-modal image de-clouding algorithm, and train this network to finally obtain a feature decoupling-based multi-modal image de-clouding model. The objective function of the feature decoupling-based multi-modal image de-clouding algorithm includes two image reconstruction losses, three feature reconstruction losses, one modality invariance loss, and two adversarial losses. The following will introduce in detail each loss term that composes the objective function:

[0021] (1) Image reconstruction loss

[0022]

[0023]

[0024] Here, the modality conversion image and the RGB reconstruction image where D cos (·,·) represents the cosine distance, and Ψ(·) represents the self-attention-based modality difference supplementation mechanism. D R is the RGB decoder.

[0025] (2) Feature reconstruction loss

[0026]

[0027]

[0028]

[0029] Here, X SRThe content features are supplemented with information via Ψ(·) to obtain content features and style features ψ l (·) represents the i-th layer of the VGG-19 network, and the corresponding weight w l are 0.125, 0.25, and 1 respectively.

[0030] (3) Modal invariance loss

[0031]

[0032] Here C SR = Ψ(C S ).

[0033] (4) Adversarial loss

[0034]

[0035]

[0036] Here D is a discriminator for discriminating the authenticity of RGB images.

[0037] The finally obtained loss function is composed of the above eight loss terms and is summarized into the following formula:

[0038]

[0039] Among them, we stipulate that and λ x , λ c , λ s and λ inv are hyperparameters for balancing each loss term and are respectively specified as 10, 30, 3000, and 10 in this model.

[0040] 5. Apply the multi-modal image de-clouding model based on feature decoupling obtained after training to the actual extreme meteorological scenario to complete the remote sensing image de-clouding task based on the fusion of synthetic aperture radar and visible light.

[0041] The present invention has the following advantages compared with the existing multi-modal image de-clouding algorithms:

[0042] 1. It has a very high fitting efficiency for the modal mapping relationship. By mining the shared content information from multi-modal images and decoupling it from the expression methods of the modalities themselves, compared with the existing methods that blindly increase the number of feature channels to promote fitting, the present invention can greatly reduce the differences between modalities, lower the difficulty of the multi-modal image fusion task, and more effectively fit the mapping relationship between modalities.

[0043] 2. The improvement in the consistency between the inside and outside of the repaired area is obvious. It uses the self-attention mechanism to convert the effective content features in the optical modality into the supplementary information required for the content features in the SAR modality. Compared with the existing methods that ensure the consistency between the inside and outside of the repaired area by splicing feature maps of multi-scale receptive fields, the present invention can utilize the content information in the optical modality more precisely, achieving a significant performance improvement at the cost of extremely small parameter quantity and computing time, while avoiding manually designing hyperparameters.

[0044] 3. The model has a small number of parameters and strong robustness. The number of parameters required by the present invention is much smaller than that of other existing methods, which can reduce the requirements for equipment. At the same time, it is very robust to the changes in the size and shape of clouds, can be trained using simulated occluded images, and can be directly applied to real cloud removal tasks without fine-tuning, which will greatly reduce the complexity of the training set and improve the practicality of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions implemented in the present invention, the following will briefly introduce each module required in the example description. Obviously, the drawings in the following description are only the flowcharts of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be extended based on this drawing.

[0046] Figure 1 is the flowchart of the remote sensing image cloud removal method based on the fusion of synthetic aperture radar and visible light proposed by the present invention;

[0047] Figure 2 is the schematic diagram of the multi-modal image cloud removal algorithm based on feature decoupling in the embodiment;

[0048] Figure 3 . is the schematic diagram of the self-attention-based modal difference supplementary mechanism in the embodiment;

[0049] Figure 4 . is the data flow diagram involved in the model training and inference stages. The data flow involved in both the training and inference stages is shown in (a). The data flow unique to the training stage is shown in (b);

[0050] Figure 5 is the schematic diagram of the cloud mask, optical image, and image to be repaired in the embodiment. Row: respectively synthesize the simulated RGB image occluded by clouds and the real RGB image occluded by clouds. Column: binary mask map, clean RGB image, and synthesized RGB image to be repaired.

[0051] Figure 6It is a schematic diagram of the cloud removal result of the model in the simulated cloud map in the embodiment. Rows: Three different ground object scenes respectively; Columns: SAR image, RGB image to be repaired, clean RGB image, and cloud-free image obtained by cloud removal and repair by our model respectively.

[0052] Figure 7 It is a schematic diagram of the cloud removal result of the model in the real cloud map in the embodiment. Rows: Three different locations in the same dataset; Columns: SAR image, RGB image to be repaired, binary cloud mask predicted by the cloud detection algorithm, clean RGB image, and cloud-free image obtained by cloud removal and repair by our model respectively.

[0053] The following further describes the implementation steps of the present invention in conjunction with the attached drawings: Specific implementation manners

[0054] Refer to Figure 1 , 2, 3, the implementation steps of the present invention are as follows:

[0055] 1. Synthesize the image to be repaired.

[0056] Draw a batch of binary mask images with random shapes. The pixel values in the images are specified as 0 or 255, where the area with pixel value 255 represents the part of the RGB image blocked by thick clouds, and the area with pixel value 0 represents the valid part of the RGB image not blocked. Use the alpha channel blending method to synthesize the binary mask M with a clean RGB image to generate a simulated RGB image X blocked by clouds. R . At the same time, select some real cloud images from the 38-Cloud dataset, and use the corresponding binary mask label M r After filtering the ground background information, then generate a real RGB image X blocked by clouds through the alpha channel blending method with the clean RGB image. Rr , and the pixel values in this mask are also 0 or 255, with the same meaning as above. As Figure 5 shown.

[0057] 2. Implement a multi-modal image cloud removal model based on feature decoupling

[0058] This algorithm consists of an encoder and a decoder. The encoder is respectively composed of a content feature extractor in the SAR modality a content feature extractor in the RGB modality and a style feature extractor in the RGB modality Composition. The content feature extractor consists of two dilated convolutions with a stride of 2 and four residual blocks built with dilated convolutions. Each layer of convolution in the content extractor is paired with an Instance Normalization layer. The style feature extractor consists of two dilated convolutions with a stride of 2, a global pooling layer, and a 1×1 convolution. Strided convolution is also used for downsampling. The decoder consists of a multi-layer perceptron, a group of residual blocks configured with AdaIN normalization layers, and a group of upsampling modules composed of alternating nearest neighbor upsampling layers and ordinary convolutions.

[0059] 3. Implement a self-attention-based modality difference supplementation mechanism

[0060] First, calculate the difference information C between the feature maps in the RGB and SAR modalities DF , and then use a gating function composed of a Sigmoid function and a 1×1 convolution to adaptively select the information in C DF to obtain the supplementary information C PC . Then, use the self-attention mechanism to model the relationship between each pixel in C S , and based on this relationship, convert the known supplementary information C PC within the valid region into the required supplementary information in the region to be repaired, and transfer the supplementary information required for each pixel in this region to the corresponding position. Finally, add the global supplementary information to C S to complete the supplementation of modality differences.

[0061] 4. Construct the final model and train it

[0062] Embed the self-attention-based modality difference supplementation mechanism into the multi-modal image cloud removal algorithm based on feature decoupling to obtain the final cloud removal network, and complete the training under the guidance of the objective function L(G,D). During training, the stochastic gradient descent method is used, the optimizer is Adam, its momentum sparsity β1 and β2 are set to 0.5 and 0.999 respectively, the weight decay factor is set to 0.0001, and it is trained for 30 rounds while keeping the learning rate at 0.0001.

[0063] 5. Application in actual tasks

[0064] As Figure 4 shown, the data flow involved in the inference phase of the model is different from that in the training phase. First, the network respectively uses the content feature extractors in the SAR and RGB modalities to extract the content features C S in the SAR image X R and the content features C S and C RExtract it, and at the same time use the style feature extractor in the RGB modality to extract the style feature s of the RGB image R 。Utilize the modality difference supplementation mechanism based on self-attention to further reduce the differences between multi-modal content features to obtain C SR ,and then couple it with s R to generate X SR 。At the same time, couple C R with s R to reconstruct the clean RGB image X R′ 。Finally, splice it according to M to obtain the final cloud-removing and restoration result X R″ 。We select three datasets containing different ground object scenes for cloud-removing and restoration experiments. The resolution of each image is 256×256. The three datasets include 100, 70, and 103 test images respectively, and cover various ground object types such as farmland, towns, mountains, and rivers. On the NVIDIA Geforce GTX 3090 GPU with 24G of video memory, a simulation experiment is carried out with the help of the Pytorch deep learning framework, and the final test results are as Figure 6 shown.

[0065] It can be seen that our method can well repair the ground information blocked by clouds, showing strong consistency with the target image in terms of structure, texture, and color, and the ground scene looks realistic enough, indicating that our method has good practicability. Moreover, our model only contains 20MB of learnable parameters, which is much smaller than other methods in this field, so the requirements for the operating environment are relatively low.

[0066] In addition, to verify that our method has strong robustness to changes in the shape and size of clouds, we input the SAR image X S and the RGB image X Rr blocked by real clouds into the model, and directly use the model trained in the above first dataset for experiments without fine-tuning. The experimental results are as Figure 7 shown.

[0067] The experimental results show that even if the characteristics such as the shape and size of the current clouds have never appeared during model training, our model can still remove them well, and the cloud-removing and restoration results maintain strong consistency with the target image in terms of structure, texture, and color, and the model performance does not show an obvious decline. This fully demonstrates that our model can be quickly put into practical application without fine-tuning in extreme meteorological observation tasks and has high practical value.

Claims

1. A remote sensing image cloud removal method based on the fusion of synthetic aperture radar and visible light, comprising the following steps: Step 1: Draw a binary mask to simulate cloud occlusion. The pixels in the mask are set to 0 or 255. 0 indicates that the current pixel is occluded by clouds, and 255 indicates that the current pixel is valid information. Then, the mask is combined with the clean RGB image through the alpha channel blending method to synthesize a simulated RGB image occluded by clouds ; At the same time, select some real cloud images from the 38-Cloud dataset, and use the corresponding binary mask labels to filter out the ground background information, and then generate a real image occluded by clouds by blending with the clean RGB image through the alpha channel blending method , and the pixel values in this mask are also 0 or 255; Step 2: Design a multi-modal image cloud removal algorithm based on feature decoupling, which can decouple the same content information expressed by the paired SAR image and RGB image from the features of the SAR modality and RGB modality respectively, so as to reduce the differences between images of different modalities and improve the efficiency of multi-modal image fusion; This algorithm consists of an encoder and a decoder: The encoder is respectively composed of a content feature extractor in the SAR mode , a content feature extractor in the RGB mode and a style feature extractor in the RGB mode ; Content Feature Extractor It consists of two dilated convolutions with a stride of 2 and four residual blocks built by dilated convolutions. The former is used for downsampling the feature map and extracting shallow features in the image; the latter is used for extracting deep semantic features. Each layer of convolution in the content extractor is paired with an Instance Normalization layer to remove the style information in the feature map. Style Feature Extractor It consists of two dilated convolutions with a stride of 2, a global pooling layer, and a convolution; the dilated convolution with a stride of 2 is used for downsampling; the global average pooling layer compresses the feature map into a feature vector that can represent the style of each channel to remove the spatial information in the image; finally, the convolution condenses and purifies the above feature vector to obtain the final style feature vector; The decoder first uses a Multilayer Perceptron (MLP) to convert the style features of the RGB image into the parameters required for the corresponding AdaIN layer, and then passes the parameters to a group of residual blocks configured with AdaIN normalization layers to perform style rendering of the content features in the RGB modality. Finally, the feature map is converted into a cloud-free RGB image by an upsampling module alternately composed of a nearest neighbor upsampling layer and a common convolution; Step 3: Design a self-attention based modal difference compensation mechanism, which can make full use of the effective content information in the RGB modality to correct the content information in the SAR modality pixel by pixel, improve the consistency of the content information extracted from the multi-modal image, and ultimately enhance the consistency inside and outside the repaired area; This step includes the following sub-steps: Sub-step 301: Calculate the difference information between the feature maps in RGB and SAR modalities , that is , where is the content feature extracted from the SAR image, is the content feature of the RGB image, is a binary mask predicted by dilated convolution to describe the invalid region in the current feature map, is the Hadamard product of matrices; Sub-step 302: Use a gating function composed of function and convolution to adaptively select the information in to obtain supplementary information , which is expressed by the formula ; Sub-step 303: Use the self-attention mechanism to model the relationships between each pixel in and convert the known supplementary information in the valid region into the required supplementary information in the region to be repaired, that is , where , , the subscript represents the pixel position of the current query, and the subscript will traverse all pixels at all positions in the attention map; , and are both linear transformation matrices, is the activation function, represents layer normalization; Step 4: Embed the self-attention based modal difference compensation mechanism into the multi-modal image cloud removal algorithm based on feature decoupling, and train this algorithm to finally obtain a multi-modal image cloud removal model based on feature decoupling; The objective function of the multi-modal image cloud removal algorithm based on feature decoupling includes two image reconstruction losses, three feature reconstruction losses, a modal invariance loss and two adversarial losses. The following will introduce each loss term that makes up the objective function in detail: (1) (2) Here is the modal conversion image , and the RGB reconstruction image , where represents the cosine distance, represents the self-attention based modal difference supplement mechanism; is the RGB decoder; (3) (4) (5) Here The content features are supplemented with information via to obtain the content features , as well as the style features . denotes the th layer of the VGG-19 network, and the corresponding weights are 0.125, 0.25, and 1 respectively; (6) Here ; (7) (8) Here is a discriminator for discriminating the authenticity of RGB images; The finally obtained loss function is composed of the above eight loss terms and is summarized into the following formula: (9) Among them, we stipulate that , , and ; , , and are hyperparameters used to balance each loss term; Step 5: Apply the multi-modal image cloud removal model based on feature decoupling obtained after training to the actual extreme meteorological scenario to complete the remote sensing image cloud removal task based on the fusion of synthetic aperture radar and visible light.

Citation Information

Patent Citations

  • De-clouding method for optical and SAR image fusion based on deep dense residual network

    CN114549385A

  • Multi-temporal remote sensing image cloud region reconstruction method

    CN115293986A