Remote sensing image cloud removal method and system based on pseudo-twin feature extraction and multi-modal fusion

By employing pseudo-twin feature extraction and multimodal fusion, and utilizing generative adversarial networks to process SAR and multi-cloud optical images, the problems of noise removal and low fusion efficiency in optical remote sensing images are solved, generating high-quality cloudless images and enabling stable monitoring of giant panda habitats.

CN122265102APending Publication Date: 2026-06-23SICHUAN RES INST OF GIANT PANDA SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN RES INST OF GIANT PANDA SCI
Filing Date
2026-03-23
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

In existing methods for removing clouds from optical remote sensing images, the salt-and-pepper speckle noise of SAR images and the differences in imaging mechanisms between optical and SAR images lead to low reconstruction accuracy. Furthermore, insufficient fusion of dual-modal data makes it difficult to effectively remove cloud interference, which affects the monitoring and analysis of key areas such as giant panda habitats.

Method used

A method based on pseudo-twin feature extraction and multimodal fusion is adopted. SAR images and cloudy optical images are processed by generative adversarial network (GAN). The pseudo-twin feature extraction branch is used to suppress speckle noise and cloud interference. Combined with multi-scale progressive decoder and multi-scale PatchGAN discriminator, the generator and discriminator are optimized to generate high-quality cloudless optical images.

Benefits of technology

The system generates high-quality cloud-free optical images, solving the problems of insufficient cloud removal data for habitats and noise interference in SAR images. It enables continuous monitoring of elements such as bamboo forests and water systems, providing stable and reliable data support for habitat protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265102A_ABST
    Figure CN122265102A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image restoration, in particular to a remote sensing image cloud removal method and system based on pseudo twin feature extraction and multi-modal fusion. The method comprises the following steps: acquiring a SAR image and a multi-cloud optical image, processing the acquired SAR image and multi-cloud optical image through a generative adversarial network (GAN) generator to obtain an adversarial generated optical image; processing the adversarial generated optical image through an adaptive multi-modal fusion strategy to obtain image fusion features; processing the image fusion features through a multi-scale progressive decoder to output a cloud-removed reconstructed optical image; inputting the cloud-removed reconstructed optical image and a real cloud-free optical image into a multi-scale PatchGAN discriminator after splicing the cloud-removed reconstructed optical image and the real cloud-free optical image with the SAR image respectively to obtain multi-scale discrimination results; and jointly optimizing the generator and the discriminator through a WGAN-GP loss until the model converges to obtain an optimal cloud removal model. The application provides a more stable and reliable data source and technical tool for animal habitat protection, ecological evaluation and management decision of high cloud coverage areas such as giant panda habitats.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image restoration technology, and in particular to a method and system for cloud removal from remote sensing images based on pseudo-twin feature extraction and multimodal fusion. Background Technology

[0002] Satellite remote sensing technology, with its advantages of wide detection range, high acquisition efficiency, and diverse data, plays a crucial role in Earth observation missions and is widely used in land use classification, environmental monitoring, disaster assessment, and other fields. Optical remote sensing, as a core means of modern Earth observation, provides convenient conditions for observing the Earth and offers indispensable data support for numerous scientific fields due to its high-resolution imaging and WYSIWYG capabilities. Giant pandas, a rare species unique to China, are distributed across the Qinling, Minshan, Qionglai, Xiangling, and Liangshan mountain ranges. Their habitat is vast and the terrain complex, with manual patrols only covering about 10% of the area. Optical satellite remote sensing can provide large-scale, periodic observations, effectively compensating for this deficiency. However, optical satellites primarily rely on receiving solar radiation reflected from the Earth's surface for imaging, a process highly susceptible to weather conditions. Cloudy weather can cause imaging signals to attenuate or disappear, making it difficult to obtain effective surface information. According to statistics, the global annual cloud cover exceeds 60%, while the main mountain ranges where giant pandas are distributed in China have a humid climate and are more frequently affected by clouds, fog, and haze. This severely reduces the availability of optical imagery data, making it difficult to identify key habitat elements such as bamboo forest patches and water systems. In addition, the data loss caused by cloud pollution will also hinder the monitoring of continuous dynamic changes in habitats.

[0003] To overcome cloud interference, many methods for declouding optical remote sensing images have emerged. Their core approaches mainly fall into two categories: one is to repair cloud areas based on the statistical characteristics and texture information of a single image; the other is to introduce multi-temporal or heterogeneous data for replacement or reconstruction. Among numerous heterogeneous data sources, Synthetic Aperture Radar (SAR), with its all-weather imaging capability and immunity to cloud interference, has become one of the most important auxiliary data sources. Existing SAR-assisted optical image declouding methods are mostly based on deep learning algorithms, and the research mainly faces two major problems: first, the inherent salt-and-pepper speckle noise in SAR images, if directly input into the network without processing, will be misjudged as real surface texture, reducing reconstruction accuracy; second, the interaction method between SAR and optical images is problematic. SAR and optical images have fundamental differences in their imaging mechanisms, and simple cascading or addition fusion methods cannot fully achieve information complementarity, resulting in low data utilization and poor image cloud area reconstruction. Therefore, this paper proposes an optical image declouding method based on the fusion of generative adversarial networks (GAN) and SAR, aiming to fully leverage the advantages of SAR data to effectively remove cloud interference and improve the availability and monitoring continuity of optical images in giant panda habitat areas. Summary of the Invention

[0004] This application provides a method and system for cloud removal from remote sensing images based on pseudo-twin feature extraction and multimodal fusion to solve the above-mentioned problems.

[0005] Firstly, this application provides a remote sensing image declouding method based on pseudo-twin feature extraction and multimodal fusion. The method includes: acquiring SAR images and cloudy optical images; inputting the SAR images and cloudy optical images into a generator of a declouding network; processing the input data in the generator through a pseudo-twin feature extraction branch, an adaptive multimodal fusion module, and a multi-scale progressive decoder to output a declouded reconstructed optical image; stitching the declouded reconstructed optical image and a truly cloudless optical image with the SAR image respectively, and then inputting the result into a multi-scale PatchGAN discriminator to obtain the discriminator output; and jointly optimizing the generator and discriminator using WGAN-GP loss until convergence to obtain the declouding model.

[0006] The above technical solutions can effectively overcome the constraints of cloudy weather on optical satellite image observation, generate high-quality cloudless optical images, and effectively solve key problems such as insufficient cloud removal data in giant panda habitat areas, SAR image speckle noise interference, and inefficient dual-modal data fusion. This makes it possible to continuously monitor and analyze key elements such as bamboo forests and water systems in giant panda habitats, and provides a more stable and reliable data source and technical tools for habitat protection, ecological assessment and management decisions.

[0007] Optionally, the generator's processing includes: based on the SAR image, processing is performed through the SAR feature extraction sub-branch in the pseudo-twin feature extraction branch to suppress speckle noise and extract structural texture features to obtain SAR feature representation; based on the multi-cloud optical image, processing is performed through the optical feature extraction sub-branch in the pseudo-twin feature extraction branch to adaptively suppress cloud interference and retain effective spectral spatial features to obtain optical feature representation; wherein, the adversarial network GAN generator is jointly trained using a loss function that includes multi-scale spatial frequency L1 loss and adversarial loss based on Wasserstein distance.

[0008] Optionally, the construction process of the pseudo-twin feature extraction branch includes: the front end of the SAR feature extraction sub-branch is equipped with a learnable frequency domain denoising head, which is used to suppress speckle noise of SAR images in the frequency domain and maintain high-frequency structural information; the optical feature extraction sub-branch uses a gated convolution strategy for feature extraction, which is used to adaptively suppress cloud area features and retain cloudless area features; the SAR feature extraction sub-branch and the optical feature extraction sub-branch are integrated to construct the pseudo-twin feature extraction branch.

[0009] Optionally, the specific implementation of the frequency domain denoising head includes: converting the input SAR image to the complex frequency domain using a Fast Fourier Transform (FFT) to obtain an amplitude spectrum and a phase spectrum; generating a soft mask from the amplitude spectrum using a learnable convolutional layer to modulate the amplitude spectrum to suppress noise, obtaining a modulated amplitude spectrum; recombining the modulated amplitude spectrum and the phase spectrum, and converting them back to the spatial domain using an Inverse Fast Fourier Transform (IFFT) to obtain preliminary denoising features.

[0010] Optionally, the specific implementation of the gated convolution strategy includes: performing gated convolution processing on the input multi-cloud optical image, and splitting the output features along the channel dimension into a multi-cloud optical image feature branch and a multi-cloud optical image gated branch; the multi-cloud optical image gated branch generates spatially adaptive gate weights through processing including the Sigmoid function; and multiplying the gate weights with the output of the multi-cloud optical image feature branch pixel by pixel to suppress cloud features and enhance cloudless features.

[0011] Optionally, the image fusion feature construction process includes: taking the SAR feature representation and the optical feature representation as input; processing them using a fusion strategy consisting of a bidirectional cross-transformer block, wherein each bidirectional cross-transformer block contains two symmetrical cross-attention strategies and a global self-attention strategy; and through the bidirectional cross-attention mechanism, enabling the SAR feature representation and the optical feature representation to perform bidirectional selection and complementary enhancement at the semantic level, thereby generating image fusion features that are reliable in cloud areas, rich in texture, and spectrally consistent.

[0012] Optionally, the specific implementation of the bidirectional cross-attention mechanism includes: in the first cross-attention stage, using the optical feature expression as the query and the SAR feature expression as the key and value, calculating the attention weight, and extracting the backscattering texture features of the corresponding spatial location from the SAR feature expression to supplement and reconstruct the missing structural information of the cloud area in the multi-cloud optical image; in the second cross-attention stage, using the SAR feature expression as the query and the optical feature expression as the key and value, calculating the attention weight, and extracting the spectral features of the corresponding location from the optical feature expression to realize the mapping from SAR structural information to spectral information; weighted summing of the output results of the two cross-attention stages, and inputting the summed result into the global self-attention strategy for information rearrangement and alignment to suppress single-path artifacts and generate the image fusion features with consistent structure and spectrum.

[0013] Optionally, the process of constructing the declouded reconstructed optical image includes: inputting the image fusion features into the multi-scale progressive decoder; performing progressive upsampling and feature reconstruction through the multi-scale progressive decoder, and simultaneously outputting reconstruction results at different scales; outputting the reconstruction results as the final declouded reconstructed optical image, and using the reconstruction results at different scales for multi-level supervised training.

[0014] Optionally, the specific implementation of the multi-scale progressive decoder includes: inputting the image fusion features into the multi-scale progressive decoder, and processing them through a progressive structure composed of multiple upsampling blocks and spatial-frequency residual blocks; wherein, the upsampling blocks utilize nearest-neighbor interpolation to enlarge the spatial resolution of the feature map, and then use 3×3 convolutions for channel dimensionality reduction and local spatial reconstruction; the spatial-frequency residual blocks include spatial and frequency branches: the spatial branch extracts local features through cascaded 3×3 convolutions to obtain local detail information; the frequency branch transforms the input features to the frequency domain via fast Fourier transform, performs cross-channel mixing and correction of frequency components through learnable 1×1 convolutions, and then transforms them back to the spatial domain via inverse fast Fourier transform to obtain global structural complementary information; finally, the local detail information and the global structural complementary information are fused.

[0015] Optionally, the specific implementation of the multi-scale discriminator includes: stitching the de-clouded reconstructed optical image or the real cloudless optical image with the corresponding SAR image along the channel dimension to obtain multi-channel conditional input features; inputting the multi-channel conditional input features and their 1 / 2 downsampled and 1 / 4 downsampled versions into the original scale discriminant branch, 1 / 2 scale discriminant branch, and 1 / 4 scale discriminant branch of the multi-scale PatchGAN discriminator, respectively; and performing progressive downsampling and feature extraction through multiple convolutional layers in each of the discriminant branches, wherein each convolutional layer uses a 4×4 convolutional kernel with a stride of 2, and is combined with the LeakyReLU activation function to capture local texture differences at the corresponding scale. The final output includes the original scale confidence map, the 1 / 2 scale confidence map, and the 1 / 4 scale confidence map. Each pixel value in each confidence map corresponds to the probability of authenticity of the local receptive field region of the input image at the corresponding scale. The WGAN-GP loss is calculated based on each confidence map, and the generator and the discriminator are optimized through adversarial training. In the multi-scale PatchGAN discriminator, all BatchNorm layers are removed from each discriminant branch structure to avoid the dynamic scaling and offset of batch statistics interfering with the Lipschitz continuity condition required by the Wasserstein distance constraint. The feature extraction path is constructed only through convolution operations and activation functions.

[0016] Secondly, this application provides a remote sensing image cloud removal system based on pseudo-twin feature extraction and multimodal fusion, the system comprising: The pseudo-twin feature extraction module is used to acquire SAR images and cloudy optical images. The SAR images and cloudy optical images are input into the adversarial network GAN generator and processed by the adversarial network GAN generator to obtain adversarial generated optical images. The cross-fusion module is used to process adversarial generative optical images through an adaptive multimodal fusion strategy. It utilizes a bidirectional cross-attention mechanism to achieve deep complementarity and enhancement of bimodal features, thereby obtaining image fusion features. The cloud removal and reconstruction module is used to process image fusion features through a multi-scale progressive decoder, perform progressive upsampling and reconstruction, and output cloud removal and reconstruction optical images. The adversarial learning module is used to stitch the de-clouded reconstructed optical image and the real cloudless optical image with the SAR image and input them into the multi-scale PatchGAN discriminator to obtain the multi-scale discrimination result. Based on the multi-scale discrimination result, the generator and discriminator are jointly optimized through WGAN-GP loss until the model converges and the optimal de-clouding model is obtained. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram illustrating an application scenario provided in one embodiment of this application; Figure 2 A flowchart of a remote sensing image declouding method based on pseudo-twin feature extraction and multimodal fusion provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a remote sensing image declouding system based on pseudo-twin feature extraction and multimodal fusion, provided as an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0020] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0021] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.

[0022] In the field of optical remote sensing image declouding, existing optical remote sensing declouding techniques utilize SAR images for assisted reconstruction, but there are three main problems: First, SAR speckle noise contaminates feature extraction, introducing artifacts; second, there are large differences in imaging between images, and existing simple fusion strategies are difficult to achieve deep semantic complementarity, resulting in insufficient data utilization; third, there is a lack of high-quality paired datasets for specific scenarios (such as giant panda habitats), which restricts the practicality and accuracy of the model in complex terrain.

[0023] Based on this, this application provides a remote sensing image declouding method and system based on pseudo-twin feature extraction and multimodal fusion, which effectively overcomes the limitations of cloudy weather on optical satellite observation, generates high-quality cloudless images, solves problems such as insufficient habitat declouding data, SAR image noise and inefficient dual-modal fusion, and realizes continuous monitoring of elements such as bamboo forests and water systems, providing stable and reliable data and technical support for habitat protection and management decisions.

[0024] Figure 1 This application provides an illustration of an application scenario. In the process of feature extraction for cloud removal from remote sensing images, the method provided in this application is applied to overcome the constraints of cloudy weather, generate high-quality cloudless optical images, and overcome key bottlenecks such as data loss, noise interference, and inefficient fusion. It enables stable and continuous monitoring of elements such as bamboo forests and water systems within habitats, providing reliable support for ecological protection and management decisions.

[0025] Specifically, the method provided in this application can be applied to any server, which interacts with satellite remote sensing observation platforms and optical satellites to acquire SAR images provided by satellite remote sensing observation platforms and multi-cloud optical images provided by optical satellites. This effectively solves key problems such as insufficient declouding data in giant panda habitat areas, speckle noise interference in SAR images, and inefficient dual-modal data fusion. It also outputs declouded reconstructed optical images to image observers, providing a more stable and reliable data source and technical tool for habitat protection, ecological assessment, and management decisions.

[0026] For specific implementation details, please refer to the following examples.

[0027] Figure 2This is a flowchart illustrating a remote sensing image cloud removal method based on pseudo-twin feature extraction and multimodal fusion, provided in one embodiment of this application. The method of this embodiment can be applied to servers in the above scenarios. Figure 2 As shown, the method includes: S201 is used to acquire SAR images and cloudy optical images. The SAR images and cloudy optical images are input into the adversarial network GAN generator, which processes them to obtain adversarial generated optical images.

[0028] SAR imagery can be remote sensing images acquired by synthetic aperture radar by emitting microwaves and receiving signals reflected from ground objects, with satellite remote sensing observation platforms as the data source. Cloudy optical imagery can be optical remote sensing images where ground object information is blurred or partially obscured due to cloud cover or fog, with optical satellites as the data source.

[0029] Specifically, in the field of remote sensing image processing, cloud cover is the main problem affecting the usability of optical images. Simply relying on the cloud-covered optical images themselves for declouding often results in poor reconstruction results when the cloud cover is thick or covers a large area, as the information below the cloud is completely lost. SAR images are not affected by weather and can provide stable surface observations, but their imaging mechanism is very different from that of optical images, and there is an information gap when using them directly. Existing methods mostly use simple addition or channel cascading, which makes it difficult to achieve deep, nonlinear information conversion and supplementation between two different modalities of data.

[0030] S202. Based on adversarial generative optical images, an adaptive multimodal fusion strategy is used to process them. A bidirectional cross-attention mechanism is used to achieve deep complementarity and enhancement of bimodal features, resulting in image fusion features.

[0031] Adaptive multimodal fusion strategies can be fusion methods that dynamically adjust fusion weights and adapt to the differences in features between different modalities. Bidirectional cross-attention mechanisms can simultaneously focus on the interrelationships between features from two different modalities, achieving bidirectional feature guidance and enhancement through attention computation. Image fusion features can be comprehensive feature representations that integrate the advantageous features of both modalities after processing by adaptive multimodal fusion strategies.

[0032] Specifically, although SAR images provide texture information of cloud areas, they may not be as good as the cloudless areas of the original optical images in terms of spectral fidelity and detail richness. The cloudless areas of the original cloudy optical images retain the most realistic ground reflection characteristics, but the cloud areas are missing. Existing fusion methods such as simple feature stitching or weighted averaging cannot deeply explore the complex relationship between these two modal data from different sources and with uneven quality, and cannot achieve "complementing each other's strengths".

[0033] S203 is used to process image fusion features through a multi-scale progressive decoder, perform progressive upsampling and reconstruction, and output cloud-reconstructed optical images.

[0034] A multi-scale progressive decoder can be a decoding network that includes multiple feature extraction modules at different scales and progressive upsampling modules. De-clouding reconstructed optical imagery can be an optical imagery that has undergone complete de-clouding processing to restore the information of ground features obscured by clouds and fog.

[0035] Specifically, the feature maps obtained through deep fusion have low resolution and contain highly abstract information. Directly upsampling them to the original image size and decoding them can easily lead to blurred details and structural distortions, which cannot meet the needs of high-precision applications. By adopting a multi-scale progressive decoding strategy, which mimics the visual reconstruction process from coarse to fine, it can more stably and accurately complete the restoration of large-scale cloud areas and the reconstruction of details, effectively avoiding artifacts and distortions that may be generated by direct single-step reconstruction.

[0036] S204 is used to stitch the de-clouded reconstructed optical image and the real cloudless optical image with the SAR image respectively and input them into the multi-scale PatchGAN discriminator to obtain the multi-scale discrimination result. Based on the multi-scale discrimination result, the generator and discriminator are jointly optimized through WGAN-GP loss until the model converges and the optimal de-clouding model is obtained.

[0037] Specifically, each branch of the PatchGAN discriminator has the same network structure, consisting of multiple convolutional layers connected in series. Features are extracted through progressive downsampling. Each convolutional layer uses a 4×4 kernel with a stride of 2 and is coupled with the LeakyReLU activation function to capture local texture differences at the corresponding scale, ultimately outputting a two-dimensional confidence map. The discriminator is optimized using the WGAN-GP loss function, which is based on the Wasserstein distance and ensures the Lipschitz continuity of the discriminator by adding a gradient penalty term. This loss function can effectively alleviate the gradient vanishing and mode collapse phenomena that are prone to occur in existing cross-entropy losses. The calculation formula is shown in Equation (1). In order to satisfy the Lipschitz continuity of the discriminator, all BatchNorm layers in the discriminator structure are removed, and the feature extraction path is constructed only through convolution operations and activation functions.

[0038]

[0039] in For the cloud removal results, For true cloudless images, , representing random linear interpolation on the real cloudless image and the declouded result.

[0040] The method provided in this embodiment can effectively overcome the constraints of cloudy weather on optical satellite image observation, generate high-quality cloudless optical images, and effectively solve key problems such as insufficient cloud removal data in giant panda habitat areas, SAR image speckle noise interference, and inefficient dual-modal data fusion. This makes it possible to continuously monitor and analyze key elements such as bamboo forests and water systems in giant panda habitats, and provides a more stable and reliable data source and technical tool for habitat protection, ecological assessment and management decisions.

[0041] In some embodiments, based on SAR imagery, the SAR feature extraction sub-branch in the pseudo-twin feature extraction branch is used to process the imagery, suppress speckle noise and extract structural texture features to obtain SAR feature representation; based on multi-cloud optical imagery, the optical feature extraction sub-branch in the pseudo-twin feature extraction branch is used to process the imagery, adaptively suppress cloud interference and retain effective spectral spatial features to obtain optical feature representation; wherein, the generator of the declouding network is obtained by joint training of a loss function including multi-scale spatial frequency L1 loss and adversarial loss based on Wasserstein distance.

[0042] The pseudo-twin feature extraction branch can be a branch structure specifically designed for the differential extraction of dual-modal features from SAR and multi-cloud optical images. The SAR feature extraction sub-branch can be one of the core sub-modules of the pseudo-twin feature extraction branch, used to achieve speckle noise suppression and structural texture feature extraction. Speckle noise can be the inherent speckle interference in SAR images, generated by the random interference between radar waves and ground targets, manifesting as randomly distributed high-grayscale noise points in the image. Structural texture features can be the key information in SAR images that characterizes the spatial structure and surface texture of ground objects. SAR feature representation can be the feature vector or feature map obtained after processing the SAR image through the SAR feature extraction sub-branch, capable of abstractly representing the effective structural texture information in the SAR image. The optical feature extraction sub-branch can be another core sub-module of the pseudo-twin feature extraction branch, used to adaptively suppress cloud interference and retain effective spectral spatial features. Cloud interference can be the obstruction and interference of effective features by cloud-covered areas in multi-cloud optical images, manifested as the high brightness and low contrast characteristics of cloud areas. Effective spectral spatial features can be spectral and spatial features reflecting the inherent properties of ground objects. Optical feature representation can accurately characterize the effective spectral spatial features in optical images. Multi-scale spatial-frequency L1 loss can be a loss function that combines multi-scale features with spatial-frequency information. Adversarial loss based on Wasserstein distance can be a core criterion for measuring the difference between the model's generated results and the real target. Joint training can simultaneously train the model using multiple loss functions, integrating the constraints of different losses through weighted fusion and other methods.

[0043] Specifically, in the feature extraction process of cloud removal from remote sensing images, existing methods do not specifically address speckle noise in SAR images and cloud interference in optical images, resulting in noise and spectral distortion in the extracted features. Furthermore, training with a single loss function is prone to causing mode collapse, which leads to blurred ground feature boundaries and loss of details in subsequent fusion and cloud removal results. To address the above issues, a pseudo-twin feature extraction branch was constructed, designed based on the different characteristics of SAR and multi-cloud optics. The purpose of the SAR branch is to extract effective high-frequency information such as texture, edge, and structure under the premise of limiting its noise, so as to provide high signal-to-noise ratio, multi-scale, and structure-preserving feature representation for subsequent cross-modal fusion. The speckle noise of SAR images comes from the coherent superposition of scattered waves during the imaging process. Its intensity is proportional to the signal amplitude and is multiplicative noise. Existing SAR denoising methods such as Lee filtering and GAMMA filtering rely on image modeling, high-frequency details are easily blurred, and parameters need to be manually adjusted, making them difficult to use for deep learning. The loss function includes two parts: multi-scale spatial-frequency L1 loss and WGAN-GP loss. The multi-scale spatial-frequency L1 loss is shown in Equation (1). This loss is achieved by calculating the spatial-frequency domain L1 loss on the three scales (original image, 1 / 2 scale, and 1 / 4 scale) of the multi-scale decoder, realizing the joint optimization of pixel accuracy and frequency domain distribution. Multi-scale supervision makes the model gradient more stable. Frequency domain L1 can effectively solve the problem of overly smooth reconstruction results and missing high-frequency information caused by pixel loss in image reconstruction tasks.

[0044]

[0045] in For the cloud removal results, For reference image, For the first Normalization factor at scale This represents the Fast Fourier Transform (FFT) that transforms the image signal to the frequency domain. Based on experience... The value is set to 0.1 to balance the dual-domain learning, and the adversarial loss based on Wasserstein distance is shown in Equation (4).

[0046] in For the cloud removal results, This is a true image without clouds.

[0047] The method provided in this embodiment accurately suppresses speckle noise in SAR images and cloud interference in optical images, while preserving complete structural texture and effective spectral features. The joint loss function ensures the stability of model training and the quality of generated images, laying a high-quality foundation for subsequent multimodal fusion. This results in clear structure and realistic spectrum in the declouded images, significantly improving the reliability of applications such as ground feature identification and topographic mapping, and enhancing the practical value of the declouding method.

[0048] In some embodiments, the front end of the SAR feature extraction sub-branch is equipped with a learnable frequency domain denoising head to suppress speckle noise in the SAR image in the frequency domain and maintain high-frequency structural information; the optical feature extraction sub-branch uses a gated convolution strategy for feature extraction to adaptively suppress cloud area features and retain cloudless area features; the SAR feature extraction sub-branch and the optical feature extraction sub-branch are integrated to construct a pseudo-twin feature extraction branch.

[0049] Frequency-domain denoising heads can suppress noise in SAR images in the frequency domain. Speckle noise, a type of noise unique to SAR images caused by microwave scattering characteristics, interferes with structural feature extraction and needs to be suppressed by frequency-domain denoising heads. High-frequency structural information refers to detailed and rapidly changing structural data in the image. Gated convolution strategies can be convolutional processing methods with gated mechanisms, using spatially adaptive gated weights to distinguish between cloud-covered and cloudless areas, thereby enhancing the targeting of optical feature extraction.

[0050] Specifically, in the feature extraction process for cloud removal from remote sensing images, if a pseudo-twin feature extraction branch is not constructed, speckle noise from SAR images will interfere with structure and texture extraction, and cloud areas in multi-cloud optical images will mix with effective spectral features, resulting in low purity and poor targeting of bimodal features. To address these issues: First, for the SAR feature extraction sub-branch, the input SAR image is processed through a learnable frequency domain denoising head. This process first maps the spatial domain image to the frequency domain via a fast Fourier transform, decomposing it into amplitude and phase spectra. Then, a trainable convolutional layer is used to analyze the cascaded spectral information and generate a soft mask. This mask is used to adaptively modulate the amplitude spectrum of a specific frequency region representing noise energy to suppress noise, while keeping the phase spectrum unchanged to maintain structural location information. Finally, the spatial domain is reconstructed through an inverse transform, yielding a preliminary denoised and structure-preserving SAR feature representation. Meanwhile, for the optical feature extraction sub-branch, the input multi-cloud optical image is processed by a gated convolution strategy. This strategy splits the output features into a main branch representing the original feature response and a gated branch used to generate spatial selection weights through specific convolution operations. The gated branch converts the features into a spatially adaptive weight map with values ​​between 0 and 1 through a processing flow including a sigmoid function. This weight map can automatically identify cloud areas (assigned lower weights) and cloudless areas (assigned higher weights). Then, this weight map is multiplied pixel by pixel with the feature map of the main branch, thereby suppressing cloud area features and enhancing effective cloudless area features, resulting in a purified optical feature expression. Finally, the outputs of these two separately optimized feature extraction sub-branches are integrated to construct the pseudo-twin feature extraction branch, providing high-quality bimodal feature input for the subsequent fusion stage of the network.

[0051] The method provided in this embodiment uses the frequency domain denoising head of the SAR sub-branch to suppress speckle noise and preserve high-frequency structure, and the gated convolution strategy of the optical sub-branch to remove cloud interference and enhance cloudless features. This efficiently extracts high-purity dual-modal features, providing high-quality input for subsequent bidirectional cross-attention fusion, reducing fusion artifacts and cloud residues, ensuring structure-spectrum consistency, and improving the reconstruction accuracy of the multi-scale decoder.

[0052] In some embodiments, the input SAR image is transformed to the complex frequency domain by Fast Fourier Transform (FFT) to obtain the amplitude spectrum and phase spectrum; the information of the amplitude spectrum and phase spectrum is concatenated, and a soft mask is generated by a learnable convolutional layer to modulate the amplitude spectrum to suppress noise, resulting in a modulated amplitude spectrum; the modulated amplitude spectrum and phase spectrum are recombined and transformed back to the spatial domain by Inverse Fast Fourier Transform (IFFT) to obtain preliminary denoising features.

[0053] The Fast Fourier Transform (FFT) is a mathematical transformation method that converts spatial domain images to the frequency domain. The complex frequency domain can be a frequency space composed of amplitude and phase spectra, representing the result of the FFT transformation. The amplitude spectrum can be the component representing signal intensity in the complex frequency domain. The phase spectrum can be the component representing signal phase in the complex frequency domain. A learnable convolutional layer can be a convolutional neural network layer whose parameters can be optimized through training. A soft mask can be a flexible weight matrix generated by the convolutional layer. The modulated amplitude spectrum can be the amplitude spectrum adjusted by the soft mask. The Inverse Fast Fourier Transform (IFFT) is a mathematical transformation method that converts frequency domain signals back to the spatial domain. Preliminary denoising features can be the SAR image features obtained after processing by a frequency domain denoising head.

[0054] Specifically, in the process of multimodal feature fusion for cloud removal from remote sensing images, speckle noise in SAR images can severely interfere with the extraction of structural texture features. If used directly for subsequent fusion, it will cause the fused features to carry noise artifacts, destroying the accuracy of cloud area structure reconstruction and spectral consistency. It will also affect the feature complementarity effect of the bidirectional cross-attention mechanism, ultimately resulting in a significant decrease in the quality of the cloud-removed images. To address the above problems, the purpose of the frequency domain denoising head is to remove speckle noise in SAR images while retaining as much high-frequency information as possible. It mainly consists of two parts: frequency domain denoising and detail enhancement. In frequency domain denoising, the input feature Fsar[8,2,128,128] is subjected to Fast Fourier Transform (FFT) to transform the feature from the spatial domain to the frequency domain Ffft[8,2,128,65]. Ffft is decomposed into amplitude spectrum (mag[8,2,128,65]) and phase spectrum (phi[8,2,128,65]). High-frequency speckle noise is suppressed through a learnable convolutional gating network (composed of convolution and sigmoid functions). The denoised amplitude spectrum (mag') and phase spectrum (phi) are recombined into a complex frequency domain representation (the combination form is...). The frequency domain denoising result [8,2,128,128] is obtained by restoring the frequency domain denoising result to the spatial domain through inverse Fourier transform (iFFT). In the spatial domain detail enhancement, the frequency domain denoising result is subtracted from the original input to obtain a residual containing high-frequency details. This residual is then further extracted and enhanced by a detail enhancement convolutional network (composed of convolution + normalization + ReLU). Finally, the enhanced details are added to the frequency domain denoising result to output the final denoised feature F. denoise 8,2,128,128.

[0055] The method provided in this embodiment accurately suppresses speckle noise in SAR images, fully preserves high-frequency structural information, provides high-quality data for SAR feature extraction, adapts to different noise distributions in different scenarios, improves the denoising versatility and robustness, ensures the effectiveness of subsequent dual-modal feature complementarity, and significantly optimizes the structural integrity, texture clarity and spectral consistency of declouded images.

[0056] In some embodiments, the input cloudy optical image is subjected to gated convolution processing, and the output features are split along the channel dimension into a cloudy optical image feature branch and a cloudy optical image gated branch; the cloudy optical image gated branch generates spatially adaptive gate weights through processing including the Sigmoid function; the gate weights are multiplied pixel by pixel with the output of the cloudy optical image feature branch to suppress cloud features and enhance features in cloudless areas.

[0057] Gated convolution processing can split the output features into feature branches and gate branches to achieve adaptive suppression of cloud features. The feature branches of multi-cloud optical images can contain spectral and structural information of the optical image. The gate branches of multi-cloud optical images can be used to generate spatially adaptive gate weights. The Sigmoid function can be an activation function with an output value between 0 and 1, used to generate gate weights for the gate branches, achieving selective enhancement or suppression of features. The gate weights can be dynamically changing weight values ​​based on the spatial location of the image, used to distinguish and process features in cloud-covered and cloudless areas. Pixel-wise multiplication can be an operation that multiplies the gate weights and the output of the feature branches pixel by pixel, used to suppress cloud features and enhance cloudless features.

[0058] Specifically, in the process of declouding remote sensing images, the features of cloud areas and cloudless areas in cloudy optical images are mixed. Existing convolutional methods with fixed weights for feature extraction are prone to mistaking cloud interference as valid information or suppressing details in cloudless areas, leading to mismatches in subsequent dual-modal fusion structures and spectral inconsistencies, which seriously affect the quality of the declouded reconstructed image. To address the above problems: In cloudy optical images, the texture of cloud-covered areas is usually noise or interference information for the reconstruction process. If cloud areas and cloudless areas are processed indiscriminately, the network will treat all information in the image as valid signals. This will not only interfere with the subsequent feature fusion process, but also cause the presence of clouds or cloud shadows in the reconstructed image. Some studies use binarized cloud masks (0 for clouds, 1 for clear sky areas) as prior information input to the network to clarify the location of cloud areas. However, this method has the following two limitations: First, clouds exist in various forms in remote sensing images. Information in thick cloud areas is completely unusable, while thin cloud areas still retain some surface information, which the binarized cloud mask will completely discard, resulting in information waste. Second, the declouding effect of the network largely depends on the segmentation accuracy of the cloud mask.To address this, this study proposes a gated convolution module, which consists of two parts: feature extraction and dual-branch feature modulation. The first part, feature extraction, uses a 3×3 convolution to initially extract and expand the input features, and employs the SE channel attention mechanism to weight the results, enhancing feature representation. The second part is a dual-branch structure, splitting the obtained features along the channel dimension into a feature branch and a gated branch. The feature branch undergoes group normalization and GELU activation to stabilize values ​​and maintain gradient flow. The gated branch expands the receptive field through a 5×5 convolution and introduces two learnable parameters, γ and β, for further signal modulation. γ is used to scale the signal, preventing the gated signal from concentrating around 0.5 after Sigmoid activation; β is used for overall signal translation, making the gated signal more inclined towards 0 or 1. Finally, the gated signal is compressed to the [0,1] interval using the Sigmoid function, obtaining spatially adaptive gate weights. By multiplying the results of the feature branch pixel by pixel, this module can autonomously learn modulation strategies for input features during supervised training, effectively suppressing the influence of cloud features. This provides a cleaner and more reliable feature representation for subsequent feature fusion and image reconstruction. Taking the input feature shape (B,C,H,W) as (8,64,128,128) as an example, the input band of the instantiated gated convolution is 64, the output band is 128, the convolution kernel is 3, and the stride is 2. The feature shape transformation is as follows: First, it passes through the ReflectionPad layer for edge filling, and the feature shape is (8,64,130,130); after passing through the convolutional layer, the feature shape is (8,256,64,64); after passing through the SE module... The weighted shape is (8,256,64,64); the obtained features are split into Feature(8,128,64,64) and Gate(8,128,64,64) according to the feature dimension; the feature(8,128,64,64) retains its shape after normalization and activation function; the Gate(8,128,64,64) retains its shape after gating branch; the Feature and Gate retain their shape after pixel-by-pixel multiplication, and the final output shape of the module is (8,128,64,64).

[0059] The gated convolution strategy provided in this embodiment adaptively distinguishes between cloud-covered and cloudless areas, accurately suppresses cloud interference, enhances effective features in cloudless areas, provides high-quality optical features for subsequent bidirectional cross-attention fusion, avoids feature confusion and artifact generation, improves the robustness of the cloud removal method, and ensures the spectral consistency and structural integrity of the final image.

[0060] In some embodiments, SAR feature representation and optical feature representation are taken as input; a fusion strategy consisting of a bidirectional cross-transformer block is used for processing, which includes two symmetrical cross-attention strategies and a global self-attention strategy; through the bidirectional cross-attention mechanism, SAR feature representation and optical feature representation are bidirectionally selected and complementaryly enhanced at the semantic level to generate image fusion features that are reliable in cloud areas, rich in texture, and spectrally consistent.

[0061] A bidirectional cross-transformer block can be a stacked module (e.g., 8 layers) containing symmetric cross-attention strategies and global self-attention strategies. Cross-attention strategies can be methods that use one modality as a query and the other as a key and value to achieve accurate matching and information extraction of bimodal features. Global self-attention strategies can be attention mechanisms that capture global dependencies of features, correct local biases, and achieve feature alignment and artifact suppression. The bidirectional cross-attention mechanism can be the core mechanism for achieving semantic-level bidirectional selection and complementary enhancement of bimodal features through two symmetric cross-attention strategies. Semantic level can be a higher-order feature fusion level based on semantic information such as land cover type and structural attributes, distinct from pixel-level / shallow-level fusion. Bidirectional selection can be a way for bimodal features to mutually filter each other's effective information, achieving complementary effects between structural texture and spectral information. Complementary enhancement can be a process of mining the advantages of bimodal features, compensating for their respective limitations, strengthening effective features, and improving the quality of fused features.

[0062] Specifically, in the dual-modal feature fusion process for cloud removal from remote sensing images, existing methods only perform shallow overlay or unidirectional supplementation, leading to semantic mismatch, insufficient complementarity, distorted cloud area reconstruction structure, spectral inconsistencies, and a tendency to generate artifacts, resulting in decision-making errors. To address these issues, this method primarily achieves bidirectional complementary fusion of SAR and multi-cloud optical features, with the input being SAR modal features F. sar and multi-cloud optical modal characteristics F opt The two feature channels have already been aligned in terms of feature channels and resolution during the encoder stage, with F... sar [8,128,64,64] and F opt [8,128,64,64] is used as the model input, and it is split into 4×4 patchF by the PatchEmbed layer. sar [8,4096,128] and F opt 8,4096,128, Spatial cross-attention and channel attention are performed in parallel for each pair of token sequences. Spatial cross-attention is based on F... sar For Query, F optFor key / value pairs, multi-head cross-attention with a window size of 8 is used to complete multi-cloud optical textures using SAR, with an output shape of Fs-08,4096,128; channel cross-attention is based on F... opt For Query, F sar For key / value pairs, similarity is calculated along the channel dimension. Optical channel weighting of SAR channels is used to achieve spectral complementation and fusion. The output shape is... The above fusion results are added together and then fed into a standard window multi-head self-attention system for further refinement to eliminate local inconsistencies after dual-path merging. Finally, the fusion result is input into the PatchUnEmbed layer to restore F. fusion 8,128,64,64.

[0063] The method provided in this embodiment achieves semantic-level complementarity of dual-modal features through a bidirectional cross-attention mechanism, accurately compensating for spectral defects in SAR images and missing cloud structures in optical images, effectively suppressing fusion artifacts, and generating reliable, textured, and spectrally consistent fused feature cloud areas, providing high-quality input for subsequent cloud removal and reconstruction, and ensuring that the cloud removal images accurately reflect the true state of ground features.

[0064] In some embodiments, in the first cross-attention stage, the optical feature representation is used as the query and the SAR feature representation is used as the key and value to calculate the attention weight. Backscatter texture features corresponding to the spatial location are extracted from the SAR feature representation to supplement and reconstruct the missing structural information of the cloud area in the cloudy optical image. In the second cross-attention stage, the SAR feature representation is used as the query and the optical feature representation is used as the key and value to calculate the attention weight. Spectral features corresponding to the location are extracted from the optical feature representation to realize the mapping from SAR structural information to spectral information. The outputs of the two cross-attention stages are weighted and summed, and the summed result is input into the global self-attention strategy for information rearrangement and alignment to suppress single-path artifacts and generate structure-spectral consistent image fusion features.

[0065] The first cross-attention stage can be a feature complementation phase prioritized in a bidirectional cross-attention mechanism, utilizing the structural advantages of SAR features to compensate for structural deficiencies in cloud areas of optical images. The second cross-attention stage can be a feature mapping phase following the first stage in a bidirectional cross-attention mechanism, utilizing the spectral advantages of optical features to assign matching spectral attributes to SAR structural information. A query can be used to initiate a feature matching request. A key can be used to establish a feature association index. A value can be used to provide valid feature information to be extracted. Attention weights can be coefficients calculated based on the correlation between the query and the key, used to measure the strength of association between different features. Backscatter texture features can be features unique to SAR images that reflect the surface structure and morphology of ground objects. Spectral features can be features unique to optical images that reflect the color and spectral attributes of ground objects. A global self-attention strategy can be a processing method used to perform global-level information integration, association, and calibration of the bidirectional cross-attention output results. Information rearrangement and alignment can be a process of adjusting the spatial distribution and semantic association of fused features through a global self-attention strategy, ensuring that structural and spectral information are consistent in spatial location and semantic logic. Single-path artifacts can be feature distortions that arise from relying solely on a single modality feature or a single cross-attention process.

[0066] Specifically, in the multimodal fusion process of cloud removal from remote sensing images, the lack of a concrete implementation of a bidirectional cross-attention mechanism leads to the inability of SAR and optical features to achieve semantic-level bidirectional complementarity. This results in inaccurate cloud structure supplementation, misalignment between structure and spectrum, and the generation of single-path artifacts, leading to poor fused feature quality. To address these issues—optical images suffer from information loss due to cloud cover, while SAR images, although penetrating clouds, lack spectral information and are affected by speckle interference—existing fusion methods (such as cascading and addition) only establish unidirectional, shallow modal associations, making it difficult to maximize the utilization of complementary information at the semantic level. The adaptive multimodal fusion module utilizes bidirectional cross-attention, allowing optical and SAR features to perform bidirectional selection and complementary enhancement at each spatial location, thereby generating a fused representation with reliable cloud areas, rich texture, and consistent spectrum. This module consists of a bidirectional cross-transformer block, containing two symmetrical cross-attention modules and a global self-attention module. The entire process can be represented as: F fusion =GSAM(CSAM S2O (F) sar ,F opt )+CCAM O2S (F) opt ,F sar In the cross-attention stage, the attention query comes from the feature to be enhanced, and the key and value come from complementary features. (In CSAM...) S2O(·) This aims to supplement SAR texture features into cloudy optical imagery to address the problem of completely invisible land surfaces in areas of thick clouds. In CSAM S2O In parentheses (·), the query is from... Key and Value come from Next, by calculating the similarity between the Query and Key, the model will search for spatial-semantic patterns in FSAR that are similar to cloudless areas in Fopt, and then extract the backscattering features of the corresponding spatial locations from the Value to replace the cloud structure information in the cloudy optical image, thus achieving SAR-guided optical image cloud structure reconstruction. In CCAM O2S In parentheses (·), the query is from... Key and Value come from Next, by calculating the similarity between Query and Key, similar radiation patterns in SAR and optical features are matched, and then the spectral features at the corresponding positions are extracted from Value to realize the mapping from SAR structural information to spectral information. The results of the two cross-attention are added together and input into the global self-attention module GSAM(·) to re-align the completed information, suppress single-path artifacts and generate a structure-spectral consistent fusion result.

[0067] The method provided in this embodiment accurately fills the structural gaps in the cloud area through bidirectional feature interaction and global calibration via a bidirectional cross-attention mechanism, achieving precise matching between structure and spectrum, effectively suppressing single-path artifacts, and generating fused feature structures with clear and consistent spectra. This provides high-quality input for subsequent cloud removal and reconstruction, making the cloud removal image more in line with actual application needs and enhancing the practical value of remote sensing imagery in various scenarios.

[0068] In some embodiments, image fusion features are input into a multi-scale progressive decoder; progressive upsampling and feature reconstruction are performed through the multi-scale progressive decoder, and reconstruction results at different scales are output simultaneously; the reconstruction results are output as the final cloud-reconstructed optical image, and multi-level supervised training is performed using the reconstruction results at different scales.

[0069] Progressive upsampling involves gradually increasing the spatial resolution of feature maps according to a certain scale progression rule, achieving a gradual restoration process from low-resolution features to high-resolution images. Feature reconstruction can restore the local details and global structure of an image through the coordinated processing of the spatial and frequency domains, transforming feature data into image information with actual semantics. Reconstruction results at different scales can be image reconstruction products at multiple resolution levels output simultaneously during the progressive upsampling process, with each scale corresponding to a different level of feature representation. Multi-level supervised training can utilize the reconstruction results at different scales output by a multi-scale progressive decoder to provide multi-dimensional and multi-level training supervision for the model, enabling the model to learn accurate feature mapping relationships at each scale.

[0070] Specifically, in the process of reconstructing ground features in remote sensing imagery, the lack of multi-scale progressive decoding steps and reliance solely on single-scale features can lead to blurred details in high-resolution images (such as missing edges of small features), structural distortion in low-resolution images (such as misaligned terrain contours), and difficulty in completely eliminating cloud remnants. This directly affects the accuracy of subsequent data analysis and fails to meet the core requirements for image quality in precise remote sensing applications. To address these issues, image fusion features obtained through deep fusion via a bidirectional cross-attention mechanism, possessing both reliable structural and accurate spectral information, are input into a specially designed multi-scale progressive decoder. This decoder employs a progressive structure, using multiple sequentially stacked processing stages (e.g., including upsampling blocks). The decoder uses a spatial-frequency residual block to perform the reconstruction task. In each stage, the decoder first uses an upsampling block (using methods such as nearest neighbor interpolation) to upscale the spatial resolution of the input feature map to a predetermined scale. Then, it performs deep feature reconstruction through the spatial-frequency residual block. This residual block innovatively uses a spatial branch (extracting local details through cascaded small-sized convolutional kernels) and a frequency branch (converting features to the frequency domain through fast Fourier transform, using learnable convolutional layers to globally mix and correct the frequency components, and then transforming them back to the spatial domain) in parallel. This collaboratively mines and fuses the local subtle textures and global structural contour information of the image, thereby continuously repairing and enhancing the image content while increasing the resolution. During the gradual upsampling and reconstruction process, the decoder will simultaneously output intermediate reconstruction results corresponding to different resolution scales (such as 1 / 4, 1 / 2 and full size of the original size). Finally, the highest resolution reconstruction result is directly output as the cloud-reconstructed optical image. During the model training phase, these reconstruction results at different scales will be compared with the corresponding downsampled version of the real cloudless image and participate in the loss calculation to implement multi-level supervised training.

[0071] The method provided in this embodiment, through progressive upsampling and spatial-frequency domain collaborative processing, not only accurately restores local details (such as ground texture and edge contours) but also ensures global structural consistency, completely removes cloud interference, improves model stability through multi-level supervised training, and adapts multi-scale output to diverse application needs, significantly improving the reliability and accuracy of remote sensing applications.

[0072] In some embodiments, image fusion features are input into a multi-scale progressive decoder and processed through a progressive structure consisting of multiple upsampling blocks and spatial-frequency residual blocks. The upsampling blocks utilize nearest-neighbor interpolation to enlarge the spatial resolution of the feature maps, followed by channel dimensionality reduction and local spatial reconstruction using 3×3 convolutions. The spatial-frequency residual blocks include spatial and frequency branches: the spatial branch extracts local features through cascaded 3×3 convolutions to obtain local detail information; the frequency branch transforms the input features to the frequency domain using a fast Fourier transform, performs cross-channel mixing and correction of frequency components through learnable 1×1 convolutions, and then transforms them back to the spatial domain using an inverse fast Fourier transform to obtain global structural complementary information; finally, the local detail information and global structural complementary information are fused.

[0073] Upsampling blocks can be functional units in multi-scale progressive decoders used to improve the spatial resolution of feature maps. Nearest neighbor interpolation can be an interpolation method in upsampling blocks used to amplify the spatial resolution of feature maps. 3×3 convolution can be a convolution operation used for channel dimensionality reduction and local spatial reconstruction. Spatial-frequency residual blocks can be core units in multi-scale progressive decoders used for feature extraction and enhancement, including spatial and frequency branches. Spatial branches can be processing paths in spatial-frequency residual blocks focused on extracting local detail information. 3×3 convolution can be a convolution operation in spatial branches used for local feature extraction. Local detail information can be data reflecting local features such as fine textures and edge contours in the image, obtained after processing by spatial branches. Frequency branches can be processing paths in spatial-frequency residual blocks focused on extracting complementary information of global structure. Fast Fourier Transform (FFT) can be a mathematical transformation method to convert spatial features to the frequency domain. Cross-channel blending and correction can be processing methods in frequency branches to optimize frequency components. The Inverse Fast Fourier Transform (IFFT) is a mathematical transformation method that converts frequency domain features back to the spatial domain. Global structural complementarity information can be obtained by processing frequency domain branches, reflecting global features such as the overall layout and large-scale structural relationships of an image.

[0074] Specifically, in the process of cloud removal and reconstruction of remote sensing images, multi-scale progressive decoders are a crucial step. Existing decoders are prone to problems such as resolution and feature imbalance, limitations in spatial or frequency domain processing, disconnect between upsampling and channel optimization, and single feature extraction methods. These issues lead to numerous artifacts, loss of detail, and structural disorder in the reconstructed images. To address these problems, in the upsampling block, nearest-neighbor interpolation is used to double the spatial resolution of the feature map. This interpolation method has no learning parameters and can maintain the local statistical characteristics of the original feature distribution, effectively avoiding the checkerboard artifact problem commonly encountered in deconvolution and pixel recombination. Subsequently, a 3×3 convolution operation is performed on the feature map after nearest-neighbor interpolation, in the channel... While reducing dimensionality, the local space is reconstructed to further optimize the quality of the feature maps, providing higher-quality input for subsequent feature fusion and reconstruction. The spatial-frequency residual block adds a parallel frequency domain branch to the existing convolutional block's spatial domain, enabling learnable correction of global frequency components and improving the model's ability to model complex cloud scenes. This residual block contains two parallel information streams, which are ultimately fused through residual addition. The spatial branch uses two cascaded 3×3 convolutions, primarily responsible for texture restoration, edge enhancement, and detail compensation within the local receptive field, providing the local modeling capabilities of existing CNN models. The frequency domain branch shares the same feature input as the spatial branch. First, the input features are converted into the frequency domain using Fourier transform. Next, the complex tensor is decomposed into real and imaginary parts and concatenated along the channel dimension. Then, two 1×1 convolutions are used to perform cross-channel mixing of the frequency components, enabling the model to learn the relationships between different frequency components. This allows for learnable correction of global frequency information. The frequency domain branch establishes the long-range dependencies of the entire image at once through FFT, effectively repairing the missing frequency components in cloud regions. This path is particularly sensitive to spectral discontinuities caused by missing information in cloud regions, and can interpolate missing frequencies in the amplitude-phase joint space, thereby suppressing ringing and block artifacts in the restored image. The outputs of the frequency domain branch and the spatial domain branch are added element-wise and then residually connected to the input to complete the dual-path modeling of local convolution and global frequency domain. This achieves complementary enhancement of local details and global structure, mitigating reconstruction errors caused by large-scale cloud occlusion.

[0075] The method provided in this embodiment effectively balances resolution improvement and feature preservation through a progressive structure and dual-branch design, while capturing local details and global structure, avoiding artifacts and information loss. This results in rich details and coherent structure in the reconstructed image, while optimizing the processing flow, reducing redundancy, improving processing efficiency and stability, and adapting to the cloud removal needs of images with different cloud cover and scale.

[0076] Figure 3 This is a schematic diagram of the structure of a remote sensing image declouding system based on pseudo-twin feature extraction and multimodal fusion provided in an embodiment of this application, as shown below. Figure 3As shown, the remote sensing image declouding system 300 based on pseudo-twin feature extraction and multimodal fusion in this embodiment includes: a pseudo-twin feature extraction module 301, a cross-fusion module 302, a declouding reconstruction module 303, and an adversarial learning module 304.

[0077] The pseudo-twin feature extraction module 301 is used to acquire SAR images and cloudy optical images. The SAR images and cloudy optical images are input into the adversarial network GAN generator and processed by the adversarial network GAN generator to obtain adversarial generated optical images. The cross-fusion module 302 is used to process the adversarial generated optical images through an adaptive multimodal fusion strategy. It uses a bidirectional cross-attention mechanism to achieve deep complementarity and enhancement of dual-modal features to obtain image fusion features. The declouding reconstruction module 303 is used to process the image fusion features through a multi-scale progressive decoder, perform progressive upsampling and reconstruction, and output declouding reconstructed optical images. The adversarial learning module 304 is used to stitch the declouding reconstructed optical images and real cloudless optical images with SAR images respectively and input them into a multi-scale PatchGAN discriminator to obtain multi-scale discrimination results. Based on the multi-scale discrimination results, the generator and discriminator are jointly optimized through WGAN-GP loss until the model converges to obtain the optimal declouding model. Optionally, the pseudo-twin feature extraction module 301, during the generator's processing, is specifically used to: based on the SAR image, process it through the SAR feature extraction sub-branch in the pseudo-twin feature extraction branch to suppress speckle noise and extract structural texture features to obtain SAR feature representation; based on the multi-cloud optical image, process it through the optical feature extraction sub-branch in the pseudo-twin feature extraction branch to adaptively suppress cloud interference and retain effective spectral spatial features to obtain optical feature representation.

[0078] Optionally, during the construction process of the pseudo-twin feature extraction module 301, the following specific uses are employed: the front end of the SAR feature extraction sub-branch is equipped with a learnable frequency domain denoising head to suppress speckle noise in the SAR image in the frequency domain and maintain high-frequency structural information; the optical feature extraction sub-branch employs a gated convolution strategy for feature extraction to adaptively suppress cloud area features and retain cloudless area features; and the SAR feature extraction sub-branch and the optical feature extraction sub-branch are integrated to construct the pseudo-twin feature extraction branch.

[0079] Optionally, the pseudo-twin feature extraction module 301, in the specific implementation of the frequency domain denoising head, is specifically used to: convert the input SAR image to the complex frequency domain through Fast Fourier Transform (FFT) to obtain the amplitude spectrum and phase spectrum; generate a soft mask from the amplitude spectrum through a learnable convolutional layer, modulate the amplitude spectrum to suppress noise, and obtain the modulated amplitude spectrum; recombine the modulated amplitude spectrum and the phase spectrum, and convert them back to the spatial domain through Inverse Fast Fourier Transform (IFFT) to obtain preliminary denoising features.

[0080] Optionally, the pseudo-twin feature extraction module 301, in the specific implementation of the gated convolution strategy, is specifically used to: perform gated convolution processing on the input multi-cloud optical image, and split the output features along the channel dimension into a multi-cloud optical image feature branch and a multi-cloud optical image gated branch; the multi-cloud optical image gated branch generates spatially adaptive gate weights through processing including the Sigmoid function; and multiplies the gate weights with the output of the multi-cloud optical image feature branch pixel by pixel to achieve suppression of cloud area features and enhancement of cloudless area features.

[0081] Optionally, the cross-fusion module 302, during the construction of the image fusion features, is specifically used to: take the SAR feature representation and the optical feature representation as input; process them using a fusion strategy composed of a bidirectional cross-Transformer block, wherein each bidirectional cross-Transformer block contains two symmetrical cross-attention strategies and a global self-attention strategy; through the bidirectional cross-attention mechanism, enable the SAR feature representation and the optical feature representation to perform bidirectional selection and complementary enhancement at the semantic level, generating image fusion features that are reliable in cloud areas, rich in texture, and spectrally consistent.

[0082] Optionally, the cross-fusion module 302, in the specific implementation of the bidirectional cross-attention mechanism, is specifically used for: in the first cross-attention stage, using the optical feature expression as the query and the SAR feature expression as the key and value, calculating the attention weight, and extracting the backscattering texture features of the corresponding spatial location from the SAR feature expression to supplement and reconstruct the missing structural information of the cloud area in the multi-cloud optical image; in the second cross-attention stage, using the SAR feature expression as the query and the optical feature expression as the key and value, calculating the attention weight, and extracting the spectral features of the corresponding location from the optical feature expression to realize the mapping from SAR structural information to spectral information; weighted summing of the output results of the two cross-attention stages, and inputting the summed result into the global self-attention strategy for information rearrangement and alignment to suppress single-path artifacts and generate the image fusion features with consistent structure and spectrum.

[0083] Optionally, the cloud removal reconstruction module 303, during the construction process of the cloud removal reconstructed optical image, is specifically used to: input the image fusion features into the multi-scale progressive decoder; perform progressive upsampling and feature reconstruction through the multi-scale progressive decoder, and simultaneously output reconstruction results at different scales; output the reconstruction results as the final cloud removal reconstructed optical image, and use the reconstruction results at different scales for multi-level supervised training.

[0084] Optionally, the cloud removal and reconstruction module 303, in the specific implementation of the multi-scale progressive decoder, is specifically used to: input the image fusion features into the multi-scale progressive decoder, and process them through a progressive structure composed of multiple upsampling blocks and spatial-frequency residual blocks; wherein, the upsampling blocks use nearest-neighbor interpolation to enlarge the spatial resolution of the feature map, and then use 1×1 convolution to perform channel dimensionality reduction and local spatial reconstruction; the spatial-frequency residual blocks include spatial branches and frequency branches: the spatial branches extract local features through cascaded 3×3 convolutions to obtain local detail information; the frequency branches convert the input features to the frequency domain through fast Fourier transform, perform cross-channel mixing and correction of frequency components through learnable 1×1 convolution, and then convert them back to the spatial domain through inverse fast Fourier transform to obtain global structural complementary information; finally, the local detail information and the global structural complementary information are fused.

[0085] Optionally, the adversarial learning module 304, in the specific implementation of the adversarial learning, is specifically used for: in the discriminator stage, stitching the de-clouded reconstructed optical image or the real cloudless optical image with the corresponding SAR image along the channel dimension to obtain multi-channel conditional input features; inputting the multi-channel conditional input features and their 1 / 2 downsampled version and 1 / 4 downsampled version into the original scale discriminant branch, 1 / 2 scale discriminant branch, and 1 / 4 scale discriminant branch of the multi-scale PatchGAN discriminator, respectively; performing progressive downsampling and feature extraction through multiple convolutional layers of each discriminant branch, wherein each convolutional layer uses a 4×4 convolutional kernel with a stride of 2, and is combined with the LeakyReLU activation function, finally outputting the original scale confidence map and the 1 / 2 scale confidence map. Figure 1 shows a 1 / 4-scale confidence map, where each pixel value in the confidence map corresponds to the probability of authenticity of the local receptive field region of the input image at the corresponding scale. Based on each confidence map, the Wasserstein distance loss and gradient penalty term are calculated. The total loss of the multi-scale PatchGAN discriminator is obtained by averaging across the three scales. Maximizing the total discriminator loss improves the discriminator's ability to distinguish between real cloudless images and generated images. Specifically, all BatchNorm layers are removed from each discriminator branch structure to avoid the dynamic scaling and offset of batch statistics interfering with the Lipschitz continuity condition required for the Wasserstein distance constraint. The feature extraction path is constructed only through convolution operations and the LeakyReLU activation function. In the generator stage, the discriminator parameters are fixed, and the cloud-reconstructed result is stitched with the SAR image and input into the multi-scale PatchGAN to obtain the multi-scale confidence map output. The Wasserstein-based adversarial loss is then calculated. By minimizing the adversarial loss, the generator is driven to generate a distribution that approximates the distribution of real cloudless optical images. Repeat the above two steps alternately until the model converges and the quality of the generated images is stable, thus obtaining the optimal cloud removal model.

[0086] The system in this embodiment can be used to execute the methods of any of the above embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.

Claims

1. A cloud removal method for remote sensing images based on pseudo-twin feature extraction and multimodal fusion, characterized in that, include: Acquire SAR images and cloudy optical images, input the SAR images and cloudy optical images into the adversarial network GAN generator, process them through the adversarial network GAN generator to obtain adversarial generated optical images. Based on adversarial generative optical images, an adaptive multimodal fusion strategy is used to process them, and a bidirectional cross-attention mechanism is used to achieve deep complementarity and enhancement of bimodal features to obtain image fusion features. Based on image fusion features, the image is processed by a multi-scale progressive decoder to perform progressive upsampling and reconstruction, and outputs a cloud-reconstructed optical image. The de-clouded reconstructed optical image and the real cloudless optical image are stitched together with the SAR image and then input into the multi-scale PatchGAN discriminator to obtain the multi-scale discrimination result. Based on the multi-scale discrimination result, the generator and discriminator are jointly optimized through WGAN-GP loss until the model converges, and the optimal de-clouding model is obtained.

2. The method according to claim 1, characterized in that, The generator's processing includes: Based on the SAR image, it is processed by the SAR feature extraction sub-branch in the pseudo-twin feature extraction branch to suppress speckle noise and extract structural texture features, thereby obtaining SAR feature representation; Based on the multi-cloud optical image, the optical feature extraction sub-branch in the pseudo-twin feature extraction branch is used to process the image, adaptively suppressing cloud interference and retaining effective spectral spatial features to obtain optical feature representation. The adversarial network GAN generator is obtained through joint training of loss functions including multi-scale spatial frequency L1 loss and Wasserstein-based adversarial loss.

3. The method according to claim 1, characterized in that, The processing procedure of the multi-scale PatchGAN discriminator includes: The cloudless optical image or de-clouded optical image is stitched together with the corresponding SAR image along the channel dimension to obtain multi-channel conditional input features. The multi-channel conditional input features and their 1 / 2 downsampled and 1 / 4 downsampled versions are respectively input into the original scale discriminant branch, 1 / 2 scale discriminant branch, and 1 / 4 scale discriminant branch of the multi-scale PatchGAN discriminator; Progressive downsampling and feature extraction are performed through multiple convolutional layers in each of the discriminative branches. Each convolutional layer uses a 4×4 kernel with a stride of 2 and is coupled with the LeakyReLU activation function to capture local texture differences at the corresponding scale. Finally, the original scale confidence map, 1 / 2 scale confidence map, and 1 / 4 scale confidence map are output. Each pixel value in each confidence map corresponds to the authenticity probability of the local receptive field region of the input image at the corresponding scale. The WGAN-GP loss is calculated based on each of the confidence maps, and the generator and the discriminator are optimized through adversarial training. In the multi-scale PatchGAN discriminator, all BatchNorm layers are removed from each discriminant branch structure to avoid the dynamic scaling and offset of batch statistics interfering with the Lipschitz continuity condition required by the Wasserstein distance constraint. The feature extraction path is constructed only through convolution operations and activation functions.

4. The method according to claim 2, characterized in that, The construction process of the pseudo-twin feature extraction branch includes: The front end of the SAR feature extraction sub-branch is equipped with a learnable frequency domain denoising head, which is used to suppress speckle noise in SAR images in the frequency domain and maintain high-frequency structural information. The optical feature extraction sub-branch employs a gated convolution strategy for feature extraction, which adaptively suppresses cloud area features while retaining cloudless area features. The pseudo-twin feature extraction branch is constructed by integrating the SAR feature extraction sub-branch and the optical feature extraction sub-branch.

5. The method according to claim 4, characterized in that, The specific implementation of the frequency domain denoising head includes: The input SAR image is transformed to the complex frequency domain using Fast Fourier Transform (FFT) to obtain the amplitude spectrum and phase spectrum. The amplitude spectrum is passed through a learnable convolutional layer to generate a soft mask, and the amplitude spectrum is modulated to suppress noise, resulting in a modulated amplitude spectrum. The modulated amplitude spectrum is recombined with the original phase spectrum and transformed back to the spatial domain by inverse fast Fourier transform (IFFT) to obtain preliminary denoising features.

6. The method according to claim 4, characterized in that, The specific implementation of the gated convolution strategy includes: The input multi-cloud optical image is subjected to gated convolution processing, and the output features are split along the channel dimension into a multi-cloud optical image feature branch and a multi-cloud optical image gated branch. The multi-cloud optical image gating branch generates spatially adaptive gating weights through processing that includes a Sigmoid function; The gating weights are multiplied pixel-by-pixel with the output of the multi-cloud optical image feature branch to suppress cloud features and enhance cloudless features.

7. The method according to claim 6, characterized in that, The process of constructing the image fusion features includes: The SAR feature representation and the optical feature representation are used as inputs; The processing is performed using a fusion strategy consisting of a bidirectional cross Transformer block, wherein each bidirectional cross Transformer block contains two symmetric cross attention strategies and a global self-attention strategy. Through the bidirectional cross-attention mechanism, the SAR feature representation and the optical feature representation are mutually selected and complemented at the semantic level to generate image fusion features that are reliable in cloud areas, rich in texture, and consistent in spectrum.

8. The method according to claim 7, characterized in that, The specific implementation of the bidirectional cross-attention mechanism includes: In the first cross-attention stage, the optical feature expression is used as the query and the SAR feature expression is used as the key and value to calculate the attention weight. Backscatter texture features corresponding to the spatial location are extracted from the SAR feature expression to supplement and reconstruct the missing structural information of the cloud area in the multi-cloud optical image. In the second cross-attention stage, the SAR feature expression is used as the query and the optical feature expression is used as the key and value. By calculating the attention weight, the corresponding spectral context is extracted from the optical feature expression to supplement the missing spectral information in the SAR mode. The outputs of the two cross-attention stages are added together, and the result is input into the global self-attention strategy for information rearrangement and alignment to suppress single-path artifacts and generate structure-spectral consistent image fusion features.

9. The method according to claim 8, characterized in that, The process of constructing the cloud-reconstructed optical image includes: The image fusion features are input into the multi-scale progressive decoder; The multi-scale progressive decoder performs stepwise upsampling and feature reconstruction, and outputs reconstruction results at different scales simultaneously. The reconstruction results are used as the final cloud-reconstructed optical image output, and multi-level supervised training is performed using the reconstruction results at different scales.

10. The method according to claim 9, characterized in that, The specific implementation of the multi-scale progressive decoder includes: The image fusion features are input into the multi-scale progressive decoder and processed through a progressive structure consisting of multiple upsampling blocks and spatial-frequency residual blocks. The upsampling block uses nearest neighbor interpolation to enlarge the spatial resolution of the feature map, and then uses 3×3 convolution to perform channel dimensionality reduction and local spatial reconstruction. The spatial-frequency residual block includes a spatial branch and a frequency branch: The spatial branch extracts local features through cascaded 3×3 convolutions to obtain local detail information; The frequency domain branch transforms the input features to the frequency domain via Fast Fourier Transform, performs cross-channel mixing and correction on the frequency components through learnable 1×1 convolution, and then transforms them back to the spatial domain via Inverse Fast Fourier Transform to obtain global structural complementary information. Finally, the local detail information is fused with the global structural complementary information.

11. A remote sensing image cloud removal system based on pseudo-twin feature extraction and multimodal fusion, characterized in that, The method applied to any one of claims 1-10 includes: The pseudo-twin feature extraction module is used to acquire SAR images and cloudy optical images. The SAR images and cloudy optical images are input into the adversarial network GAN generator and processed by the adversarial network GAN generator to obtain adversarial generated optical images. The cross-fusion module is used to process adversarial generative optical images through an adaptive multimodal fusion strategy. It utilizes a bidirectional cross-attention mechanism to achieve deep complementarity and enhancement of bimodal features, thereby obtaining image fusion features. The cloud removal and reconstruction module is used to process image fusion features through a multi-scale progressive decoder, perform progressive upsampling and reconstruction, and output cloud removal and reconstruction optical images. The adversarial learning module is used to stitch the de-clouded reconstructed optical image and the real cloudless optical image with the SAR image and input them into the multi-scale PatchGAN discriminator to obtain the multi-scale discrimination result. Based on the multi-scale discrimination result, the generator and discriminator are jointly optimized through WGAN-GP loss until the model converges and the optimal de-clouding model is obtained.