Optical remote sensing image cloud removing method and device
Patent Information
- Application Number
- CN202611093534.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-22
- Publication Date
- 2026-09-25
AI Technical Summary
这种处理范式与云退化的空间变化特性存在本质冲突
(1)退化感知路由提升融合精度。通过从含云输入中自动估计局部退化程度并据此进行补丁级路由,使模型容量流向退化最严重的区域,而不是在空间上均匀分配。这种空间自适应的处理方式有效避免了在晴空区域过度引入SAR信息以及在厚云区域引入不足的问题,从而提升了重建精度和光谱一致性。
Smart Images

Figure CN122820474A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular to a method and apparatus for removing clouds from optical remote sensing images. Background Technology
[0002] Optical satellite remote sensing imagery serves as a crucial data foundation for various Earth observation missions, including precision agriculture, disaster response, and climate monitoring. However, cloud contamination frequently impacts the availability of optical data, as over half of the Earth's land surface is obscured by clouds at any given time. This contamination disrupts the temporal continuity of optical observations and propagates errors into downstream analytical tasks. In tasks such as land cover classification, change detection, and semantic segmentation, missing or damaged pixels severely undermine the reliability of interpretation results.
[0003] Early cloud removal methods primarily relied on multi-temporal synthesis or sparse representation. These methods essentially assume the availability of cloud-free reference observations from adjacent imaging dates, an assumption often unattainable in real-world scenarios. With the development of deep learning, the field has gradually shifted towards cross-modal reconstruction of SAR and optics, with synthetic aperture radar (SAR) becoming an important auxiliary mode due to its all-weather imaging capabilities. However, SAR images themselves suffer from severe speckle noise, limited spectral information, and significant domain differences with optical reflectivity, making simple stitching or shallow fusion insufficient to achieve ideal results. Existing methods uniformly apply a single fusion strategy across the entire spatial domain, completely ignoring the severity of local optical signal degradation. In fact, cloud contamination within a single optical scene exhibits significant spatial heterogeneity. Some areas are completely obscured by thick clouds, requiring reconstruction based solely on structural priors provided by SAR. Other areas are covered by thin clouds or haze, still carrying some optical information, while adjacent areas may be completely unobstructed. A spatially unified fusion mechanism inevitably introduces excessive SAR information into clear-sky areas, producing structural artifacts, while introducing insufficient SAR information into thick cloud areas, resulting in blurry and structurally incomplete reconstructions. This processing paradigm is fundamentally in conflict with the spatial variation characteristics of cloud degradation. Summary of the Invention
[0004] This application provides a method and apparatus for removing clouds from optical remote sensing images, which can detect the activity level of applications and provide data reference for application optimization.
[0005] In a first aspect, this application provides a method for cloud removal from optical remote sensing images, including: Acquire synthetic aperture radar (SAR) images and cloud-containing optical remote sensing images of the same observation area; Features of the SAR image and the cloud-containing optical remote sensing image are extracted separately using a SAR encoder and an optical encoder, and the two types of features are stitched together along the channel dimension to obtain joint features. Input the cloud-containing optical remote sensing image into the degradation estimation head to obtain the global degradation descriptor and spatial degradation map; The joint features are divided into non-overlapping patches, and the dense weight vector of each patch relative to multiple candidate experts is calculated based on the global degradation descriptor and the spatial degradation map corresponding to each patch position. Sparse selection is performed on the dense weight vector, and the highest K values in the dense weight vector are retained to obtain the sparse gating weights corresponding to each patch; the candidate experts corresponding to the retained values in the sparse gating weights are the activated experts. Each patch is input into the activated expert for processing to obtain the output of the activated expert. The activated experts include one or more of the following: SAR structure expert, optical texture expert, and cross-modal bridging expert. The outputs of each activated expert are weighted and aggregated according to sparse gating weights to obtain the fusion features of each patch. The fusion features of each patch are then rearranged according to their spatial location to obtain the complete fusion feature map. The fully fused feature map is input into the decoder to generate a cloud-free remote sensing image.
[0006] For example, the step of inputting the fully fused feature map into the decoder to generate a cloud-free remote sensing image includes: The complete fused feature map is input into the global refinement module, which uses window multi-head self-attention and shift window multi-head self-attention to extract features from the complete fused feature map and obtain the refined tensor. The refined tensor is compressed according to the following formula to obtain the compressed features;
[0007] in, For the obtained compression features, μ and σ are the conditional distributions q(Z|F) and q(Z|F) respectively. ref The mean and standard deviation of ), where ε is the noise independently sampled from the standard normal distribution N(0, I), and the sign is... This represents element-wise multiplication; The refined tensor is input into the decoder, which includes a main fusion decoder and two auxiliary decoders; The main fusion decoder reconstructs the cloud-free remote sensing image based on the refined tensor, while the two auxiliary decoders reconstruct the SAR image and the cloud-containing image respectively through compressed features.
[0008] For example, the above method further includes: The main fusion loss is calculated based on the de-clouded remote sensing image and the corresponding real cloudless image of the observation area. The auxiliary reconstruction loss is calculated based on the SAR reconstructed image, the cloud-containing reconstructed image, and the original SAR image and the cloud-containing optical remote sensing image. Obtain the degradation score of each patch from the spatial degradation graph corresponding to each patch, and calculate the routing loss of the patch on each candidate expert based on the degradation score; Calculate the load balancing loss based on the average probability of each candidate expert being activated. Calculate the KL divergence loss based on the conditional distribution of the refining tensor; Optimize network parameters based on the main fusion loss, the auxiliary reconstruction loss, the routing loss, the load balancing loss, and the KL divergence loss.
[0009] For example, calculating the routing loss of the patch on each candidate expert based on the degradation score includes: The group-level prediction probability of patch routing to each candidate expert is calculated using the following formula;
[0010] Spatial degradation graph corresponding to the patch In the patch The degradation fraction obtained by average aggregation within the range, To control the positive hyperparameters of the sharpness of target distribution in the SAR and optical groups, 4 (1- )exist A value close to 0.5 is relatively large and is used to emphasize cross-modal bridging experts in the transition region. To prevent the normalization denominator from being zero, the stability constant, For the normalized first Group-level prediction probability of each candidate expert; The routing loss is calculated based on the group-level prediction probabilities of each candidate expert using the following formula:
[0011] This represents the set of all patch locations that participate in route supervision within a training batch. To determine the number of experts, c iterates through the three candidate experts: sar, opt, and bridge. For the first Group-level prediction probability of each candidate expert It is the stability constant for logarithmic operations.
[0012] For example, the step of inputting each patch into the activated expert for processing to obtain the output of the activated expert includes: Low-frequency components are extracted from joint features using a low-pass operator, and high-frequency residuals are calculated based on the low-frequency components. The low-pass operator is implemented by average pooling and bilinear upsampling. For the activated SAR structure expert, the macroscopic structure prior is extracted by the Swing Transformer block of the SAR structure expert with low-frequency components, and then after residual transformation, it is combined with the channel space attention module of the SAR structure expert with high-frequency residual input. The features extracted by the channel space attention module are merged with the joint features to obtain the output of the SAR structure expert. For an activated optical texture expert, the high-frequency residual is input into the SwinTransformer block of the optical texture expert to extract texture features. After residual transformation, the features are input into the channel spatial attention module of the optical texture expert along with the low-frequency component. The features extracted by the channel spatial attention module are merged with the joint features to obtain the output of the optical texture expert. For the activated cross-modal bridging expert, the joint features are divided into SAR sub-features and optical sub-features along the channel dimension, and flattened into N spatially labeled context features. Calculate the query, key, and value of SAR sub-features and optical sub-features respectively, and calculate the bidirectional cross-attention feature of the context feature based on the query, key, and value of SAR sub-features and optical sub-features; The bidirectional cross-attention features are concatenated along the channel dimension. Based on the spatial label, the concatenated features are restored into a spatial feature map. The features of the spatial feature map are extracted through the channel spatial attention module of the cross-modal bridging expert. The extracted features are then merged with the joint features to obtain the output of the cross-modal bridging expert.
[0013] For example, the weighted aggregation of the outputs of each activated expert based on sparse gating weights includes: The outputs of each activated expert are weighted and aggregated according to the following formula based on the sparse gating weights:
[0014] Where E represents the number of candidate experts. Indicates the first One candidate expert on the patch The output, This represents the sparse gating weight of the candidate expert at position p. This indicates the fusion characteristics of the patch; inactive experts meet the requirements. .
[0015] Exemplarily, the method further includes: The candidate experts are two SAR structure experts, two optical texture experts, and two cross-modal bridging experts. Experts of the same type use the same network structure and do not share learnable parameters. The sparse gating weights are K-sparse, where K is 2.
[0016] For example, the step of calculating the main fusion loss based on the de-clouded remote sensing image and the corresponding real cloudless image of the observation area includes: The formula for calculating the main fusion loss is:
[0017] This represents the perceptual consistency loss derived from the pre-trained VGG19 network. , , , To balance the empirical weighting coefficients of each penalty term, J is the cloud-removed remote sensing image output by the main fusion decoder. This is a true cloudless image. This represents the consistency loss of the features extracted by the pre-trained SSIM network. It is the Sobel gradient loss of the image;
[0018] This represents the Sobel gradient operator. These represent the number of channels, height, and width of the image, respectively. The absolute value is used to calculate the element-wise L1 difference between the predicted edge and the ground truth edge.
[0019] For example, the step of inputting the fully fused feature map into the decoder to generate a cloud-free remote sensing image includes: The joint features are input into the SAR conditional skip gating to obtain the gating map. The gating map is then multiplied element-wise with the features of the cloud-containing optical remote sensing image extracted by the optical encoder to obtain the gating optical skip features. SAR conditional skip gating is processed by convolution, batch normalization, GELU activation, and convolution. The resulting gated graph is a gated graph with values between 0 and 1. The gated optical skip features are stitched together with the joint features, and the stitched features are input together with the complete fused feature map into the decoder to generate a cloud-free remote sensing image.
[0020] Secondly, this application provides an optical remote sensing image cloud removal device, comprising: The image acquisition module is used to acquire synthetic aperture radar (SAR) images and cloud-containing optical remote sensing images of the same observation area; The shallow feature extraction module is used to extract features from the SAR image and the cloud-containing optical remote sensing image respectively through the SAR encoder and the optical encoder, and to stitch the two features along the channel dimension to obtain joint features; The degradation perception module is used to input cloud-containing optical remote sensing images into the degradation estimation head to obtain global degradation descriptors and spatial degradation maps; The patch routing module is used to divide the joint features into non-overlapping patches and calculate the dense weight vector of each patch relative to multiple candidate experts based on the global degradation descriptor and the spatial degradation map corresponding to each patch position. The expert pool module is used to perform sparse selection on the dense weight vector, retaining the highest K values in the dense weight vector to obtain the sparse gating weights corresponding to each patch; the candidate experts corresponding to the retained values in the sparse gating weights are the activated experts. Expert decoupling processing is used to input each patch into the activated expert for processing and obtain the output of the activated expert. The activated experts include one or more of SAR structure experts, optical texture experts, and cross-modal bridging experts. The feature fusion module is used to weight and aggregate the outputs of each activated expert based on sparse gating weights to obtain the fused features of each patch, and rearrange the fused features of each patch according to spatial position to obtain a complete fused feature map. The decoding and reconstruction module is used to input the complete fused feature map into the decoder to generate a cloud-free remote sensing image.
[0021] Thirdly, this application provides an electronic device including a memory and one or more processors. The memory stores one or more computer programs, each including instructions that, when executed by the processor, cause the electronic device to perform the optical remote sensing image cloud removal method as described in the first aspect.
[0022] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the optical remote sensing image cloud removal method as described in the first aspect.
[0023] Fifthly, this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the optical remote sensing image cloud removal method as described in the first aspect.
[0024] The beneficial effects that the aforementioned optical remote sensing image cloud removal device, electronic device, computer-readable storage medium, and computer program product can achieve are as follows: (1) Degradation-aware routing improves fusion accuracy. By automatically estimating the local degradation level from cloud-containing inputs and performing patch-level routing accordingly, the model capacity is directed to the most degraded areas, rather than being evenly distributed spatially. This spatially adaptive approach effectively avoids the problems of over-introducing SAR information in clear-sky areas and under-introducing it in thick cloud areas, thereby improving reconstruction accuracy and spectral consistency.
[0025] (2) Heterogeneous experts achieve functional specialization. The expert pool constructed in this scheme includes SAR structure experts, optical texture experts and cross-modal bridging experts. These different experts use different special operators in frequency domain processing and information flow direction, realizing explicit functional specialization. This overcomes the shortcomings of traditional homogeneous expert architecture that only relies on parameter differentiation, enabling the same backbone network to apply different operators to thick clouds, clear skies and transition patches.
[0026] (3) High robustness in complex environments. It exhibits high robustness in complex scenarios with strong cloud pollution, complex ground object interference, and severe degradation of optical signals. Through spatial adaptive fusion of degradation perception, this application can achieve stable reconstruction results with fewer parameters and less computation, thus improving the reliability and stability of cloud removal algorithms in practical Earth observation applications. Attached Figure Description
[0027] Figure 1 A schematic flowchart illustrating the optical remote sensing image cloud removal method provided in this application embodiment; Figure 2 Network architecture diagram of the optical remote sensing image cloud removal method provided in the embodiments of this application Figure 1 ; Figure 3 Network architecture diagram of the optical remote sensing image cloud removal method provided in the embodiments of this application Figure 2 ; Figure 4 Network architecture diagram of the optical remote sensing image cloud removal method provided in the embodiments of this application Figure 3 ; Figure 5 Network architecture diagram of the optical remote sensing image cloud removal method provided in the embodiments of this application Figure 4 ; Figure 6 Network architecture diagram of the optical remote sensing image cloud removal method provided in the embodiments of this application Figure 5 . Detailed Implementation
[0028] To facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. For example, "first chip" and "second chip" are only used to distinguish different chips and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" do not necessarily imply that they are different. It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being better or more advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner. In the embodiments of this application, "at least one" means one or more, and "more than one" means two or more.
[0029] It should be noted that "at the time of..." in the embodiments of this application can be either at the instant when a certain situation occurs, or for a period of time after the occurrence of a certain situation. The embodiments of this application do not make specific limitations on this.
[0030] The implementation of this embodiment will now be described in detail with reference to the accompanying drawings.
[0031] This embodiment provides a method for removing clouds from optical remote sensing images, which can monitor the activity of various business functions of the application and provide data references for the maintenance and updating of the application.
[0032] For example, this optical remote sensing image cloud removal method can be applied to various electronic devices such as computers (PCs), tablets, virtual reality / augmented reality devices, wearable devices, industrial computers, and vehicle-mounted systems; it can also be applied to servers, cloud environments, server clusters, etc., and this embodiment does not impose any special limitations on it.
[0033] Figure 1 A schematic flowchart of the optical remote sensing image cloud removal method provided in this application embodiment is shown.
[0034] like Figure 1 As shown, the method for removing clouds from optical remote sensing images may include the following steps: Step 101: Acquire synthetic aperture radar (SAR) images and cloud-containing optical remote sensing images of the same observation area.
[0035] Step 102: Extract features from the SAR image and the cloud-containing optical remote sensing image using the SAR encoder and optical encoder respectively, and then stitch the two features along the channel dimension to obtain joint features.
[0036] Step 103: Input the cloud-containing optical remote sensing image into the degradation estimation head to obtain the global degradation descriptor and spatial degradation map.
[0037] Step 104: Divide the joint features into non-overlapping patches, and calculate the dense weight vector of each patch relative to multiple candidate experts based on the global degradation descriptor and the spatial degradation map corresponding to each patch position.
[0038] Step 105: Perform sparse selection on the dense weight vector, retain the highest K values in the dense weight vector, and obtain the sparse gating weights corresponding to each patch; the candidate experts corresponding to the retained values in the sparse gating weights are the activated experts.
[0039] Step 106: Input each patch into the activated expert for processing to obtain the output of the activated expert. The activated expert includes one or more of the following: SAR structure expert, optical texture expert, and cross-modal bridging expert.
[0040] Step 107: Based on the sparse gating weights, the outputs of each activated expert are weighted and aggregated to obtain the fusion features of each patch. The fusion features of each patch are then rearranged according to their spatial positions to obtain the complete fusion feature map.
[0041] Step 108: Input the complete fused feature map into the decoder to generate a cloud-free remote sensing image.
[0042] Understandably, the above steps can be the image processing flow of a neural network model. Therefore, this application also provides an optical remote sensing image cloud removal model. Figure 2 The overall architecture of the model is shown, as follows: Figure 2 As shown, the model mainly includes a SAR encoder, an optical encoder, a degradation-aware patch routing module, a heterogeneous frequency decoupling expert pool, and a task-aware information flow optimization module. The degradation-aware patch routing module is used to implement steps 103 to 105, the heterogeneous frequency decoupling expert pool includes three different types of experts, and the task-aware information flow optimization module is used to implement step 108.
[0043] Specifically, in step 108, the decoder includes three paths: one main fusion decoder and two auxiliary decoders. The three decoders decode based on the complete fusion feature map, obtaining three reconstructed images. The main fusion decoder reconstructs a cloud-free remote sensing image, while the auxiliary decoders reconstruct a SAR reconstructed image and a cloud-containing reconstructed image. The specific process is as follows: the complete fusion feature map is input into the global refinement module. The global refinement module uses window multi-head self-attention and shifted window multi-head self-attention to extract features from the complete fusion feature map, obtaining a refined tensor; the refined tensor is compressed according to the following formula to obtain compressed features.
[0044] in, For the obtained compression features, μ and σ are the conditional distributions q(Z|F) and q(Z|F) respectively. ref The mean and standard deviation of ), where ε is the noise independently sampled from the standard normal distribution N(0, I), and the sign is... This represents element-wise multiplication; the refined tensor is input into the decoder, which includes a main fusion decoder and two auxiliary decoders; the main fusion decoder reconstructs the cloud-free remote sensing image based on the refined tensor, and the two auxiliary decoders reconstruct the SAR reconstructed image and the cloud-containing reconstructed image respectively through compressed features.
[0045] Since cloud removal deals with spatially heterogeneous degenerate distributions, this scheme constructs the cloud removal problem as a dynamic hybrid expert problem. The entire network is trained end-to-end without explicit cloud mask supervision. This technique consists of three tightly coupled modules, with the forward process sequentially going through three stages: local degradation estimation, specialized feature transformation, and task-aware representation routing.
[0046] The first stage is degradation-aware patch routing. The degradation estimation head in this module evaluates cloud-containing observations, inferring a global degradation vector and a spatial degradation map. These two signals jointly drive a patch-level router, dynamically assigning local image regions to corresponding experts based on their degradation level. The entire process does not rely on any external cloud annotations. Considering that pixel-by-pixel routing would incur excessive computational overhead and produce severe spatial discontinuities, while image-level routing would lose its local adaptive capability, this implementation sets the routing strategy at the patch level.
[0047] The second stage is a heterogeneous frequency decoupling expert pool. This expert pool replaces the structurally identical experts in traditional routing designs, comprising three types of experts with different functions: SAR structure experts for thick cloud reconstruction, optical texture experts for high-frequency detail restoration in thin cloud regions, and cross-modal bridging experts responsible for modal transition harmonization. SAR backscattering and optical reflectivity occupy different frequency subbands. SAR is related to low-frequency terrain geometry, while optical reflectivity carries most of the high-frequency texture details. Based on this, this application introduces a frequency decoupling design, allowing different experts to process data in their respective preferred frequency subbands.
[0048] The third stage is task-aware information flow optimization. This module resolves the representation conflict between the auxiliary reconstruction task and the main fusion task. The auxiliary reconstruction task requires minimum sufficient statistics to drive the network to compress the latent space and discard task-irrelevant variations, while the main fusion task requires less compressed representations to retain complete cross-modal complementary information. Forcing these two divergent objectives to share the same single representation can cause task interference. This invention explicitly splits the feature trajectory in two through variational information bottlenecks and prevents shallow features contaminated by clouds from propagating into the synthesis path through skip gating under SAR conditions.
[0049] In feature modeling, this application also introduces a degradation-guided routing regularization term. Unlike mask-supervised routing, this regularization term does not use external cloud annotations. Instead, it derives a soft group-level target distribution from the degradation map predicted by the degradation estimation head, encouraging SAR experts to activate in heavily degraded regions, optical experts in clear or lightly degraded regions, and bridging experts in transitional regions. The lightweight frequency decoupling expert design has lower computational costs and is suitable for deployment in resource-constrained environments.
[0050] In step 101, SAR images and cloud-containing optical remote sensing images are acquired.
[0051] In step 102, the SAR image With cloud-containing optical remote sensing images The encoded features are obtained by SAR encoder and optical encoder respectively, and then concatenated along the channel dimension to form joint feature F.
[0052] Meanwhile, in step 103, only the optical remote sensing image is... The degradation estimation head is fed in to obtain the global degradation descriptor. and spatial degradation diagram .
[0053] First, the optical remote sensing image is modeled as a spatially varying hybrid signal containing both an implicit clear-sky signal J and an atmospheric scattering component C. Let Ω be the spatial domain of the image, and the observed cloud-containing optical image... The model is a spatial variation mixture of the implicit clear-sky signal J and the atmospheric scattering component C, as shown in formula (1): (1) in, Indicates spatial location index, The value ranges from 0 to 1, representing the local degradation state at that location. When When it approaches 1, it indicates that the pixel Dominated by the cloud, when A value close to 0 indicates that the pixel corresponds to a clear sky observation. exist The continuous changes above result in spatial heterogeneous degradation observed in real-world scenarios. Local degradation state. In network implementation, it is represented by a global descriptor. and local degradation value Common characteristics.
[0054] To achieve parameterization under degradation conditions, degradation-aware patch routing designs a degradation estimation head that integrates cloud-containing remote sensing images. Mapping to the joint latent space yields the global degeneracy descriptor. and spatial degradation diagram . Figure 3 The architecture diagram of the degradation-aware patch routing module is shown. Figure 3 As shown, remote sensing images containing clouds The input degradation estimation head applies four convolutions with a stride of 2, batch normalization, and GELU activation blocks to the input cloud-containing remote sensing image. It then splits into two branches. One branch passes through global average pooling and a fully connected layer, followed by sigmoid activation to generate a global degradation descriptor. Another branch generates a spatial degradation map through convolution and sigmoid activation. For example, the global descriptor dimension is set to 6, corresponding to the total number of experts, establishing a correspondence between descriptor channels and each expert. The degradation estimation head is trained without explicit cloud mask supervision and optimized through downstream reconstruction and routing objectives. Here, The 6-dimensional designation is intended to enable the formation of a global routing prior for six experts after the embedding function and projection matrix are applied. A single dimension is not directly equivalent to the final weight of a particular expert; the final weight is a combination of factors. and calculate.
[0055] Since it is not feasible to learn a separate set of parameters for each spatial location, this embodiment achieves this conditional mapping through a local mixture of experts with a fixed set of parameters. The aggregation coefficient is determined by the local degenerate state, as shown in formula (2): (2) in, Indicates the first One expert, Indicates that the expert pool and Coupled gate functions, The total number of experts. The aggregation coefficient in the formula. These weights are not pre-defined, but are specifically implemented by the degradation-aware patch routing module described later. This module calculates a set of sparse gating weights for each patch based on the local degradation state. , That is, the specific value of the aggregation coefficient in formula (2), and the weighted summation of the outputs of each expert in formula (9) to obtain the fusion feature. Equations (3) to (9) represent the step-by-step implementation of this polymerization process. Furthermore, the local degenerate state in equation (2) In network implementation, it is represented by a global descriptor. and local degradation value Common representation, gate function The numerical value corresponds to the formula (4) to obtain the first sparse gating weights Therefore, formula (2) gives an abstract expression of the mixture of degenerate conditions, formulas (3) and (4) are responsible for obtaining the aggregation coefficient, formulas (5) to (8) are responsible for calculating the output of each expert, and formula (9) completes the actual aggregation.
[0056] In step 104, the joint feature F is divided into non-overlapping patches to be processed. Spatial degradation diagram By aligning the patch grid with interpolation or regional averaging, the spatial degradation map portion corresponding to each patch location is obtained. , Then it will be broadcast to all patch locations.
[0057] Based on the global degradation descriptor and the spatial degradation graph corresponding to each patch location, each patch location... Get a cross A dense weight vector of available experts As shown in formula (3): (3) in, The projection matrix is learnable. and Indicates an embedded function. This indicates splicing along the channel dimension.
[0058] For formula 3, the direct input is... and Its output is the same as A dense weight vector with one-to-one positional correspondence; subsequently, this weight controls... Which experts will be included?
[0059] Dense routing allocates redundant capacity and dilutes the specialization of experts. In step 105, sparsity is forced during computation while maintaining the backpropagation gradient flow, employing a soft Top-K routing strategy with hard sparsity selection to construct a binary mask. Only retain dense weight vectors The highest K values are selected, and then normalized. The normalized sparse gating weights are... Calculate using formula (4): (4) in, It represents the Hadamah accumulation. To ensure stable small positive numbers, gradients are propagated through selected softmax weights, while discrete Top-K indices are used for hard routing decisions.
[0060] In one implementation, the candidate experts are two SAR structure experts, two optical texture experts, and two cross-modal bridging experts. Experts of the same type use the same network structure, do not share learnable parameters, and have sparse gating weights of Ksparse, where K is 2.
[0061] Traditional hybrid expert architectures typically employ a homogeneous pool of experts who share the same functionalities, differing only in the optimized parameters. This uniform paradigm is unsuitable for cloud degradation tasks because different degradation states require different computational processing. To address this, this application constructs a structurally heterogeneous expert pool consisting of six experts divided into three functionally specialized subsets: SAR structural experts, optical texture experts, and cross-modal bridging experts. Two experts are deployed for each category, enabling each subset to further specialize its processing for submodalities within its target region.
[0062] Six experts receive the same input: a patch from the joint features corresponding to the current position. However, only the experts selected by the Top-K routes perform valid computations. To maintain consistency with subsequent routing regularization terms, the six experts are uniformly numbered: the 1st and 2nd are SAR structure experts, the 3rd and 4th are optical texture experts, and the 5th and 6th are cross-modal bridging experts. Two experts of the same type use the same network structure and processing logic, but have independent and non-shared learnable parameters; therefore, they are not sequential or fixed pairings, but rather two parallel candidate operators. During training, under the combined effect of reconstruction loss, routing regularization, and load balancing, two experts of the same type can further learn different sub-modes in their target regions. For example, two SAR structure experts can adapt to different land cover structures in thick cloud regions, and two optical texture experts can adapt to different texture statistics in clear sky or slightly degraded regions. During inference, each patch activates only K=2 experts. The two experts can come from the same type or different types, and their outputs are ultimately determined according to their respective... Weighted summation. The superscript 'i' in the formula refers to the global ID of the corresponding expert.
[0063] In step 106, the six experts receive the same input: a patch from the joint features corresponding to the current position. However, only the output of the experts selected by the Top-K routes participates in the final fusion. The expert execution process is as follows: Low-frequency components are extracted from joint features using a low-pass operator, and high-frequency residuals are calculated based on these low-frequency components. The low-pass operator is implemented using average pooling and bilinear upsampling. For the activated SAR structure expert, the low-frequency components are input into the SAR structure expert's Swing Transformer block to extract macroscopic structural priors. After residual transformation, these priors are combined with the high-frequency residuals and input into the SAR structure expert's channel spatial attention module. The features extracted by the channel spatial attention module are then merged with the joint features to obtain the output of the SAR structure expert. For the activated optical texture expert, the high-frequency residuals are input into the optical texture expert's Swing Transformer block. The Transformer block extracts texture features, which are then transformed using residuals and input to the channel spatial attention module of the optical texture expert along with the low-frequency components. The features extracted by the channel spatial attention module are merged with the joint features to obtain the output of the optical texture expert. For the activated cross-modal bridging expert, the joint features are divided into SAR sub-features and optical sub-features along the channel dimension and flattened into N spatially labeled context features. The queries, keys, and values of the SAR sub-features and optical sub-features are calculated respectively. The queries, keys, and values of the SAR sub-features and optical sub-features are used to calculate the bidirectional cross-attention features of the context features. The bidirectional cross-attention features are concatenated along the channel dimension, and the concatenated features are restored to a spatial feature map based on the spatial labels. The features of the spatial feature map are extracted by the channel spatial attention module of the cross-modal bridging expert, and the extracted features are merged with the joint features to obtain the output of the cross-modal bridging expert.
[0064] Figure 4 The architecture of the expert pool is shown, such as Figure 4 As shown, frequency decomposition is performed first. It should be noted that this frequency decomposition applies to the joint feature F of the input expert pool, which is the feature obtained by concatenating the SAR and optically coded features mentioned above. First, low-frequency components are extracted using a low-pass operator. The low-pass operator is composed of a step size of The average pooling is followed by bilinear upsampling, and then the high-frequency residuals are derived. equal minus .
[0065] For SAR structure experts, only Feeding the data into the Swing Transformer block to synthesize macroscopic structural priors while bypassing [other mechanisms] Branching is used to reduce high-frequency interference, followed by mitigation of boundary artifacts during feature integration via a channel spatial attention module, and global residual connections are used to preserve the fidelity of the input features. The formula is expressed as: (5) Where F represents the joint features of the input expert. and They represent the low-frequency component and the high-frequency residual obtained by the low-pass operator decomposition, respectively. This indicates that the expert used the Swing Transformer block to synthesize macroscopic structure priors in the low-frequency subband. Represents the residual block. The channel spatial attention module is used to mitigate boundary artifacts during feature integration. The superscript i indicates that the parameters of each module are unique to the i-th expert and are not shared among them. This represents the output of the SAR structure expert. In the formula, for... Branches are bypassed directly without Swing processing to reduce high-frequency interference, and global residual connections at the ends are used to preserve the fidelity of input features.
[0066] Optical texture experts through high-frequency subband Applying self-attention to achieve texture restoration while preserving As a structural anchor point, as shown in formula (6): (6) in, Let F represent the output of the i-th optical texture expert; F represents the joint features input to that expert. and These represent the low-frequency component and the high-frequency residual obtained by the low-pass operator decomposition, respectively. This indicates the expert's proprietary SwingTransformer block, which operates on high-frequency residuals. To restore texture; Represents residual transformation, This indicates a channel-space joint attention refinement module; low-frequency components. As structural anchors, they participate in reconstruction together with the processed high-frequency components, with the outer layer added... This constitutes a global residual connection. For this type of expert, the global numbering... Choose 3 or 4.
[0067] For transitional regions such as fog or cloud edges, relying on a single modality is insufficient to achieve ideal fidelity. This embodiment designs a cross-modal bridging expert to achieve bidirectional feature interaction, Divided into SAR sub-features along the channel dimension and optical sub-characteristics Flatten it into N spatial markers and calculate expert-specific query, key, and value projections, with bidirectional cross-attention as shown in Equation 7: (7) in, , , These represent expert-specific queries, key-value projections, and key-value projections, respectively. The embedding dimension for each attention head. Furthermore, This represents the cross-attention output obtained by using SAR sub-features as queries and optical sub-features as keys and values, indicating that the SAR location reads complementary information from the optical sub-features. This represents the cross-attention output with opposite directions, where the query, key, and value are obtained by linear projection of the sub-features of the corresponding modality through the linear projections of each expert. The superscript T indicates matrix transpose, and division by the square root of d is used to scale the dot product to stabilize the values. Softmax normalizes along the dimension of the key.
[0068] The joint feature F is composed of SAR-coded features and optically-coded features concatenated along the channel dimension. Let the total number of channels be C. During partitioning, the first half of F, C / 2 channels, is taken as SAR sub-features according to channel order. The latter half, C / 2 channels, are taken as optical sub-features. These two correspond to the outputs of two different modal encoders, thus decoupling SAR and optical information in the channel dimension. The specific embodiment described above ensures that the two encoded features have the same number of channels, so they can be divided into C / 2 equal parts; if different channel configurations are used, then... and As a boundary, satisfying The division principle is still based on extracting two types of sub-features from the modal channel intervals retained during splicing.
[0069] The representations obtained from mutual attention are reshaped, spliced, and refined by the channel space attention module to obtain the expert's output, as shown in formula (8): (8) Note that the arrow subscript is used to mark the query and the queried modality: with Generate queries, with When generating keys and values, the output retains the number of SAR-marked locations, but injects the optical context; the reverse branch works similarly.
[0070] Concat concatenates the cross-attention results from two directions along the channel dimension, while Reshape restores the N labels to a spatial feature map. This represents the unique channel—the spatial joint attention refinement module—of the i-th bridging expert. This represents the output of the bridging expert, with the outer layer F being the global residual term. For this type of expert, the global number i is either 5 or 6.
[0071] The three types of experts share the same outer structure of residual and channel-space attention, differing only in their internal frequency domain operations or cross-modal operations. This is reflected in the patch-level sparse gating graph derived from the degradation-aware patch routing. Under the guidance of [the relevant authority], expert outputs are aggregated in a sparse weighted manner.
[0072] In step 107, the outputs of each activated expert are weighted and aggregated according to the following formula based on the sparse gating weights: (9) Where E represents the number of candidate experts. Indicates the first One candidate expert on the patch The output, This represents the sparse gating weight of the candidate expert at position p. This indicates the fusion characteristics of the patch; inactive experts meet the requirements. .
[0073] Sparse Gating Weights That is, the patch-level gating tensor is passed backward, and the outputs of each expert in the heterogeneous expert pool are weighted and aggregated in formula (9), thereby transforming the result of local degradation estimation into the fusion strategy that actually takes effect on each patch. After the expert processing is completed, the weighted outputs of each patch are rearranged according to their original spatial positions, thereby restoring the complete fusion feature map. .
[0074] Equations (1.5), (1.6), and (1.8) define the transformations of the input features by the SAR structure expert, optical texture expert, and cross-modal bridging expert, respectively, thus providing the specific forms of each expert operator in Equation (1.9). Equation (1.9) then uses the gating weights obtained from Equation (1.4) to further refine these transformations. The outputs of each expert are weighted and summed. In other words, when the i-th expert belongs to the SAR type, its output is calculated by formula (1.5); when it belongs to the optical type, it is calculated by formula (1.6); and when it belongs to the bridging type, it is calculated by formulas (1.7) and (1.8). These three together constitute the aggregated term of formula (1.9). In formula (1.9), E=6. Indicates the first An expert on the patch The output, This indicates the normalized sparse weight of the expert at position p. This indicates the aggregation result at that position. Since the experts not selected by Top-K meet the criteria... Although formula (1.9) is written as a summation over all 6 experts, only 2 terms actually produce non-zero contributions.
[0075] because It is K-sparse with K equal to 2, and each spatial patch has only two experts generating non-zero weights. Combining the functional heterogeneity among expert subsets, the heterogeneous frequency-decoupled expert pool achieves functional specialization along the degenerate spectrum while maintaining computational efficiency.
[0076] After selective aggregation is completed in the heterogeneous expert pool, composite characterization is performed. (i.e., the complete fused feature map) is routed to the decoding path.
[0077] In step 108, most multi-task architectures feed this shared latent state into all decoders, implicitly assuming that a single representation can optimally serve each downstream objective. This embodiment improves upon the shortcomings of this design. The auxiliary reconstruction task requires a minimum sufficient statistic, while the main fusion task requires a less compressed representation to preserve full-spectrum cross-modal complementary information. Forcing these two divergent objectives to share the same single representation would cause task interference.
[0078] To reconcile this representational conflict, a task-aware information flow optimization module is introduced to explicitly split the feature trajectory into two parts. Figure 5 The architecture diagram of the task-aware information flow optimization module is shown. (Refer to...) Figure 5 For the main fusion path, the uncompressed refined tensor Equal to the global refinement module The output is directly transmitted to preserve high-frequency details. For the auxiliary reconstruction path, a variational information bottleneck is applied, which... Projection is based on the mean and variance The parameters are given by a random latent distribution, and then a compressed representation is synthesized using a reparameterization technique. The formula is as follows: (10) in, For the obtained compression features, μ and σ are the conditional distributions q(Z|F) and q(Z|F) respectively. ref The mean and standard deviation of ), where ε is the noise independently sampled from the standard normal distribution N(0, I), and the sign is... This indicates element-wise multiplication.
[0079] The global refining module It consists of a pair of Swing Transformer blocks, which respectively employ windowed multi-head self-attention (W-MSA) and shifted windowed multi-head self-attention (SW-MSA) to aggregate the fused features. Global context modeling and refinement are performed, and the output is the uncompressed refined tensor mentioned above. , recorded as equal Acting on Mean μ and log-variance Each by acting on The two learnable projection heads on the top output, and then... Obtain the standard deviation Sampling from the standard normal distribution Then, reparameterization is performed according to formula (10). This introduces stochastic compression into the forward process and allows the gradient to pass through... and Backpropagation is performed to the refining module and encoder.
[0080] The main fusion path is not used Instead, it directly uses the uncompressed version. Auxiliary paths use This is to avoid two types of tasks competing for the same potential representation.
[0081] The resulting compression characterization The data is passed along the auxiliary reconstruction path and fed into the SAR reconstruction decoder and optical reconstruction decoder to reconstruct the original SAR input and the cloud-containing optical input. The reconstruction error is constrained by the auxiliary reconstruction loss, while the mean of this random latent distribution is... With variance Constrained by information bottleneck losses, thus limiting With refined tensor The mutual information between the input and the reconstructed embedding is explicitly constrained by this stochastic bottleneck, which acts as an information filter to separate the structural prior from mode-specific noise.
[0082] Outside of the deep latent space, U-shaped networks typically employ lateral skip connections to pass multi-scale spatial priors to the decoder. However, in cloud removal tasks, shallow optical features often carry unfiltered cloud contamination, which propagates directly to the decoder through skip connections, thus degrading the quality of the synthesized output. To reduce this leakage, this application designs a skip gating mechanism under SAR conditions.
[0083] Figure 6 The architecture diagram of SAR conditional skip gating is shown, such as... Figure 6As shown, the joint feature obtained by concatenating the features extracted by the optical encoder (shallow features of the cloud-containing optical remote sensing image) and the features extracted by the SAR encoder (shallow features of the SAR image) is used as the input to the SAR conditional skip gating. The joint feature is input into the SAR conditional skip gating to obtain a gating map. The gating map is then multiplied element-wise with the features of the cloud-containing optical remote sensing image extracted by the optical encoder to obtain the gated optical skip features. The SAR conditional skip gating undergoes convolution processing, batch normalization, GELU activation, and further convolution processing. The resulting gating map has values between 0 and 1. The gated optical skip features are then concatenated with the joint feature. The concatenated features and the complete fused feature map are then input into the decoder to generate a cloud-free remote sensing image.
[0084] SAR conditional skip gating processes the input features sequentially with 3×3 convolution, batch normalization (BN), GELU activation, and 3×3 convolution to obtain the gated map.
[0085] Specifically, the concatenated features serve as the input to the decoder, which can be understood as consisting of two parts: one part is the deep features, where the main fusion decoder uses the uncompressed refined tensor. The auxiliary decoders for SAR and optical signals use compressed features. The other part is the shallow skip features provided by the encoder. The shallow optical features (i.e. the features extracted by the optical encoder) are first subjected to SAR conditional gating to suppress cloud-contaminated information, and then spliced with the SAR shallow features and injected into the corresponding scale of the main fusion decoder.
[0086] Using the all-weather penetration of SAR modes as an invariant structural reference, this gating measure of paired shallow embedding and The local consistency between them is shown in Equation (11): (11) in, This indicates Sigmoid activation, and BN indicates batch normalization. and These represent SAR shallow hop features and optical shallow hop features at the same scale, respectively, with symbols... Indicates channel splicing. and This represents a two-layer learnable mapping. This represents a gated graph with values between 0 and 1. This represents the optical jump characteristics after gating. When the learned bimodal response difference is significant, indicating that there may be heavy cloud cover at that location, Approaching zero, thus effectively suppressing contaminated optical components in skip connections while preserving SAR structural reference. Conversely, in clear-sky areas, gating remains active to preserve optical texture. This selective adjustment ensures that the fusion decoder receives only clean and task-relevant features. In Equation (11), the gating does not directly determine cloud areas through a fixed difference formula, but rather uses the stitched dual-modal local response to learn a consistency relationship; when the learned response determines that the optical features are significantly inconsistent with the SAR structure, By taking a smaller value at the corresponding position, contaminated optical skipping information is suppressed.
[0087] This embodiment also includes a multi-objective loss function to supervise the main fusion task, auxiliary reconstruction, and internal representation dynamics. Specifically, it includes: calculating the main fusion loss based on the de-clouded remote sensing image and the corresponding real cloudless image of the observation area; calculating the auxiliary reconstruction loss based on the SAR reconstructed image, the cloud-containing reconstructed image, and the original SAR image and the cloud-containing optical remote sensing image; obtaining the degradation score of each patch from the spatial degradation map corresponding to each patch, and calculating the routing loss of the patch on each candidate expert based on the degradation score; calculating the load balancing loss based on the average probability of each candidate expert being activated; calculating the KL divergence loss based on the conditional distribution of the refined tensor; and optimizing the network parameters based on the main fusion loss, the auxiliary reconstruction loss, the routing loss, the load balancing loss, and the KL divergence loss.
[0088] The total objective loss consists of the main fusion loss, auxiliary reconstruction loss, routing loss, load balancing loss, and KL divergence loss, expressed by the formula: (12) The main fusion loss is used to minimize the reconstructed image. with truth value The differences between them. The auxiliary reconstruction loss for SAR and cloud-included optical input, For the group-level routing loss of degradation guidance, L bal For expert load balancing losses, KL divergence loss for variational information bottleneck; , , and These are the non-negative weights of the four auxiliary items mentioned above, which balance the contributions of each item.
[0089] To further preserve high-frequency boundaries in the clean reference image, this embodiment incorporates a Sobel-based gradient loss to provide explicit pixel-level supervision signals for image edges: (13) For the cloud-free image generated by the fusion decoder, J is the corresponding cloudless ground truth value, c, h, and w are the channel, row, and column indices, respectively, C, H, and W are the corresponding dimension sizes, and the absolute value is used to calculate the element-wise L1 difference between the predicted edge and the ground truth edge.
[0090] (14) in, This represents the perceptual consistency loss derived from deep features of a pre-trained VGG19 network. , , , To balance the empirical weighting coefficients of each penalty term, the first term constrains the pixel radiance value. Constraining local structural similarity, Constrain high-level semantics to maintain consistency with perception. Constrain high-frequency edges; , , and The contribution of each of the four losses is controlled separately.
[0091] In the auxiliary reconstruction loss, the auxiliary decoder represents the compression bottleneck. The original SAR input and the cloud-inclusive optical input are reconstructed in the middle, prompting the encoder to retain the necessary structural priors, as shown in Equation (15): (15) and These are the original SAR input and the original optical input with clouds, respectively. and respectively with The input consists of the outputs of two auxiliary decoders; the two L1 errors together constitute the input. This loss allows the compressed representation to retain the information necessary to complete cross-modal structure reconstruction, but it does not directly replace the main fusion output. .
[0092] To align routing behavior with the estimated degradation state, this embodiment introduces a degradation-guided routing regularization term. This term does not use external cloud annotations but instead derives a soft group-level target from the degradation map predicted by the degradation estimation head. Let the SAR expert group consist of experts 1 and 2, the optical expert group consist of experts 3 and 4, and the bridging expert group consist of experts 5 and 6. For each patch... First, from the patch-level degradation graph Obtain its degradation score The value ranges from 0 to 1, with a larger value indicating more severe cloud degradation. Subsequently, a soft target distribution is constructed on the three expert groups.
[0093] Specifically, the routing loss of the patch on each candidate expert is calculated based on the degradation score, including: The group-level prediction probability (i.e., soft target distribution) of patch routing to each candidate expert is calculated using the following formula. (16) Spatial degradation graph corresponding to the patch In the patch The degradation fraction obtained by average aggregation within the range, To control the positive hyperparameters of the sharpness of target distribution in the SAR and optical groups, 4 (1- )exist A value close to 0.5 is relatively large and is used to emphasize cross-modal bridging experts in the transition region. To prevent the normalized denominator from being zero, the stability constant, For the normalized first Group-level prediction probability of each candidate expert.
[0094] This design encourages SAR experts to activate in heavily degraded regions, optical experts to activate in clear or slightly degraded regions, and bridging experts to activate in transitional regions.
[0095] Group-level routing probabilities are densely gated before the Top-K mask. The calculation is as shown in Formula 17: (17) in, , and Each is determined by the dense weights within the corresponding expert set. , Summing them up yields the result, thus the three elements constitute a patch. Predicted probabilities across the three functional groups.
[0096] The routing loss is calculated based on the group-level prediction probabilities of each candidate expert using the following formula: (18) This represents the set of all patch locations that participate in route supervision within a training batch. To determine the number of experts, c iterates through the three candidate experts: sar, opt, and bridge. For the first Group-level prediction probability of each candidate expert It is the stability constant for logarithmic operations.
[0097] Here, sg represents the stopping gradient operation applied to the target branch to prevent the soft target distribution from being directly optimized by the routing loss. Dense gating is used for this regularization term to stabilize the optimization, while sparse Top-K gating is used for expert execution.
[0098] To avoid expert collapse, a switch-type load balancer is used, as shown in formula (19): (19) in, The patch ratio for the i-th expert in the Top-K set. Let be the average dense route probability assigned to the i-th expert in this batch. The total number of experts is 6. This reflects the frequency with which the i-th expert is actually selected. Reflecting the The average probability obtained by an expert in the dense routing phase; here, the expert level... Group-level probabilities, different from those in formula (1.17) , and The sum of their products can simultaneously constrain the hard selection frequency and the soft assignment probability, reducing the risk of routing becoming concentrated in the hands of a few experts in the long run.
[0099] Random paths in variational information bottlenecks, minimizing the potential distribution With standard Gaussian prior The KL divergence between them is shown in formula (20): (20) Where d is the dimension of the latent space. This term constrains... and The mutual information between them prompts the reconstruction branch to retain only the structural information relevant to the task. Indicates by Parameterized approximate posterior distribution, This represents the standard Gaussian prior. and The potential space is respectively The mean and variance of the dimension, This represents the KL divergence. This term directly affects the formula (10). and And together with the reconstruction loss of formula (15) constrain it. .
[0100] In the specific implementation, the framework of the above model is based on PyTorch. Radiometric normalization is applied to both SAR and optical inputs before network input to bring them into comparable dynamic ranges. The digital quantization values of the optical image are truncated to the upper limit of 5000 and then linearly scaled to the standard dynamic range of 0 to 1. The VV channel of the SAR backscattering coefficient is truncated to -25 to 0 dB, and the VH channel is truncated to -32.5 to 0 dB, and then proportionally min-max normalized to the 0 to 1 interval. During the training phase, spatial data augmentation such as random flipping, multi-angle rotation, and mesh distortion are used to alleviate overfitting. The network parameters are optimized using the Adam optimizer, with an initial learning rate of 0.0001, a batch size of 4, and a maximum training epoch of 200. An adaptive scheduling mechanism is used to monitor the validation loss. When validation performance stagnates for three consecutive epochs, the learning rate is reduced to one-fifth of its original value, and further reductions are made when the validation loss stagnates for six consecutive epochs or the learning rate decays below 5 × 10⁻⁶. -8 Training is terminated via an early stop protocol to obtain the trained optical remote sensing image cloud removal model.
[0101] Furthermore, this embodiment also provides an optical remote sensing image cloud removal device, which can be used to perform the above-described optical remote sensing image cloud removal method. The optical remote sensing image cloud removal device specifically includes: an image acquisition module, used to acquire synthetic aperture radar (SAR) images and cloud-containing optical remote sensing images of the same observation area; a shallow feature extraction module, used to extract features from the SAR image and the cloud-containing optical remote sensing image respectively through a SAR encoder and an optical encoder, and to concatenate the two features along the channel dimension to obtain joint features; a degradation perception module, used to input the cloud-containing optical remote sensing image into a degradation estimation head to obtain a global degradation descriptor and a spatial degradation map; a patch routing module, used to divide the joint features into non-overlapping patches, and to calculate the dense weight vector of each patch relative to multiple candidate experts based on the global degradation descriptor and the spatial degradation map corresponding to each patch position; and an expert pool module, used to process the dense weight vector. The dense weight vector is used for sparse selection, retaining the K highest values to obtain the sparse gating weights for each patch. The candidate experts corresponding to the retained values in the sparse gating weights are the activated experts. Expert decoupling processing is used to input each patch into the activated experts for processing to obtain the output of the activated experts. The activated experts include one or more of SAR structure experts, optical texture experts, and cross-modal bridging experts. The feature fusion module is used to weight and aggregate the outputs of each activated expert according to the sparse gating weights to obtain the fused features of each patch, and rearrange the fused features of each patch according to spatial location to obtain a complete fused feature map. The decoding and reconstruction module is used to input the complete fused feature map into the decoder to generate a cloud-free remote sensing image.
[0102] The specific details of each module or unit in the aforementioned optical remote sensing image cloud removal device have been described in detail in the corresponding optical remote sensing image cloud removal method, so they will not be repeated here.
[0103] This application also provides an electronic device, including a processor and a memory, wherein the memory stores one or more computer programs, and the one or more computer programs include instructions that, when executed by the electronic device, implement the above-described method. Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication portion of a network interface card such as a LAN card or modem, and / or installed from a removable medium such as a disk, optical disk, magneto-optical disk, or semiconductor memory. When the computer program is executed by a processor (CPU), it performs the functions defined in the embodiments of this application.
[0104] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0106] The units described in the embodiments of this disclosure can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the unit itself.
[0107] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which include instructions that, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.
[0108] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0109] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for cloud removal in optical remote sensing images, characterized in that, include: Acquire synthetic aperture radar (SAR) images and cloud-containing optical remote sensing images of the same observation area; Features of the SAR image and the cloud-containing optical remote sensing image are extracted separately using a SAR encoder and an optical encoder, and the two types of features are stitched together along the channel dimension to obtain joint features. Input the cloud-containing optical remote sensing image into the degradation estimation head to obtain the global degradation descriptor and spatial degradation map; The joint features are divided into non-overlapping patches, and the dense weight vector of each patch relative to multiple candidate experts is calculated based on the global degradation descriptor and the spatial degradation map corresponding to each patch position. Sparse selection is performed on the dense weight vector, retaining the highest K values in the dense weight vector to obtain the sparse gating weights corresponding to each patch. The candidate experts corresponding to the values retained in the sparse gating weights are the activated experts; Each patch is input into the activated expert for processing to obtain the output of the activated expert. The activated experts include one or more of the following: SAR structure expert, optical texture expert, and cross-modal bridging expert. The outputs of each activated expert are weighted and aggregated according to sparse gating weights to obtain the fusion features of each patch. The fusion features of each patch are then rearranged according to their spatial location to obtain the complete fusion feature map. The fully fused feature map is input into the decoder to generate a cloud-free remote sensing image.
2. The method for removing clouds from optical remote sensing images according to claim 1, characterized in that, The process of inputting the fully fused feature map into the decoder to generate a cloud-free remote sensing image includes: The complete fused feature map is input into the global refinement module, which uses window multi-head self-attention and shift window multi-head self-attention to extract features from the complete fused feature map and obtain the refined tensor. The refined tensor is compressed according to the following formula to obtain the compressed features; in, For the obtained compression features, μ and σ are the conditional distributions q(Z|F) and q(Z|F) respectively. ref The mean and standard deviation of ), where ε is the noise independently sampled from the standard normal distribution N(0, I), and the sign is... This represents element-wise multiplication; The refined tensor is input into the decoder, which includes a main fusion decoder and two auxiliary decoders; The main fusion decoder reconstructs the cloud-free remote sensing image based on the refined tensor, while the two auxiliary decoders reconstruct the SAR image and the cloud-containing image respectively through compressed features.
3. The method for removing clouds from optical remote sensing images according to claim 2, characterized in that, Also includes: The main fusion loss is calculated based on the de-clouded remote sensing image and the corresponding real cloudless image of the observation area. The auxiliary reconstruction loss is calculated based on the SAR reconstructed image, the cloud-containing reconstructed image, and the original SAR image and the cloud-containing optical remote sensing image. Obtain the degradation score of each patch from the spatial degradation graph corresponding to each patch, and calculate the routing loss of the patch on each candidate expert based on the degradation score; Calculate the load balancing loss based on the average probability of each candidate expert being activated. Calculate the KL divergence loss based on the conditional distribution of the refining tensor; Optimize network parameters based on the main fusion loss, the auxiliary reconstruction loss, the routing loss, the load balancing loss, and the KL divergence loss.
4. The method for removing clouds from optical remote sensing images according to claim 3, characterized in that, The calculation of the routing loss of the patch on each candidate expert based on the degradation score includes: The group-level prediction probability of patch routing to each candidate expert is calculated using the following formula; Spatial degradation graph corresponding to the patch In the patch The degradation fraction obtained by average aggregation within the range, To control the positive hyperparameters of the sharpness of target distribution in the SAR and optical groups, 4 (1- )exist A value close to 0.5 is relatively large and is used to emphasize cross-modal bridging experts in the transition region. To prevent the normalization denominator from being zero, the stability constant, For the normalized first Group-level prediction probability of each candidate expert; The routing loss is calculated based on the group-level prediction probabilities of each candidate expert using the following formula: This represents the set of all patch locations that participate in route supervision within a training batch. To determine the number of experts, c iterates through the three candidate experts: sar, opt, and bridge. For the first Group-level prediction probability of each candidate expert It is the stability constant for logarithmic operations.
5. The method for removing clouds from optical remote sensing images according to claim 1, characterized in that, The process of inputting each patch into the activated expert for processing and obtaining the output of the activated expert includes: Low-frequency components are extracted from joint features using a low-pass operator, and high-frequency residuals are calculated based on the low-frequency components. The low-pass operator is implemented by average pooling and bilinear upsampling. For the activated SAR structure expert, the macroscopic structure prior is extracted by the Swing Transformer block of the SAR structure expert with low-frequency components, and then after residual transformation, it is combined with the channel space attention module of the SAR structure expert with high-frequency residual input. The features extracted by the channel space attention module are merged with the joint features to obtain the output of the SAR structure expert. For an activated optical texture expert, the high-frequency residual is input into the expert's Swing Transformer block to extract texture features. After residual transformation, these features are input into the optical texture expert's channel spatial attention module along with the low-frequency component. The features extracted by the channel spatial attention module are then merged with the joint features to obtain the output of the optical texture expert. For the activated cross-modal bridging expert, the joint features are divided into SAR sub-features and optical sub-features along the channel dimension, and flattened into N spatially labeled context features. Calculate the query, key, and value of SAR sub-features and optical sub-features respectively, and calculate the bidirectional cross-attention feature of the context feature based on the query, key, and value of SAR sub-features and optical sub-features; The bidirectional cross-attention features are concatenated along the channel dimension. Based on the spatial label, the concatenated features are restored into a spatial feature map. The features of the spatial feature map are extracted through the channel spatial attention module of the cross-modal bridging expert. The extracted features are then merged with the joint features to obtain the output of the cross-modal bridging expert.
6. The method for removing clouds from optical remote sensing images according to claim 1, characterized in that, The weighted aggregation of the outputs of each activated expert based on sparse gating weights includes: The outputs of each activated expert are weighted and aggregated according to the following formula based on the sparse gating weights: Where E represents the number of candidate experts. Indicates the first One candidate expert on the patch The output, This represents the sparse gating weight of the candidate expert at position p. This indicates the fusion characteristics of the patch; inactive experts meet the requirements. .
7. The method for removing clouds from optical remote sensing images according to claim 1, characterized in that, The method further includes: The candidate experts are two SAR structure experts, two optical texture experts, and two cross-modal bridging experts. Experts of the same type use the same network structure and do not share learnable parameters. The sparse gating weights are K-sparse, where K is 2.
8. The method for removing clouds from optical remote sensing images according to claim 3, characterized in that, The calculation of the main fusion loss based on the de-clouded remote sensing image and the corresponding real cloudless image of the observation area includes: The formula for calculating the main fusion loss is: This represents the perceptual consistency loss derived from the pre-trained VGG19 network. , , , To balance the empirical weighting coefficients of each penalty term, J is the cloud-removed remote sensing image output by the main fusion decoder. This is a true cloudless image. This represents the consistency loss of the features extracted by the pre-trained SSIM network. It is the Sobel gradient loss of the image; This represents the Sobel gradient operator. These represent the number of channels, height, and width of the image, respectively. The absolute value is used to calculate the element-wise L1 difference between the predicted edge and the ground truth edge.
9. The method for removing clouds from optical remote sensing images according to claim 1, characterized in that, The process of inputting the fully fused feature map into the decoder to generate a cloud-free remote sensing image includes: The joint features are input into the SAR conditional skip gating to obtain the gating map. The gating map is then multiplied element-wise with the features of the cloud-containing optical remote sensing image extracted by the optical encoder to obtain the gating optical skip features. SAR conditional skip gating is processed by convolution, batch normalization, GELU activation, and convolution. The resulting gated graph is a gated graph with values between 0 and 1. The gated optical skip features are stitched together with the joint features, and the stitched features are input together with the complete fused feature map into the decoder to generate a cloud-free remote sensing image.
10. A cloud removal device for optical remote sensing images, characterized in that, include: The image acquisition module is used to acquire synthetic aperture radar (SAR) images and cloud-containing optical remote sensing images of the same observation area; The shallow feature extraction module is used to extract features from the SAR image and the cloud-containing optical remote sensing image respectively through the SAR encoder and the optical encoder, and to stitch the two features along the channel dimension to obtain joint features; The degradation perception module is used to input cloud-containing optical remote sensing images into the degradation estimation head to obtain global degradation descriptors and spatial degradation maps; The patch routing module is used to divide the joint features into non-overlapping patches and calculate the dense weight vector of each patch relative to multiple candidate experts based on the global degradation descriptor and the spatial degradation map corresponding to each patch position. The expert pool module is used to perform sparse selection on the dense weight vector, retaining the highest K values in the dense weight vector to obtain the sparse gating weights corresponding to each patch. The candidate experts corresponding to the values retained in the sparse gating weights are the activated experts; Expert decoupling processing is used to input each patch into the activated expert for processing and obtain the output of the activated expert. The activated experts include one or more of SAR structure experts, optical texture experts, and cross-modal bridging experts. The feature fusion module is used to weight and aggregate the outputs of each activated expert according to the sparse gating weights to obtain the fused features of each patch, and rearrange the fused features of each patch according to their spatial positions to obtain a complete fused feature map. The decoding and reconstruction module is used to input the complete fused feature map into the decoder to generate a cloud-free remote sensing image.