A remote sensing image cloud removal method and system based on optical remote sensing images and synthetic aperture radar images
By employing feature extraction and adaptive fusion mechanisms from optical and SAR branches, the problems of modal differences and radiometric consistency in cross-modal fusion declouding methods between optical remote sensing images and synthetic aperture radar images are solved, achieving high-quality remote sensing image declouding reconstruction and improving image quality under complex cloud conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HENAN UNIVERSITY
- Filing Date
- 2026-03-26
- Publication Date
- 2026-07-21
AI Technical Summary
Existing cross-modal fusion declouding methods for optical remote sensing images and synthetic aperture radar images suffer from problems such as difficulty in reconciling modal differences, insufficient detail recovery, and poor radiometric consistency, resulting in poor declouding performance, especially under complex cloud conditions where the performance still needs improvement.
The design incorporates a feature extraction module and an adaptive fusion mechanism. Features are extracted through optical and SAR branches respectively, and a hierarchical cross-modal adaptive fusion module is used to generate adaptive weights based on the semantic and radiometric differences of features from different modalities for fusion. Combined with a residual multi-path aggregation module and a multi-head convolutional self-attention mechanism, high-quality cloud removal and reconstruction of remote sensing images is achieved.
It effectively balances local details and global structure, improves the stability and reconstruction quality of cloud removal and reconstruction, enhances radiometric consistency and spectral fidelity, and improves the cloud removal effect of images under complex cloud conditions.
Smart Images

Figure CN122434779A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image processing technology, and in particular to a method and system for cloud removal from remote sensing images based on optical remote sensing images and synthetic aperture radar images. Background Technology
[0002] Optical remote sensing images, with their rich spectral information and texture details, can provide periodic and large-scale surface information, playing an irreplaceable role in fields such as urban planning, land use monitoring, and environmental assessment. However, optical remote sensing images are highly susceptible to cloud interference during acquisition, leading to image quality degradation, the creation of unobservable areas of surface information, and breaks in temporal continuity, severely impacting subsequent image analysis and application.
[0003] To address these issues, existing technologies have proposed various cloud removal methods, mainly including single-source image thin cloud removal, multi-temporal optical image interpolation cloud removal, and cross-sensor fusion cloud removal from heterogeneous images. Among these, synthetic aperture radar (SAR) images, with their all-weather, all-day imaging capabilities and ability to penetrate cloud layers to obtain surface structure information, have become a key technological support for cross-sensor fusion cloud removal. However, optical and SAR images exhibit significant differences in imaging mechanisms and cross-modal gaps, resulting in issues such as missing details and radiometric inconsistencies even after registration processing.
[0004] Existing SAR-assisted optical image declouding methods suffer from incomplete information recovery and blurred details in areas covered by thick clouds. While fusion-based methods utilize modal complementarity to some extent, they lack differentiated representation mechanisms for different modal characteristics and deep feature interaction mechanisms, resulting in insufficient modal complementarity and difficulty in maintaining radiometric consistency and structural integrity. Therefore, the declouding performance under complex cloud conditions still needs improvement. Thus, a remote sensing image declouding scheme that can effectively balance local details and global structure and adaptively adjust the contribution of multimodal features is urgently needed. Summary of the Invention
[0005] To address the problems of poor cloud removal performance in existing cross-modal fusion methods of optical and SAR images, which suffer from difficulties in reconciling modal differences, insufficient detail recovery, or poor radiometric consistency, this invention provides a novel remote sensing image cloud removal method and system based on optical remote sensing images and synthetic aperture radar images. This invention achieves high-quality cloud removal and reconstruction of remote sensing images under complex cloud conditions by designing a feature extraction module and an adaptive fusion mechanism.
[0006] In a first aspect, the present invention provides a method for cloud removal from remote sensing images based on optical remote sensing images and synthetic aperture radar images, comprising:
[0007] Step 1: Acquire paired optical remote sensing images and SAR images;
[0008] Step 2: Input the optical remote sensing image into the optical branch to extract optical features, and input the SAR image into the SAR branch to extract SAR features;
[0009] Step 3: Input the optical features and the SAR features into the hierarchical cross-modal adaptive fusion module to generate adaptive weights based on the semantic and radiometric differences of different modal features, and fuse the optical features and the SAR features based on the adaptive weights to obtain fused features;
[0010] Step 4: Process the fused features using an activation function to obtain a cloudless optical remote sensing image.
[0011] Furthermore, the optical branch includes sequentially connected convolutional layers and multiple stacked residual multipath aggregation modules (RMPABs); wherein, the RMPAB includes a patch embedding layer and a Transformer layer; the Transformer layer includes multiple parallel convolutional paths, a multi-head attention mechanism, layer normalization, and a feedforward network; wherein, the receptive fields of each convolutional path are different; each convolutional path is used to extract features of the feature map processed by the patch embedding layer under its corresponding receptive field;
[0012] Correspondingly, inputting the optical remote sensing image into the optical branch to extract optical features includes: inputting the optical remote sensing image into a convolutional layer to extract initial optical features; inputting the initial optical features into a patch embedding layer for patch embedding processing to obtain a patch feature map; simultaneously inputting the patch feature map into multiple parallel convolutional paths to extract features under different receptive fields; using a multi-head attention mechanism to aggregate features under different receptive fields to obtain aggregated features; concatenating the aggregated features with the patch feature map, then processing the concatenated features through layer normalization and a feedforward network (FFN), and finally fusing them with the output of the FFN through residual connections to obtain the final optical features.
[0013] Furthermore, the SAR branch comprises multiple stacked residual convolutional blocks.
[0014] Furthermore, the hierarchical cross-modal adaptive fusion module includes a cross-stacked upsampling module and an adaptive weight estimation module (AWEM); the feature fusion process of the AWEM includes:
[0015] The SAR features and optical features are stitched together along the channel dimension to obtain the stitched features. The hop connection features of the optical branch and the hop connection features of the SAR branch are extracted and then globally averaged and pooled respectively.
[0016] The pooled splicing features The first combined feature is obtained by combining the skip connection features of the optical branch; the pooled splicing features are then combined. The second combined feature is obtained by combining it with the skip connection feature of the SAR branch;
[0017] The first combined feature and the second combined feature are respectively input into their respective multilayer perceptrons. The outputs of the two multilayer perceptrons are concatenated and then processed by a normalized exponential function to obtain the weights of the optical features and SAR features respectively.
[0018] Based on weights, optical features, SAR features, and stitching features are fused to obtain fused features.
[0019] Furthermore, it also includes: pre-training the optical branch, SAR branch, and hierarchical cross-modal adaptive fusion module jointly, and the loss function used during the training process. for:
[0020]
[0021] in, It is a pixel-level error loss; It is a multi-scale structural similarity loss; It is spectral angle loss.
[0022] Secondly, embodiments of the present invention provide a remote sensing image declouding system based on optical remote sensing images and synthetic aperture radar images, comprising:
[0023] Image acquisition unit, used to acquire paired optical remote sensing images and SAR images;
[0024] The feature extraction unit is used to input the optical remote sensing image into the optical branch to extract optical features, and to input the SAR image into the SAR branch to extract SAR features;
[0025] The fusion unit is used to input the optical features and the SAR features into the hierarchical cross-modal adaptive fusion module to generate adaptive weights based on the semantic and radiometric differences of different modal features, and to fuse the optical features and the SAR features based on the adaptive weights to obtain fused features;
[0026] The reconstruction unit is used to process the fused features through an activation function to obtain a cloudless optical remote sensing image.
[0027] This invention introduces an adaptive weight estimation mechanism to dynamically adjust the contribution of different modal features, taking into account both local detail restoration and global structural consistency, thereby improving the stability and reconstruction quality of cloud removal reconstruction of remote sensing images under complex cloud conditions.
[0028] Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in the first aspect.
[0029] Fourthly, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the first aspect.
[0030] The beneficial effects of this invention are as follows:
[0031] (1) Make full use of the modal complementary information of optical and SAR images. This invention constructs a dual-branch feature extraction structure with optical and SAR branches, respectively, to perform feature modeling on the spectral information of optical images and the structural texture information of SAR images, effectively alleviating the problem of information loss in single-modal cloud removal methods under thick cloud or large-scale cloud coverage conditions.
[0032] (2) A hierarchical cross-modal adaptive fusion mechanism is introduced to improve the rationality of fusion. The optical features and SAR features are dynamically weighted and fused through the adaptive weight estimation module. The contribution of different modal features is adaptively adjusted according to their effectiveness in the current scene, avoiding modal interference and information redundancy caused by fixed weight fusion, and improving the stability and robustness of cross-modal fusion.
[0033] (3) Balancing local detail restoration with global structural consistency. The optical branch introduces a residual multipath aggregation module and a multi-head convolutional self-attention mechanism, which can simultaneously capture local detail features and global contextual information at different scales; combined with the stable structural constraints provided by SAR images, it effectively improves the integrity of the cloud-reconstructed images in terms of edges, textures and ground features.
[0034] (4) Improve the radiometric consistency and spectral fidelity of reconstructed images. By using pixel-level error loss, structural similarity loss and spectral angle loss in combination for model training, the model is guided to pay attention to numerical accuracy, structural similarity and spectral consistency during the reconstruction process, thereby improving the reliability of declouded images in subsequent quantitative analysis and applications. Attached Figure Description
[0035] Figure 1 A flowchart illustrating a remote sensing image declouding method based on optical remote sensing images and synthetic aperture radar images provided in an embodiment of the present invention;
[0036] Figure 2 A framework diagram of the remote sensing image cloud removal model provided in an embodiment of the present invention;
[0037] Figure 3 This is a schematic diagram of the RMPAB structure in the optical branch provided in an embodiment of the present invention;
[0038] Figure 4 This is a schematic diagram of the structure of the AWEM in the hierarchical cross-modal adaptive fusion module provided in an embodiment of the present invention;
[0039] Figure 5 This is a schematic diagram illustrating the cloud removal effect on the SEN12MS-CR dataset provided in an embodiment of the present invention;
[0040] Figure 6 This is a schematic diagram illustrating the cloud removal effect on the LuojiaSET-OSFCR dataset provided in an embodiment of the present invention;
[0041] Figure 7 A schematic diagram illustrating the cloud removal performance metrics for different cloud ranges on the SEN12MS-CR dataset, provided in an embodiment of the present invention.
[0042] Figure 8 This is a schematic diagram illustrating the cloud removal performance metrics for different cloud ranges on the LuojiaSET-OSFCR dataset, provided in an embodiment of the present invention.
[0043] Figure 9 A schematic diagram of a remote sensing image declouding system based on optical remote sensing images and synthetic aperture radar images provided in an embodiment of the present invention;
[0044] Figure 10 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0046] This invention is applicable to the repair of quality degradation problems caused by cloud cover in optical remote sensing images, and can be widely used in scenarios that rely on high-quality remote sensing images, such as urban planning, natural disaster monitoring, and land use surveys.
[0047] like Figure 1As shown, the embodiments of the present invention provide a method for cloud removal from remote sensing images based on optical remote sensing images and synthetic aperture radar images, such as... Figure 2 As shown, this invention constructs a remote sensing image cloud removal model, which includes a dual-branch feature extraction module and a hierarchical cross-modal adaptive fusion module; the method includes the following steps:
[0048] S101: Acquire paired optical remote sensing images and SAR images;
[0049] Specifically, optical remote sensing images may contain multi-band data, and SAR images may be single-polarized or dual-polarized data. Furthermore, to ensure the acquired images conform to the input format for subsequent feature extraction and to improve cloud removal, before step S102, the optical remote sensing images and SAR images should be registered, cropped, and normalized to map pixel values to the [0,1] range, resulting in standardized optical and SAR images. It is understood that, in addition to the above steps, other commonly used enhancement or normalization processes can be applied to the input images to suit the subsequent feature extraction module.
[0050] S102: Input the optical remote sensing image into the optical branch to extract optical features, and input the SAR image into the SAR branch to extract SAR features;
[0051] Specifically, this embodiment uses a dual-branch feature extraction module that includes a SAR branch and an optical branch to extract features from the standardized optical image and the standardized SAR image, respectively.
[0052] The SAR branch comprises multiple residual convolutional blocks. The feature extraction process of this branch includes: inputting a standardized SAR image into the residual convolutional blocks; extracting surface structure and texture features through multiple residual convolutional blocks; and outputting a SAR feature map containing surface structure and spatial texture information. Each residual convolutional block contains a convolutional layer, a batch normalization layer, and a non-linear activation function. The input features are added to the output features through residual connections, thereby enhancing the feature representation ability and mitigating the gradient vanishing problem.
[0053] The optical branch comprises sequentially connected convolutional layers and multiple stacked residual multipath aggregation modules (RMPABs). The feature extraction process of this branch mainly includes: first, inputting a standardized optical image into the convolutional layers to extract initial optical features; then, inputting the initial optical features into the RMPABs, which extract spectral information and local-global contextual features to obtain the optical feature map. .
[0054] In one implementation, such as Figure 3As shown, it includes two equivalent subgraphs, with the left subgraph being an expansion of the right subgraph; the RMPAB feature extraction process includes:
[0055] First, the initial optical features are processed by Patch embedding to obtain a Patch feature map; wherein, Patch embedding refers to dividing the input feature map into several local regions and performing feature mapping.
[0056] Subsequently, multiple (e.g., 3) parallel convolutional paths are used to extract feature information under different receptive fields (e.g., small, medium, and large receptive fields); the output features corresponding to each path are represented as follows:
[0057]
[0058] in, This represents the input Patch feature map. Indicates the first Scale features extracted from the path This indicates feature extraction operations with different receptive fields. In this embodiment, three parallel convolutional paths were used to extract feature information from small, medium, and large receptive fields, respectively.
[0059] Next, a multi-head convolutional self-attention module (MHCA) is used to aggregate features from small, medium, and large receptive fields to enhance the interaction between features at different spatial locations, resulting in aggregated features. The aggregation feature can be represented as:
[0060]
[0061] in, These represent the characteristic information under the small, medium, and large receptive fields, respectively. This indicates an aggregation operation.
[0062] Finally, the aggregated features are concatenated with the patch feature map. The concatenated features are then processed by layer normalization and a feedforward network (FFN), and finally fused with the output of the FFN through residual connections to obtain an optical feature map containing spectral information and local-global context information. .
[0063] S103: Optical feature map and SAR feature map The input is a hierarchical cross-modal adaptive fusion module to generate adaptive weights based on the semantic and radiometric differences of different modal features, and the optical feature map is then processed based on these adaptive weights. and SAR feature map The fusion is performed to obtain a fused feature map. ;
[0064] Specifically, the hierarchical cross-modal adaptive fusion module includes a cross-stacked upsampling module and an adaptive weight estimation module (AWEM); wherein, the upsampling module uses deconvolution layers, interpolation, or the PixelShuffle method to enlarge the feature map to the original input image size. Figure 4 As shown, the feature fusion process of the AWEM includes:
[0065] S1031: SAR features With optical characteristics The splicing is performed along the channel dimension to obtain the spliced features. The data is then subjected to global average pooling (GAP); simultaneously, skip connection features of the optical branches are extracted. Skip connection characteristics of SAR branches Each of these is then subjected to global average pooling (GAP).
[0066] Specifically, in this embodiment, such as Figure 2 As shown, suppose the optical branch includes three residual convolutional modules, and the extracted skip connection features are... This includes the outputs of the first two residual convolutional modules. Similarly, assuming the SAR branch includes one convolutional layer and three RMPABs, the extracted skip connection features... This includes the outputs of the first two RMPABs.
[0067] S1032: Concatenate the pooled features The first combined feature is obtained by combining the skip connection features of the optical branch; the pooled splicing features are then combined. The second combined feature is obtained by combining it with the skip connection feature of the SAR branch;
[0068] S1033: Input the first combined feature and the second combined feature into their respective multilayer perceptrons (MLPs). After concatenating the outputs of the two MLPs, perform normalized exponential function processing to obtain the weights of the optical features and SAR features. and ; Weights representing optical features The weights of SAR features are represented.
[0069] Specifically, the weights are processed using a normalized exponential function to ensure that the fusion weights of different modal features meet preset constraints (such as...). ).
[0070] S1034: Based on weights, optical features, SAR features, and stitching features are fused to obtain a fused feature map. The fused feature map can be represented as:
[0071]
[0072] in, To fuse feature maps, This indicates the fusion strategy.
[0073] For example, a linear weighted fusion strategy can be used, in which case the fusion features can be expressed as:
[0074]
[0075] in, To fuse feature maps;
[0076] It is understandable that, in addition to linear weighted fusion, convolutional fusion, attention mechanisms, or other nonlinear fusion strategies can also be used.
[0077] S104: Merge feature maps The image is mapped to the [0,1] range using an activation function (such as Sigmoid) to obtain a cloudless optical remote sensing image.
[0078] The cloud removal method for remote sensing images provided in this invention has the following advantages:
[0079] (1) Make full use of the modal complementary information of optical and SAR images. This invention constructs a dual-branch feature extraction structure with optical and SAR branches, respectively, to perform feature modeling on the spectral information of optical images and the structural texture information of SAR images, effectively alleviating the problem of information loss in single-modal cloud removal methods under thick cloud or large-scale cloud coverage conditions.
[0080] (2) A hierarchical cross-modal adaptive fusion mechanism is introduced to improve the rationality of fusion. The optical features and SAR features are dynamically weighted and fused through the adaptive weight estimation module. The contribution of different modal features is adaptively adjusted according to their effectiveness in the current scene, avoiding modal interference and information redundancy caused by fixed weight fusion, and improving the stability and robustness of cross-modal fusion.
[0081] (3) Balancing local detail restoration with global structural consistency. The optical branch introduces a residual multipath aggregation module and a multi-head convolutional self-attention mechanism, which can simultaneously capture local detail features and global contextual information at different scales; combined with the stable structural constraints provided by SAR images, it effectively improves the integrity of the cloud-reconstructed images in terms of edges, textures and ground features.
[0082] (4) Improve the radiometric consistency and spectral fidelity of reconstructed images. By using pixel-level error loss, structural similarity loss and spectral angle loss in combination for model training, the model is guided to pay attention to numerical accuracy, structural similarity and spectral consistency during the reconstruction process, thereby improving the reliability of declouded images in subsequent quantitative analysis and applications.
[0083] It should be noted that when using Figure 2 Before applying the model shown to cloud removal from remote sensing images, it needs to be trained to learn how to extract optical and SAR features, how to fuse features from the two modalities, and how to perform cloud removal and reconstruction based on the fused features. The training process uses a joint loss function to optimize model performance.
[0084]
[0085] in, It is a pixel-level error loss that guides the model to approximate a real cloudless image; It is a multi-scale structural similarity loss that preserves the structural integrity of the image. It is spectral angle loss, which improves spectral fidelity.
[0086] In addition, various optimization algorithms (such as AdamW) can be used to train the model during the training process, and parameters such as learning rate, batch size and weight decay can be adjusted according to the implementation example; training stability can be optimized using learning rate scheduling strategies, gradient pruning and other methods.
[0087] To verify the effectiveness of the present invention, the following experiment was conducted:
[0088] 1. Experimental Environment
[0089] Hardware specifications: CPU is Intel(R) Xeon(R) Gold 5218R CPU @ 2.10GHz, memory size is 64GB, GPU is NVIDIA Quadro GV100;
[0090] Software platform: Python version 3.8.20, CUDA version 12.8, and models are built and trained based on the PyTorch 2.2.0 deep learning framework.
[0091] 2. Experimental Dataset
[0092] To evaluate the cloud removal effect of this invention on remote sensing images, the SEN12MS-CR and LuojiaSET-OSFCR datasets were selected for experimental research. Detailed information on the two datasets is shown in Table 1.
[0093] Table 1 Dataset Details
[0094]
[0095] 3. Experimental Setup
[0096] The batch size of the dataset is shown in Table 1. The learning rate is uniformly set to 0.0001, and the AdamW algorithm is used to optimize the training parameters. Each experiment is executed for 50 epochs.
[0097] To comprehensively measure the performance of depth estimation, the following metrics are used for quantitative evaluation:
[0098]
[0099]
[0100]
[0101]
[0102]
[0103] 4. Comparison of different cloud removal models
[0104] Several typical SAR-assisted cloud removal models from recent years were selected and compared in detail with the model of this invention. The results are shown in Table 2.
[0105] Table 2 Comparison of cloud removal performance between the two datasets
[0106]
[0107] As can be seen from the experimental results in the table, the HCFN model proposed in this invention exhibits excellent performance on both datasets, and its metrics are superior to many existing mainstream SAR-assisted cloud removal models.
[0108] Figure 5 This section presents a qualitative comparison of the cloud removal method with five other methods on the SEN12MS-CR dataset. The first three columns represent the visualization effects of the original dataset, namely SAR images, clouded optical images, and true cloudless optical images. The next five columns show the cloud removal effects of the other five methods, and the last column shows the cloud removal effect of the method of this invention. Visually, it can be seen that the method of this invention has the best effect.
[0109] Figure 6 This section presents a qualitative comparison of the cloud removal method with five other methods on the LuojiaSET-OSFCR dataset. The first three columns represent the visualization effects of the original dataset, namely SAR images, clouded optical images, and real cloudless optical images. The next five columns show the cloud removal effects of the other five methods, and the last column shows the cloud removal effect of the method of this invention. Visually, it can be seen that the method of this invention has the best effect.
[0110] Figure 7 This is an analysis of the cloud removal effect at different cloud coverage rates on the SEN12MS-CR dataset. The horizontal axis represents the cloud coverage rate, increasing from 0-20% to 80%-100%, and the vertical axis represents the numerical value of the indicator. The four graphs correspond to four indicators. The upward arrows indicate that the larger the value, the better the effect, and the downward arrows indicate that the smaller the value, the better the effect. The six colors correspond to the five comparison methods and the method of this invention. The thicker the cloud layer, the worse the cloud removal effect.
[0111] Figure 8 This is an analysis of the cloud removal effect at different cloud coverage rates on the LuojiaSET-OSFCR dataset. The horizontal axis represents the cloud coverage rate, increasing from 0-20% to 80%-100%, and the vertical axis represents the numerical value of the indicator. The four graphs correspond to four indicators. The upward arrows indicate that the larger the value, the better the effect, and the downward arrows indicate that the smaller the value, the better the effect. The six colors correspond to the five comparison methods and the method of this invention. The thicker the cloud layer, the worse the cloud removal effect.
[0112] like Figure 9 As shown, this embodiment of the invention provides a remote sensing image declouding system based on optical remote sensing images and synthetic aperture radar images, including: an image acquisition unit, a feature extraction unit, a fusion unit, and a reconstruction unit.
[0113] Specifically, the image acquisition unit is used to acquire paired optical remote sensing images and SAR images; the feature extraction unit is used to input the optical remote sensing images into the optical branch to extract optical features, and input the SAR images into the SAR branch to extract SAR features; the fusion unit is used to input the optical features and the SAR features into the hierarchical cross-modal adaptive fusion module to generate adaptive weights based on the semantic and radiometric differences of different modal features, and to fuse the optical features and the SAR features based on the adaptive weights to obtain fused features; the reconstruction unit is used to process the fused features through an activation function to obtain a cloudless optical remote sensing image.
[0114] It should be noted that the remote sensing image cloud removal system provided in this embodiment of the invention is for implementing the above method. Its specific functions can be referred to in the above method embodiments, and will not be repeated here.
[0115] Figure 10 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 10As shown, the electronic device may include a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004. The processor 1001, communication interface 1002, and memory 1003 communicate with each other via the communication bus 1004. The processor 1001 can call logical instructions in the memory 1003 to execute an adaptive dynamic receptive field infrared target detection method for multi-view scenarios of unmanned aerial vehicles (UAVs). This method includes: constructing a detection network, including: replacing the C3K2 module in YOLOv11 with a multi-dilation rate convolutional module (MFDB), adding a multi-scale pooling attention module (MPA) after the C2PSA module, and adding a multi-scale feature and channel modulation module (MFCM) at the end of the original neck network; inputting the infrared image to be detected into the detection network to obtain detection results, the detection results including the selected target, the target category, and the corresponding probability.
[0116] Furthermore, when the logical instructions in the aforementioned memory 1003 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0117] This invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute a remote sensing image cloud removal method based on optical remote sensing images and synthetic aperture radar images provided in the above-described method embodiments.
[0118] This invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements a remote sensing image cloud removal method based on optical remote sensing images and synthetic aperture radar images provided in the above-described method embodiments.
[0119] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for cloud removal from remote sensing images based on optical remote sensing images and synthetic aperture radar images, characterized in that, include: Step 1: Acquire paired optical remote sensing images and SAR images; Step 2: Input the optical remote sensing image into the optical branch to extract optical features, and input the SAR image into the SAR branch to extract SAR features; Step 3: Input the optical features and the SAR features into the hierarchical cross-modal adaptive fusion module to generate adaptive weights based on the semantic and radiometric differences of different modal features, and fuse the optical features and the SAR features based on the adaptive weights to obtain fused features; Step 4: Process the fused features using an activation function to obtain a cloudless optical remote sensing image.
2. The cloud removal method for remote sensing images based on optical remote sensing images and synthetic aperture radar images according to claim 1, characterized in that, The optical branch includes sequentially connected convolutional layers and multiple stacked residual multipath aggregation modules (RMPABs); wherein, the RMPAB includes a patch embedding layer and a Transformer layer; the Transformer layer includes multiple parallel convolutional paths, a multi-head attention mechanism, layer normalization, and a feedforward network; wherein, the receptive fields of each convolutional path are different; each convolutional path is used to extract features of the feature map processed by the patch embedding layer in its corresponding receptive field; Correspondingly, inputting the optical remote sensing image into the optical branch to extract optical features includes: inputting the optical remote sensing image into a convolutional layer to extract initial optical features; inputting the initial optical features into a patch embedding layer for patch embedding processing to obtain a patch feature map; simultaneously inputting the patch feature map into multiple parallel convolutional paths to extract features under different receptive fields; using a multi-head attention mechanism to aggregate features under different receptive fields to obtain aggregated features; concatenating the aggregated features with the patch feature map, then processing the concatenated features through layer normalization and a feedforward network (FFN), and finally fusing them with the output of the FFN through residual connections to obtain the final optical features.
3. The cloud removal method for remote sensing images based on optical remote sensing images and synthetic aperture radar images according to claim 1, characterized in that, The SAR branch comprises multiple stacked residual convolutional blocks.
4. The cloud removal method for remote sensing images based on optical remote sensing images and synthetic aperture radar images according to claim 1, characterized in that, The hierarchical cross-modal adaptive fusion module includes a cross-stacked upsampling module and an adaptive weight estimation module (AWEM). The feature fusion process of the AWEM includes: The SAR features and optical features are stitched together along the channel dimension to obtain the stitched features. The hop connection features of the optical branch and the hop connection features of the SAR branch are extracted and then globally averaged and pooled respectively. The pooled splicing features The first combined feature is obtained by combining the skip connection features of the optical branch; the pooled splicing features are then combined. The second combined feature is obtained by combining it with the skip connection feature of the SAR branch; The first combined feature and the second combined feature are respectively input into their respective multilayer perceptrons. The outputs of the two multilayer perceptrons are concatenated and then processed by a normalized exponential function to obtain the weights of the optical features and SAR features respectively. Based on weights, optical features, SAR features, and stitching features are fused to obtain fused features.
5. The cloud removal method for remote sensing images based on optical remote sensing images and synthetic aperture radar images according to claim 1, characterized in that, Also includes: The optical branch, SAR branch, and hierarchical cross-modal adaptive fusion module are pre-trained jointly, and the loss function used during training is... for: in, It is a pixel-level error loss; It is a multi-scale structural similarity loss; It is spectral angle loss.
6. A remote sensing image declouding system based on optical remote sensing images and synthetic aperture radar images, characterized in that, include: Image acquisition unit, used to acquire paired optical remote sensing images and SAR images; The feature extraction unit is used to input the optical remote sensing image into the optical branch to extract optical features, and to input the SAR image into the SAR branch to extract SAR features; The fusion unit is used to input the optical features and the SAR features into the hierarchical cross-modal adaptive fusion module to generate adaptive weights based on the semantic and radiometric differences of different modal features, and to fuse the optical features and the SAR features based on the adaptive weights to obtain fused features; The reconstruction unit is used to process the fused features through an activation function to obtain a cloudless optical remote sensing image.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 7.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.